---
title: "MM-RAG (Multimodal RAG) (ES)"
source: "https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)"
wiki: "systems-analysis.info/int"
article: "MM-RAG_(Multimodal_RAG)_(ES)"
language: "es"
categories:
  - "Category:Large language models"
  - "Category:Prompt engineering"
  - "Category:Spanish"
revision_id: 4045
wiki_created_at: 2026-09-06T23:29:42Z
wiki_modified_at: 2026-09-06T23:29:42Z
downloaded_at: 2026-09-07T23:00:32Z
---

# MM-RAG (Multimodal RAG) (ES)

**MM-RAG** (del inglés *Multimodal Retrieval-Augmented Generation*) es una extensión del paradigma clásico de RAG en la que los modelos grandes de lenguaje (LLM) utilizan no solo texto para generar respuestas, sino también datos visuales (imágenes, diagramas, tablas, gráficos). La recuperación multimodal permite encontrar y vincular evidencia en diversas representaciones, reduciendo el riesgo de alucinaciones al basarse en fuentes externas con referencias precisas a fragmentos de páginas y áreas (*bounding boxes*)<sup>[\[1\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_note-lewis2020-1)[\[2\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_note-visrag2024-2)</sup>.

MM-RAG es especialmente útil para documentos donde una parte significativa del significado se presenta en forma no textual (diseño de página, diagramas, estructuras de tablas). En tales casos, el RAG textual clásico a menudo pierde elementos importantes del contexto<sup>[\[3\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_note-docvqa2021-3)[\[4\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_note-textvqa2019-4)</sup>.

## Contexto y problema que resuelve

El RAG clásico opera con pasajes de texto y no percibe las estructuras visuales (la disposición de los elementos, los pies de foto, los ejes de los gráficos). MM-RAG cierra estas brechas: extrae elementos estructurados (texto, tablas, imágenes con coordenadas), los indexa en un espacio vectorial y combina evidencia de diferentes modalidades<sup>[\[5\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_note-layoutlm2020-5)[\[6\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_note-clip2021-6)</sup>.

## Arquitectura de MM-RAG

El pipeline de MM-RAG añade al RAG clásico etapas de procesamiento de datos visuales y alineación de modalidades: **ingesta → indexación → recuperación multimodal → fusión y reclasificación → generación con trazabilidad**.

1.  **Recopilación y preprocesamiento (Ingestion).** La entrada consiste en PDF/escaneos/imágenes. Se realiza OCR y **análisis de diseño** (layout analysis) de la página para identificar zonas: párrafos, encabezados, tablas, imágenes y sus coordenadas. Las herramientas típicas son los modelos de la familia LayoutLM y las bibliotecas instrumentales como LayoutParser; la validación y el entrenamiento a menudo se basan en los conjuntos de datos PubLayNet y DocLayNet<sup>[\[7\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_note-publaynet2019-7)[\[8\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_note-doclaynet2022-8)[\[9\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_note-layoutparser2021-9)</sup>.
2.  **Segmentación en regiones.** Se extraen objetos visuales (diagramas, tablas, ilustraciones, pies de foto). Para una mayor robustez, se utilizan modelos sin OCR (OCR-free, por ejemplo, Donut) o pipelines combinados OCR+VLM<sup>[\[10\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_note-donut2021-10)</sup>.
3.  **Indexación (Vector Index).** Los fragmentos de texto (chunks) y los elementos visuales (imágenes o sus descripciones) se convierten en representaciones vectoriales y se almacenan en una base de datos vectorial. Para el espacio unificado *texto↔imagen* se utilizan CLIP o SigLIP; para producción son convenientes los índices multimodales/multivectoriales (un objeto, varios vectores)<sup>[\[6\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_note-clip2021-6)[\[11\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_note-siglip2023-11)[\[12\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_note-milvus-mv-12)</sup>.
4.  **Recuperación multimodal y reclasificación.** Se realiza una combinación de búsqueda textual y visual; los candidatos (párrafos, tablas, imágenes/regiones) se fusionan y se reclasifican con un modelo más "pesado" (cross-encoder/LLM-reranker) para mejorar la precisión<sup>[\[13\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_note-cohere-rerank-13)</sup>.
5.  **Empaquetado del contexto y generación.** Los fragmentos seleccionados se introducen en un LLM/VLM. Si el modelo es multimodal (por ejemplo, GPT-4V/4o), las imágenes pueden ser ingresadas directamente; si se utiliza un LLM textual, las imágenes se convierten previamente en descripciones detalladas<sup>[\[14\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_note-gpt4v-14)[\[15\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_note-gpt4o-15)</sup>.
6.  **Trazabilidad y citación.** La respuesta se acompaña de citas clicables vinculadas no solo al documento/página, sino también a la región específica (coordenadas). Esto aumenta el nivel de *grounding* (anclaje en la evidencia) y la confianza del usuario<sup>[\[2\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_note-visrag2024-2)</sup>.

## Evaluación de calidad y métricas

La eficacia de MM‑RAG se evalúa en los niveles de extracción, recuperación y generación.

- **Calidad de la extracción de datos visuales.** Precisión del OCR (WER/CER), calidad del análisis de diseño (mAP/Precision/Recall) en los conjuntos de datos DocLayNet/PubLayNet<sup>[\[8\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_note-doclaynet2022-8)[\[7\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_note-publaynet2019-7)</sup>.
- **Calidad de la recuperación.** Métricas estándar de recuperación de información (IR): Recall@K, Precision@K, MRR; para la multimodalidad, se evalúan por separado para cada modalidad y de forma combinada.
- **Calidad de la respuesta (end‑to‑end).** Métricas automáticas de *faithfulness* (fidelidad)/*groundedness* (anclaje en la evidencia) y evaluación humana. En la práctica, se utilizan frameworks como RAGAS/TruLens/DeepEval<sup>[\[16\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_note-ragas-16)</sup>.
- **Benchmarks.**
  - **DocVQA**: preguntas sobre imágenes de documentos<sup>[\[3\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_note-docvqa2021-3)</sup>.
  - **TextVQA**: preguntas que requieren leer texto dentro de las imágenes<sup>[\[4\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_note-textvqa2019-4)</sup>.
  - **InfographicVQA**: preguntas sobre infografías<sup>[\[17\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_note-infographicvqa2021-17)</sup>.
  - **ChartQA**: preguntas sobre diagramas que requieren razonamiento lógico<sup>[\[18\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_note-chartqa2022-18)</sup>.
  - **MMDocRAG**: un benchmark de RAG multimodal para DocQA (documentos de varias páginas, cadenas de evidencia cross-modales)<sup>[\[19\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_note-mmdocrag2025-19)</sup>.

## Tabla comparativa de componentes

| Componente               | Variantes de implementación                                               | Ventajas                                                                                                        | Desventajas / Riesgos                                                                              | Cuándo elegir                                                                       |
|--------------------------|---------------------------------------------------------------------------|-----------------------------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------|-------------------------------------------------------------------------------------|
| OCR                      | Tesseract / PaddleOCR / API en la nube                                    | Locales: privacidad y control; en la nube: alta precisión "de fábrica".                                         | Errores en diseños complejos; API: costo y requisitos de cumplimiento (compliance).                | Datos privados: OCR local; máxima precisión: en la nube (si está permitido).        |
| Análisis de diseño       | Reglas / Modelo de ML (LayoutLM, LayoutParser)                            | Las reglas son simples para plantillas uniformes; el ML es robusto ante la diversidad.                          | Las reglas fallan con nuevos diseños; el ML requiere recursos/datos.                               | Formularios estandarizados: reglas; corpus heterogéneo: ML.                         |
| Vectorización (imágenes) | CLIP / SigLIP / descripciones sin OCR (Donut/Pix2Struct)                  | Espacio latente común *texto↔imagen* (CLIP/SigLIP); los modelos sin OCR eliminan la dependencia del OCR.        | CLIP no lee el texto dentro de las imágenes; las descripciones pueden distorsionar el significado. | CLIP/SigLIP para búsqueda multimodal básica; sin OCR para escaneos de baja calidad. |
| Fusión de resultados     | Ordenamiento por puntuación (score) / cuotas por modalidad / LLM-reranker | El reclasificador (reranker) aumenta notablemente la precisión en la selección del contexto.                    | Aumento de la latencia y el costo.                                                                 | Escenarios de alta precisión; métodos simples para pruebas de concepto (PoC).       |
| Almacenamiento/Índice    | Vector único / multivector (texto+imagen) / híbrido (BM25+vector)         | El multivector cubre diferentes representaciones de un mismo objeto; el híbrido rescata palabras clave/códigos. | Complejidad del esquema y las actualizaciones.                                                     | Sistemas de producción con datos mixtos y SLAs estrictos.                           |

Comparación de componentes y enfoques clave en MM‑RAG

## Consideraciones prácticas

- **Búsqueda híbrida** (BM25 + vector): es el estándar de facto para mejorar la exhaustividad (recall) y la precisión con términos/códigos específicos<sup>[\[20\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_note-weaviate-hybrid-20)</sup>.
- **Reclasificación** (reranking) con un cross-encoder/LLM: ahorra tokens al descartar candidatos "basura" antes de la generación<sup>[\[13\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_note-cohere-rerank-13)</sup>.
- **Recuperadores VLM modernos** (por ejemplo, ColPali): demuestran ventajas en documentos visualmente ricos gracias a la indexación directa de las páginas como imágenes<sup>[\[21\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_note-colpali2024-21)</sup>.

## Véase también

- [Retrieval-Augmented Generation](https://systems-analysis.info/int/Retrieval-augmented_generation_(RAG)_(ES) "Retrieval-augmented generation (RAG) (ES)")
- [Base de datos vectorial](https://systems-analysis.info/int/Base_de_datos_vectorial "Base de datos vectorial")
- [Embedding](https://systems-analysis.info/int/Embedding_(NLP)_(ES) "Embedding (NLP) (ES)")

## Bibliografía

- Lewis, P. et al. (2020). *Retrieval‑Augmented Generation for Knowledge‑Intensive NLP Tasks*. NeurIPS. <a href="https://arxiv.org/abs/2005.11401" class="external text" rel="nofollow">arXiv:2005.11401</a>.
- Gao, L. et al. (2023). *Precise Zero‑Shot Dense Retrieval without Relevance Labels (HyDE)*. ACL 2023. <a href="https://arxiv.org/abs/2212.10496" class="external text" rel="nofollow">arXiv:2212.10496</a>.
- Mei, L., Mo, S., Yang, Z., Chen, C. (2025). *A Survey of Multimodal Retrieval‑Augmented Generation*. <a href="https://arxiv.org/abs/2504.08748" class="external text" rel="nofollow">arXiv:2504.08748</a>.
- Abootorabi, M.M. et al. (2025). *Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval‑Augmented Generation*. Findings of ACL 2025. <a href="https://aclanthology.org/2025.findings-acl.861/" class="external text" rel="nofollow">ACL Anthology</a>.
- Yu, S. et al. (2024). *VisRAG: Vision‑based Retrieval‑augmented Generation on Multi‑modality Documents*. <a href="https://arxiv.org/abs/2410.10594" class="external text" rel="nofollow">arXiv:2410.10594</a>.
- Cho, J. et al. (2024). *M3DocRAG: Multi‑modal Retrieval is What You Need for Multi‑document QA*. <a href="https://arxiv.org/abs/2411.04952" class="external text" rel="nofollow">arXiv:2411.04952</a>.
- Tanaka, R. et al. (2025). *VDocRAG: Retrieval‑Augmented Generation over Visually‑Rich Documents*. CVPR 2025. <a href="https://arxiv.org/abs/2504.09795" class="external text" rel="nofollow">arXiv:2504.09795</a> • <a href="https://openaccess.thecvf.com/content/CVPR2025/papers/Tanaka_VDocRAG_Retrieval-Augmented_Generation_over_Visually-Rich_Documents_CVPR_2025_paper.pdf" class="external text" rel="nofollow">CVF Open Access</a>.
- Dong, K. et al. (2025). *MMDocRAG: Benchmarking Retrieval‑Augmented Multimodal Generation for Document Question Answering*. <a href="https://arxiv.org/abs/2505.16470" class="external text" rel="nofollow">arXiv:2505.16470</a>.
- Wasserman, N. et al. (2025). *REAL‑MM‑RAG: A Real‑World Multi‑Modal Retrieval Benchmark*. <a href="https://arxiv.org/abs/2502.12342" class="external text" rel="nofollow">arXiv:2502.12342</a>.
- Faysse, M. et al. (2024). *ColPali: Efficient Document Retrieval with Vision Language Models*. <a href="https://arxiv.org/abs/2407.01449" class="external text" rel="nofollow">arXiv:2407.01449</a>.
- Radford, A. et al. (2021). *Learning Transferable Visual Models From Natural Language Supervision (CLIP)*. ICML 2021. <a href="https://arxiv.org/abs/2103.00020" class="external text" rel="nofollow">arXiv:2103.00020</a>.
- Tschannen, M. et al. (2025). *SigLIP 2: Multilingual Vision‑Language Encoders with Improved Semantic Understanding, Localization, and Dense Features*. <a href="https://arxiv.org/abs/2502.14786" class="external text" rel="nofollow">arXiv:2502.14786</a>.
- Xu, Y. et al. (2020). *LayoutLM: Pre‑training of Text and Layout for Document Image Understanding*. KDD 2020. <a href="https://dl.acm.org/doi/10.1145/3394486.3403172" class="external text" rel="nofollow">DOI</a> • <a href="https://arxiv.org/abs/1912.13318" class="external text" rel="nofollow">arXiv:1912.13318</a>.
- Huang, Y. et al. (2022). *LayoutLMv3: Pre‑training for Document AI with Unified Text and Image Masking*. <a href="https://arxiv.org/abs/2204.08387" class="external text" rel="nofollow">arXiv:2204.08387</a>.
- Zhong, X., Tang, J., Jimeno‑Yepes, A.J. (2019). *PubLayNet: Largest Dataset Ever for Document Layout Analysis*. <a href="https://arxiv.org/abs/1908.07836" class="external text" rel="nofollow">arXiv:1908.07836</a>.
- Pfitzmann, B. et al. (2022). *DocLayNet: A Large Human‑Annotated Dataset for Document‑Layout Analysis*. <a href="https://arxiv.org/abs/2206.01062" class="external text" rel="nofollow">arXiv:2206.01062</a>.
- Kim, G. et al. (2021). *OCR‑free Document Understanding Transformer (Donut)*. <a href="https://arxiv.org/abs/2111.15664" class="external text" rel="nofollow">arXiv:2111.15664</a>.
- Shen, Z. et al. (2021). *LayoutParser: A Unified Toolkit for Deep Learning Based Document Image Analysis*. <a href="https://arxiv.org/abs/2103.15348" class="external text" rel="nofollow">arXiv:2103.15348</a>.
- Singh, A. et al. (2019). *Towards VQA Models That Can Read (TextVQA)*. CVPR 2019. <a href="https://arxiv.org/abs/1904.08920" class="external text" rel="nofollow">arXiv:1904.08920</a>.
- Mathew, M. et al. (2021/2022). *DocVQA / InfographicVQA: Datasets for VQA on Document Images and Infographics*. WACV 2021 / WACV 2022. <a href="https://openaccess.thecvf.com/content/WACV2021/html/Mathew_DocVQA_A_Dataset_for_VQA_on_Document_Images_WACV_2021_paper.html" class="external text" rel="nofollow">CVF</a> • <a href="https://arxiv.org/abs/2104.12756" class="external text" rel="nofollow">arXiv:2104.12756</a>.
- Masry, A. et al. (2022). *ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning*. Findings of ACL 2022. <a href="https://aclanthology.org/2022.findings-acl.177/" class="external text" rel="nofollow">ACL</a> • <a href="https://arxiv.org/abs/2203.10244" class="external text" rel="nofollow">arXiv:2203.10244</a>.
- Liu, F. et al. (2022). *DePlot: One‑shot Visual Language Reasoning by Plot‑to‑Table Translation*. <a href="https://arxiv.org/abs/2212.10505" class="external text" rel="nofollow">arXiv:2212.10505</a>.
- Wang, P. et al. (2024). *Qwen2‑VL: Enhancing Vision‑Language Model’s Capabilities in OCR and Chart QA*. <a href="https://arxiv.org/abs/2409.12191" class="external text" rel="nofollow">arXiv:2409.12191</a>.
- Es, S. et al. (2023). *RAGAS: Automated Evaluation of Retrieval‑Augmented Generation*. <a href="https://arxiv.org/abs/2309.15217" class="external text" rel="nofollow">arXiv:2309.15217</a>.

## Referencias

1.  <span id="cite_note-lewis2020-1">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_ref-lewis2020_1-0) Lewis, P., Perez, E., et al. (2020). *Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks*. NeurIPS. <a href="https://arxiv.org/abs/2005.11401" class="external text" rel="nofollow">arXiv:2005.11401</a>.</span>
2.  <span id="cite_note-visrag2024-2">↑ <sup>[2.0](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_ref-visrag2024_2-0)</sup> <sup>[2.1](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_ref-visrag2024_2-1)</sup> Yu, S. et al. (2024). *VisRAG: Vision-based Retrieval-Augmented Generation on Multi-modality Documents*. <a href="https://arxiv.org/abs/2407.06437" class="external text" rel="nofollow">arXiv:2407.06437</a>.</span>
3.  <span id="cite_note-docvqa2021-3">↑ <sup>[3.0](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_ref-docvqa2021_3-0)</sup> <sup>[3.1](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_ref-docvqa2021_3-1)</sup> Mathew, M. et al. (2021). *DocVQA: A Dataset for VQA on Document Images*. WACV. <a href="https://arxiv.org/abs/2007.00398" class="external text" rel="nofollow">arXiv:2007.00398</a>.</span>
4.  <span id="cite_note-textvqa2019-4">↑ <sup>[4.0](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_ref-textvqa2019_4-0)</sup> <sup>[4.1](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_ref-textvqa2019_4-1)</sup> Singh, A. et al. (2019). *TextVQA: Towards VQA Models That Can Read*. CVPR. <a href="https://arxiv.org/abs/1904.08920" class="external text" rel="nofollow">arXiv:1904.08920</a>.</span>
5.  <span id="cite_note-layoutlm2020-5">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_ref-layoutlm2020_5-0) Xu, Y. et al. (2020). *LayoutLM: Pre-training of Text and Layout for Document Image Understanding*. KDD. <a href="https://dl.acm.org/doi/10.1145/3394486.3403172" class="external text" rel="nofollow">DOI</a>.</span>
6.  <span id="cite_note-clip2021-6">↑ <sup>[6.0](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_ref-clip2021_6-0)</sup> <sup>[6.1](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_ref-clip2021_6-1)</sup> Radford, A. et al. (2021). *Learning Transferable Visual Models From Natural Language Supervision (CLIP)*. ICML. <a href="https://arxiv.org/abs/2103.00020" class="external text" rel="nofollow">arXiv:2103.00020</a>.</span>
7.  <span id="cite_note-publaynet2019-7">↑ <sup>[7.0](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_ref-publaynet2019_7-0)</sup> <sup>[7.1](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_ref-publaynet2019_7-1)</sup> Zhong, X., Tang, J., Yepes, A. J. (2019). *PubLayNet: Largest Dataset for Document Layout Analysis*. ICDAR. <a href="https://arxiv.org/abs/1908.07836" class="external text" rel="nofollow">arXiv:1908.07836</a>.</span>
8.  <span id="cite_note-doclaynet2022-8">↑ <sup>[8.0](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_ref-doclaynet2022_8-0)</sup> <sup>[8.1](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_ref-doclaynet2022_8-1)</sup> Pfitzmann, B. et al. (2022). *DocLayNet: A Large Human‑Annotated Dataset for Document‑Layout Analysis*. KDD. <a href="https://dl.acm.org/doi/10.1145/3534678.3539043" class="external text" rel="nofollow">DOI</a> / <a href="https://arxiv.org/abs/2206.01062" class="external text" rel="nofollow">arXiv:2206.01062</a>.</span>
9.  <span id="cite_note-layoutparser2021-9">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_ref-layoutparser2021_9-0) Shen, Z. et al. (2021). *LayoutParser: A Unified Toolkit for DL‑based Document Image Analysis*. <a href="https://arxiv.org/abs/2103.15348" class="external text" rel="nofollow">arXiv:2103.15348</a>.</span>
10. <span id="cite_note-donut2021-10">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_ref-donut2021_10-0) Kim, G. et al. (2021). *Donut: OCR‑free Document Understanding Transformer*. <a href="https://arxiv.org/abs/2111.15664" class="external text" rel="nofollow">arXiv:2111.15664</a>.</span>
11. <span id="cite_note-siglip2023-11">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_ref-siglip2023_11-0) Zhai, X. et al. (2023). *Sigmoid Loss for Language‑Image Pre‑Training (SigLIP)*. ICCV. <a href="https://arxiv.org/abs/2303.15343" class="external text" rel="nofollow">arXiv:2303.15343</a>.</span>
12. <span id="cite_note-milvus-mv-12">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_ref-milvus-mv_12-0) Milvus Docs. *Multi‑Vector Hybrid Search*. <a href="https://milvus.io/docs/multi-vector-search.md" class="external text" rel="nofollow">milvus.io/docs/multi-vector-search.md</a>.</span>
13. <span id="cite_note-cohere-rerank-13">↑ <sup>[13.0](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_ref-cohere-rerank_13-0)</sup> <sup>[13.1](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_ref-cohere-rerank_13-1)</sup> Cohere Docs. *Rerank API*. <a href="https://docs.cohere.com/reference/rerank" class="external text" rel="nofollow">docs.cohere.com/reference/rerank</a>.</span>
14. <span id="cite_note-gpt4v-14">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_ref-gpt4v_14-0) OpenAI. *GPT‑4V(ision) System Card*. (2023). <a href="https://cdn.openai.com/papers/GPTV_System_Card.pdf" class="external text" rel="nofollow">PDF</a>.</span>
15. <span id="cite_note-gpt4o-15">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_ref-gpt4o_15-0) OpenAI. *Hello GPT‑4o*. (2024). <a href="https://openai.com/index/hello-gpt-4o/" class="external text" rel="nofollow">openai.com/index/hello-gpt-4o/</a>.</span>
16. <span id="cite_note-ragas-16">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_ref-ragas_16-0) Es, S. et al. (2023). *RAGAS: Automated Evaluation of Retrieval Augmented Generation*. <a href="https://arxiv.org/abs/2309.15217" class="external text" rel="nofollow">arXiv:2309.15217</a>.</span>
17. <span id="cite_note-infographicvqa2021-17">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_ref-infographicvqa2021_17-0) Mathew, M. et al. (2021). *InfographicVQA: Understanding Infographics via Question Answering*. ICDAR. <a href="https://arxiv.org/abs/2104.12756" class="external text" rel="nofollow">arXiv:2104.12756</a>.</span>
18. <span id="cite_note-chartqa2022-18">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_ref-chartqa2022_18-0) Masry, A. et al. (2022). *ChartQA: A Benchmark for Question Answering about Charts*. ACL (Findings). <a href="https://arxiv.org/abs/2103.16435" class="external text" rel="nofollow">arXiv:2103.16435</a>.</span>
19. <span id="cite_note-mmdocrag2025-19">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_ref-mmdocrag2025_19-0) Dong, K. et al. (2025). *Benchmarking Retrieval‑Augmented Multimodal Generation for Document QA (MMDocRAG)*. <a href="https://arxiv.org/abs/2505.16470" class="external text" rel="nofollow">arXiv:2505.16470</a>.</span>
20. <span id="cite_note-weaviate-hybrid-20">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_ref-weaviate-hybrid_20-0) Weaviate Docs. *Hybrid search*. <a href="https://docs.weaviate.io/weaviate/concepts/search/hybrid-search" class="external text" rel="nofollow">docs.weaviate.io/.../hybrid-search</a>.</span>
21. <span id="cite_note-colpali2024-21">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES)#cite_ref-colpali2024_21-0) Faysse, M. et al. (2024). *ColPali: Efficient Document Retrieval with Vision‑Language Models*. <a href="https://arxiv.org/abs/2407.01449" class="external text" rel="nofollow">arXiv:2407.01449</a>.</span>
