---
title: "MM-RAG (Multimodales RAG)"
source: "https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)"
wiki: "systems-analysis.info/int"
article: "MM-RAG_(Multimodales_RAG)"
language: "de"
categories:
  - "Category:German"
  - "Category:Large language models"
  - "Category:Prompt engineering"
revision_id: 4067
wiki_created_at: 2026-09-06T23:30:02Z
wiki_modified_at: 2026-09-06T23:30:02Z
downloaded_at: 2026-09-07T23:00:42Z
---

# MM-RAG (Multimodales RAG)

**MM-RAG** (von engl. *Multimodal Retrieval-Augmented Generation*) ist eine Erweiterung des klassischen RAG-Paradigmas, bei dem [LLMs](https://systems-analysis.info/int/Gro%C3%9Fe_Sprachmodelle "Große Sprachmodelle") zur Beantwortung von Anfragen nicht nur Text, sondern auch visuelle Daten (Bilder, Diagramme, Tabellen, Grafiken) verwenden. Das multimodale Retrieval ermöglicht es, Belege in verschiedenen Darstellungsformen zu finden und zu verknüpfen, was das Risiko von Halluzinationen reduziert, da es sich auf externe Quellen mit präzisen Verweisen auf Seitenfragmente und Bereiche (*Bounding Boxes*) stützt<sup>[\[1\]](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_note-lewis2020-1)[\[2\]](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_note-visrag2024-2)</sup>.

MM-RAG ist besonders nützlich für Dokumente, bei denen ein wesentlicher Teil der Bedeutung in nicht-textueller Form vorliegt (Seitenlayout, Diagramme, Tabellenstrukturen). In solchen Fällen verliert das klassische textbasierte RAG oft wichtige Kontextelemente<sup>[\[3\]](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_note-docvqa2021-3)[\[4\]](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_note-textvqa2019-4)</sup>.

## Kontext und Problemstellung

Klassisches RAG arbeitet mit Textpassagen und erkennt keine visuellen Strukturen (Anordnung von Elementen, Bildunterschriften, Achsen von Grafiken). MM-RAG schließt diese Lücken: Es extrahiert strukturierte Elemente (Text, Tabellen, Bilder mit Koordinaten), indexiert sie in einem Vektorraum und kombiniert Belege aus verschiedenen Modalitäten<sup>[\[5\]](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_note-layoutlm2020-5)[\[6\]](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_note-clip2021-6)</sup>.

## Architektur von MM-RAG

Die MM-RAG-Pipeline erweitert das klassische RAG um Schritte zur Verarbeitung visueller Daten und zur Angleichung der Modalitäten: **Ingestion → Indexierung → Multimodales Retrieval → Fusion und Reranking → Generierung mit Tracing**.

1.  **Ingestion und Vorverarbeitung.** Als Eingabe dienen PDF-Dateien, Scans oder Bilder. Es werden OCR und eine **Layoutanalyse** der Seite durchgeführt, um Bereiche wie Absätze, Überschriften, Tabellen, Bilder und deren Koordinaten zu identifizieren. Typische Werkzeuge sind Modelle der LayoutLM-Familie und Bibliotheken wie LayoutParser; die Validierung und das Training stützen sich oft auf die Datensätze PubLayNet und DocLayNet<sup>[\[7\]](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_note-publaynet2019-7)[\[8\]](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_note-doclaynet2022-8)[\[9\]](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_note-layoutparser2021-9)</sup>.
2.  **Segmentierung in Regionen.** Visuelle Objekte (Diagramme, Tabellen, Illustrationen, Bildunterschriften) werden extrahiert. Für eine höhere Robustheit werden OCR-freie Modelle (z. B. Donut) oder kombinierte OCR+VLM-Pipelines eingesetzt<sup>[\[10\]](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_note-donut2021-10)</sup>.
3.  **Indexierung (Vector Index).** Text-Chunks und visuelle Elemente (Bilder oder deren Beschreibungen) werden in Vektoreinbettungen umgewandelt und in einer Vektordatenbank gespeichert. Für den kombinierten *Text↔Bild*-Raum werden CLIP oder SigLIP verwendet; für Produktionsumgebungen eignen sich multimodale/multivektorale Indizes (ein Objekt – mehrere Vektoren)<sup>[\[6\]](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_note-clip2021-6)[\[11\]](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_note-siglip2023-11)[\[12\]](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_note-milvus-mv-12)</sup>.
4.  **Multimodales Retrieval und Reranking.** Es wird eine Kombination aus textbasierter und visueller Suche durchgeführt; die Kandidaten (Absätze, Tabellen, Bilder/Regionen) werden zusammengeführt und von einem komplexeren Modell (Cross-Encoder/LLM-Reranker) neu geordnet, um die Genauigkeit zu erhöhen<sup>[\[13\]](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_note-cohere-rerank-13)</sup>.
5.  **Kontext-Zusammenstellung und Generierung.** Die ausgewählten Fragmente werden an ein LLM/VLM übergeben. Wenn das Modell multimodal ist (z. B. GPT‑4V/4o), können Bilder direkt eingegeben werden; bei einem rein textbasierten LLM werden Bilder vorab in detaillierte Beschreibungen umgewandelt<sup>[\[14\]](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_note-gpt4v-14)[\[15\]](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_note-gpt4o-15)</sup>.
6.  **Tracing und Zitation.** Die Antwort wird von klickbaren Zitaten begleitet, die nicht nur auf das Dokument/die Seite verweisen, sondern auch auf eine bestimmte Region (Koordinaten). Dies erhöht das *Grounding* und das Vertrauen der Nutzer<sup>[\[2\]](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_note-visrag2024-2)</sup>.

## Qualitätsbewertung und Metriken

Die Effektivität von MM-RAG wird auf den Ebenen der Extraktion, des Retrievals und der Generierung bewertet.

- **Qualität der Extraktion visueller Daten.** OCR-Genauigkeit (WER/CER), Qualität der Layoutanalyse (mAP/Precision/Recall) auf Datensätzen wie DocLayNet/PubLayNet<sup>[\[8\]](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_note-doclaynet2022-8)[\[7\]](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_note-publaynet2019-7)</sup>.
- **Qualität des Retrievals.** Standardmetriken des Information Retrieval: Recall@K, Precision@K, MRR; für die Multimodalität werden diese separat für die einzelnen Modalitäten und deren Kombination bewertet.
- **Qualität der Antwort (End-to-End).** Automatische Metriken wie *Faithfulness*/*Groundedness* und menschliche Bewertung. In der Praxis werden Frameworks wie RAGAS/TruLens/DeepEval verwendet<sup>[\[16\]](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_note-ragas-16)</sup>.
- **Benchmarks.**
  - **DocVQA**: Fragen zu Dokumentenbildern<sup>[\[3\]](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_note-docvqa2021-3)</sup>.
  - **TextVQA**: Fragen, die das Lesen von Text in Bildern erfordern<sup>[\[4\]](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_note-textvqa2019-4)</sup>.
  - **InfographicVQA**: Fragen zu Infografiken<sup>[\[17\]](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_note-infographicvqa2021-17)</sup>.
  - **ChartQA**: Fragen zu Diagrammen, die logisches Denken erfordern<sup>[\[18\]](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_note-chartqa2022-18)</sup>.
  - **MMDocRAG**: Ein Benchmark für multimodales RAG für DocQA (mehrseitige Dokumente, modalitätsübergreifende Beweisketten)<sup>[\[19\]](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_note-mmdocrag2025-19)</sup>.

## Vergleichstabelle der Komponenten

| Komponente              | Implementierungsvarianten                                         | Vorteile                                                                                                                  | Nachteile / Risiken                                                                            | Wann zu wählen                                                                                           |
|-------------------------|-------------------------------------------------------------------|---------------------------------------------------------------------------------------------------------------------------|------------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------|
| OCR                     | Tesseract / PaddleOCR / Cloud-APIs                                | Lokale Lösungen bieten Datenschutz und Kontrolle; Cloud-Lösungen bieten hohe Genauigkeit "out of the box".                | Fehler bei komplexem Layout; APIs verursachen Kosten und unterliegen Compliance-Anforderungen. | Bei vertraulichen Daten – lokale OCR; für maximale Genauigkeit – Cloud (falls zulässig).                 |
| Layoutanalyse           | Regelbasiert / ML-Modell (LayoutLM, LayoutParser)                 | Regeln sind für einheitliche Vorlagen einfach; ML ist robust gegenüber Variationen.                                       | Regeln versagen bei neuen Layouts; ML erfordert Ressourcen/Daten.                              | Einheitliche Formulare – regelbasiert; heterogener Korpus – ML.                                          |
| Vektorisierung (Bilder) | CLIP / SigLIP / OCR-freie Beschreibungen (Donut/Pix2Struct)       | Gemeinsamer latenter Raum *Text↔Bild* (CLIP/SigLIP); OCR-freie Ansätze eliminieren die Abhängigkeit von OCR.              | CLIP liest keinen Text innerhalb von Bildern; Beschreibungen können die Bedeutung verfälschen. | CLIP/SigLIP für die grundlegende multimodale Suche; OCR-freie Ansätze für Scans von schlechter Qualität. |
| Fusion der Ergebnisse   | Sortierung nach Score / Quoten pro Modalität / LLM-Reranker       | Ein Reranker verbessert die Genauigkeit der Kontextvariable deutlich.                                                     | Erhöhte Latenz und Kosten.                                                                     | Szenarien mit hohen Präzisionsanforderungen; einfache Methoden für PoCs.                                 |
| Speicher/Index          | Einzelner Vektor / Multivektor (Text+Bild) / Hybrid (BM25+Vektor) | Ein Multivektor deckt verschiedene Repräsentationen eines Objekts ab; hybride Ansätze sind gut für Schlüsselwörter/Codes. | Komplexere Architektur und Updates.                                                            | Produktionssysteme mit gemischten Daten und strengen SLAs.                                               |

Vergleich der Schlüsselkomponenten und Ansätze in MM-RAG

## Praktische Hinweise

- **Hybride Suche** (BM25 + Vektor) ist der De-facto-Standard zur Verbesserung von Recall und Präzision bei spezifischen Begriffen/Codes<sup>[\[20\]](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_note-weaviate-hybrid-20)</sup>.
- **Reranking** mit einem Cross-Encoder/LLM spart Tokens, indem irrelevante Kandidaten vor der Generierungsphase verworfen werden<sup>[\[13\]](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_note-cohere-rerank-13)</sup>.
- **Moderne VLM-Retriever** (z. B. ColPali) zeigen Vorteile bei visuell reichhaltigen Dokumenten durch die direkte Indexierung von Bildseiten<sup>[\[21\]](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_note-colpali2024-21)</sup>.

## Siehe auch

- Retrieval-Augmented Generation
- [Vektordatenbank](https://systems-analysis.info/int/Vektordatenbanken "Vektordatenbanken")
- Embedding
- GraphRAG

## Literatur

- Lewis, P. et al. (2020). *Retrieval‑Augmented Generation for Knowledge‑Intensive NLP Tasks*. NeurIPS. <a href="https://arxiv.org/abs/2005.11401" class="external text" rel="nofollow">arXiv:2005.11401</a>.
- Gao, L. et al. (2023). *Precise Zero‑Shot Dense Retrieval without Relevance Labels (HyDE)*. ACL 2023. <a href="https://arxiv.org/abs/2212.10496" class="external text" rel="nofollow">arXiv:2212.10496</a>.
- Mei, L., Mo, S., Yang, Z., Chen, C. (2025). *A Survey of Multimodal Retrieval‑Augmented Generation*. <a href="https://arxiv.org/abs/2504.08748" class="external text" rel="nofollow">arXiv:2504.08748</a>.
- Abootorabi, M.M. et al. (2025). *Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval‑Augmented Generation*. Findings of ACL 2025. <a href="https://aclanthology.org/2025.findings-acl.861/" class="external text" rel="nofollow">ACL Anthology</a>.
- Yu, S. et al. (2024). *VisRAG: Vision‑based Retrieval‑augmented Generation on Multi‑modality Documents*. <a href="https://arxiv.org/abs/2410.10594" class="external text" rel="nofollow">arXiv:2410.10594</a>.
- Cho, J. et al. (2024). *M3DocRAG: Multi‑modal Retrieval is What You Need for Multi‑document QA*. <a href="https://arxiv.org/abs/2411.04952" class="external text" rel="nofollow">arXiv:2411.04952</a>.
- Tanaka, R. et al. (2025). *VDocRAG: Retrieval‑Augmented Generation over Visually‑Rich Documents*. CVPR 2025. <a href="https://arxiv.org/abs/2504.09795" class="external text" rel="nofollow">arXiv:2504.09795</a> • <a href="https://openaccess.thecvf.com/content/CVPR2025/papers/Tanaka_VDocRAG_Retrieval-Augmented_Generation_over_Visually-Rich_Documents_CVPR_2025_paper.pdf" class="external text" rel="nofollow">CVF Open Access</a>.
- Dong, K. et al. (2025). *MMDocRAG: Benchmarking Retrieval‑Augmented Multimodal Generation for Document Question Answering*. <a href="https://arxiv.org/abs/2505.16470" class="external text" rel="nofollow">arXiv:2505.16470</a>.
- Wasserman, N. et al. (2025). *REAL‑MM‑RAG: A Real‑World Multi‑Modal Retrieval Benchmark*. <a href="https://arxiv.org/abs/2502.12342" class="external text" rel="nofollow">arXiv:2502.12342</a>.
- Faysse, M. et al. (2024). *ColPali: Efficient Document Retrieval with Vision Language Models*. <a href="https://arxiv.org/abs/2407.01449" class="external text" rel="nofollow">arXiv:2407.01449</a>.
- Radford, A. et al. (2021). *Learning Transferable Visual Models From Natural Language Supervision (CLIP)*. ICML 2021. <a href="https://arxiv.org/abs/2103.00020" class="external text" rel="nofollow">arXiv:2103.00020</a>.
- Tschannen, M. et al. (2025). *SigLIP 2: Multilingual Vision‑Language Encoders with Improved Semantic Understanding, Localization, and Dense Features*. <a href="https://arxiv.org/abs/2502.14786" class="external text" rel="nofollow">arXiv:2502.14786</a>.
- Xu, Y. et al. (2020). *LayoutLM: Pre‑training of Text and Layout for Document Image Understanding*. KDD 2020. <a href="https://dl.acm.org/doi/10.1145/3394486.3403172" class="external text" rel="nofollow">DOI</a> • <a href="https://arxiv.org/abs/1912.13318" class="external text" rel="nofollow">arXiv:1912.13318</a>.
- Huang, Y. et al. (2022). *LayoutLMv3: Pre‑training for Document AI with Unified Text and Image Masking*. <a href="https://arxiv.org/abs/2204.08387" class="external text" rel="nofollow">arXiv:2204.08387</a>.
- Zhong, X., Tang, J., Jimeno‑Yepes, A.J. (2019). *PubLayNet: Largest Dataset Ever for Document Layout Analysis*. <a href="https://arxiv.org/abs/1908.07836" class="external text" rel="nofollow">arXiv:1908.07836</a>.
- Pfitzmann, B. et al. (2022). *DocLayNet: A Large Human‑Annotated Dataset for Document‑Layout Analysis*. <a href="https://arxiv.org/abs/2206.01062" class="external text" rel="nofollow">arXiv:2206.01062</a>.
- Kim, G. et al. (2021). *OCR‑free Document Understanding Transformer (Donut)*. <a href="https://arxiv.org/abs/2111.15664" class="external text" rel="nofollow">arXiv:2111.15664</a>.
- Shen, Z. et al. (2021). *LayoutParser: A Unified Toolkit for Deep Learning Based Document Image Analysis*. <a href="https://arxiv.org/abs/2103.15348" class="external text" rel="nofollow">arXiv:2103.15348</a>.
- Singh, A. et al. (2019). *Towards VQA Models That Can Read (TextVQA)*. CVPR 2019. <a href="https://arxiv.org/abs/1904.08920" class="external text" rel="nofollow">arXiv:1904.08920</a>.
- Mathew, M. et al. (2021/2022). *DocVQA / InfographicVQA: Datasets for VQA on Document Images and Infographics*. WACV 2021 / WACV 2022. <a href="https://openaccess.thecvf.com/content/WACV2021/html/Mathew_DocVQA_A_Dataset_for_VQA_on_Document_Images_WACV_2021_paper.html" class="external text" rel="nofollow">CVF</a> • <a href="https://arxiv.org/abs/2104.12756" class="external text" rel="nofollow">arXiv:2104.12756</a>.
- Masry, A. et al. (2022). *ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning*. Findings of ACL 2022. <a href="https://aclanthology.org/2022.findings-acl.177/" class="external text" rel="nofollow">ACL</a> • <a href="https://arxiv.org/abs/2203.10244" class="external text" rel="nofollow">arXiv:2203.10244</a>.
- Liu, F. et al. (2022). *DePlot: One‑shot Visual Language Reasoning by Plot‑to‑Table Translation*. <a href="https://arxiv.org/abs/2212.10505" class="external text" rel="nofollow">arXiv:2212.10505</a>.
- Wang, P. et al. (2024). *Qwen2‑VL: Enhancing Vision‑Language Model’s Capabilities in OCR and Chart QA*. <a href="https://arxiv.org/abs/2409.12191" class="external text" rel="nofollow">arXiv:2409.12191</a>.
- Es, S. et al. (2023). *RAGAS: Automated Evaluation of Retrieval‑Augmented Generation*. <a href="https://arxiv.org/abs/2309.15217" class="external text" rel="nofollow">arXiv:2309.15217</a>.

## Einzelnachweise

1.  <span id="cite_note-lewis2020-1">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_ref-lewis2020_1-0) Lewis, P., Perez, E., et al. (2020). *Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks*. NeurIPS. <a href="https://arxiv.org/abs/2005.11401" class="external text" rel="nofollow">arXiv:2005.11401</a>.</span>
2.  <span id="cite_note-visrag2024-2">↑ <sup>[2.0](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_ref-visrag2024_2-0)</sup> <sup>[2.1](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_ref-visrag2024_2-1)</sup> Yu, S. et al. (2024). *VisRAG: Vision-based Retrieval-Augmented Generation on Multi-modality Documents*. <a href="https://arxiv.org/abs/2407.06437" class="external text" rel="nofollow">arXiv:2407.06437</a>.</span>
3.  <span id="cite_note-docvqa2021-3">↑ <sup>[3.0](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_ref-docvqa2021_3-0)</sup> <sup>[3.1](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_ref-docvqa2021_3-1)</sup> Mathew, M. et al. (2021). *DocVQA: A Dataset for VQA on Document Images*. WACV. <a href="https://arxiv.org/abs/2007.00398" class="external text" rel="nofollow">arXiv:2007.00398</a>.</span>
4.  <span id="cite_note-textvqa2019-4">↑ <sup>[4.0](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_ref-textvqa2019_4-0)</sup> <sup>[4.1](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_ref-textvqa2019_4-1)</sup> Singh, A. et al. (2019). *TextVQA: Towards VQA Models That Can Read*. CVPR. <a href="https://arxiv.org/abs/1904.08920" class="external text" rel="nofollow">arXiv:1904.08920</a>.</span>
5.  <span id="cite_note-layoutlm2020-5">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_ref-layoutlm2020_5-0) Xu, Y. et al. (2020). *LayoutLM: Pre-training of Text and Layout for Document Image Understanding*. KDD. <a href="https://dl.acm.org/doi/10.1145/3394486.3403172" class="external text" rel="nofollow">DOI</a>.</span>
6.  <span id="cite_note-clip2021-6">↑ <sup>[6.0](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_ref-clip2021_6-0)</sup> <sup>[6.1](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_ref-clip2021_6-1)</sup> Radford, A. et al. (2021). *Learning Transferable Visual Models From Natural Language Supervision (CLIP)*. ICML. <a href="https://arxiv.org/abs/2103.00020" class="external text" rel="nofollow">arXiv:2103.00020</a>.</span>
7.  <span id="cite_note-publaynet2019-7">↑ <sup>[7.0](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_ref-publaynet2019_7-0)</sup> <sup>[7.1](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_ref-publaynet2019_7-1)</sup> Zhong, X., Tang, J., Yepes, A. J. (2019). *PubLayNet: Largest Dataset for Document Layout Analysis*. ICDAR. <a href="https://arxiv.org/abs/1908.07836" class="external text" rel="nofollow">arXiv:1908.07836</a>.</span>
8.  <span id="cite_note-doclaynet2022-8">↑ <sup>[8.0](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_ref-doclaynet2022_8-0)</sup> <sup>[8.1](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_ref-doclaynet2022_8-1)</sup> Pfitzmann, B. et al. (2022). *DocLayNet: A Large Human‑Annotated Dataset for Document‑Layout Analysis*. KDD. <a href="https://dl.acm.org/doi/10.1145/3534678.3539043" class="external text" rel="nofollow">DOI</a> / <a href="https://arxiv.org/abs/2206.01062" class="external text" rel="nofollow">arXiv:2206.01062</a>.</span>
9.  <span id="cite_note-layoutparser2021-9">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_ref-layoutparser2021_9-0) Shen, Z. et al. (2021). *LayoutParser: A Unified Toolkit for DL‑based Document Image Analysis*. <a href="https://arxiv.org/abs/2103.15348" class="external text" rel="nofollow">arXiv:2103.15348</a>.</span>
10. <span id="cite_note-donut2021-10">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_ref-donut2021_10-0) Kim, G. et al. (2021). *Donut: OCR‑free Document Understanding Transformer*. <a href="https://arxiv.org/abs/2111.15664" class="external text" rel="nofollow">arXiv:2111.15664</a>.</span>
11. <span id="cite_note-siglip2023-11">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_ref-siglip2023_11-0) Zhai, X. et al. (2023). *Sigmoid Loss for Language‑Image Pre‑Training (SigLIP)*. ICCV. <a href="https://arxiv.org/abs/2303.15343" class="external text" rel="nofollow">arXiv:2303.15343</a>.</span>
12. <span id="cite_note-milvus-mv-12">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_ref-milvus-mv_12-0) Milvus Docs. *Multi‑Vector Hybrid Search*. <a href="https://milvus.io/docs/multi-vector-search.md" class="external text" rel="nofollow">milvus.io/docs/multi-vector-search.md</a>.</span>
13. <span id="cite_note-cohere-rerank-13">↑ <sup>[13.0](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_ref-cohere-rerank_13-0)</sup> <sup>[13.1](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_ref-cohere-rerank_13-1)</sup> Cohere Docs. *Rerank API*. <a href="https://docs.cohere.com/reference/rerank" class="external text" rel="nofollow">docs.cohere.com/reference/rerank</a>.</span>
14. <span id="cite_note-gpt4v-14">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_ref-gpt4v_14-0) OpenAI. *GPT‑4V(ision) System Card*. (2023). <a href="https://cdn.openai.com/papers/GPTV_System_Card.pdf" class="external text" rel="nofollow">PDF</a>.</span>
15. <span id="cite_note-gpt4o-15">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_ref-gpt4o_15-0) OpenAI. *Hello GPT‑4o*. (2024). <a href="https://openai.com/index/hello-gpt-4o/" class="external text" rel="nofollow">openai.com/index/hello-gpt-4o/</a>.</span>
16. <span id="cite_note-ragas-16">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_ref-ragas_16-0) Es, S. et al. (2023). *RAGAS: Automated Evaluation of Retrieval Augmented Generation*. <a href="https://arxiv.org/abs/2309.15217" class="external text" rel="nofollow">arXiv:2309.15217</a>.</span>
17. <span id="cite_note-infographicvqa2021-17">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_ref-infographicvqa2021_17-0) Mathew, M. et al. (2021). *InfographicVQA: Understanding Infographics via Question Answering*. ICDAR. <a href="https://arxiv.org/abs/2104.12756" class="external text" rel="nofollow">arXiv:2104.12756</a>.</span>
18. <span id="cite_note-chartqa2022-18">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_ref-chartqa2022_18-0) Masry, A. et al. (2022). *ChartQA: A Benchmark for Question Answering about Charts*. ACL (Findings). <a href="https://arxiv.org/abs/2103.16435" class="external text" rel="nofollow">arXiv:2103.16435</a>.</span>
19. <span id="cite_note-mmdocrag2025-19">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_ref-mmdocrag2025_19-0) Dong, K. et al. (2025). *Benchmarking Retrieval‑Augmented Multimodal Generation for Document QA (MMDocRAG)*. <a href="https://arxiv.org/abs/2505.16470" class="external text" rel="nofollow">arXiv:2505.16470</a>.</span>
20. <span id="cite_note-weaviate-hybrid-20">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_ref-weaviate-hybrid_20-0) Weaviate Docs. *Hybrid search*. <a href="https://docs.weaviate.io/weaviate/concepts/search/hybrid-search" class="external text" rel="nofollow">docs.weaviate.io/.../hybrid-search</a>.</span>
21. <span id="cite_note-colpali2024-21">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodales_RAG)#cite_ref-colpali2024_21-0) Faysse, M. et al. (2024). *ColPali: Efficient Document Retrieval with Vision‑Language Models*. <a href="https://arxiv.org/abs/2407.01449" class="external text" rel="nofollow">arXiv:2407.01449</a>.</span>
