---
title: "MM-RAG (Multimodal RAG) (NL)"
source: "https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)"
wiki: "systems-analysis.info/int"
article: "MM-RAG_(Multimodal_RAG)_(NL)"
language: "nl"
categories:
  - "Category:Dutch"
  - "Category:Large language models"
  - "Category:Prompt engineering"
revision_id: 4054
wiki_created_at: 2026-09-06T23:29:50Z
wiki_modified_at: 2026-09-06T23:29:50Z
downloaded_at: 2026-09-07T23:00:36Z
---

# MM-RAG (Multimodal RAG) (NL)

**MM-RAG** (Engels: *Multimodal Retrieval-Augmented Generation*) — een uitbreiding van het klassieke RAG-paradigma, waarbij LLM's voor het beantwoorden van vragen niet alleen tekst gebruiken, maar ook visuele gegevens (afbeeldingen, schema's, tabellen, grafieken). Multimodale retrieval maakt het mogelijk bewijzen te vinden en te koppelen in verschillende representaties, waardoor het risico op hallucinaties afneemt doordat men steunt op externe bronnen met nauwkeurige verwijzingen naar paginafragmenten en gebieden (*bounding boxes*)<sup>[\[1\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_note-lewis2020-1)[\[2\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_note-visrag2024-2)</sup>.

MM-RAG is vooral nuttig voor documenten waarbij een aanzienlijk deel van de betekenis in niet-tekstuele vorm is weergegeven (pagina-indeling, diagrammen, tabelstructuren). In dergelijke gevallen verliest klassieke tekstuele RAG vaak belangrijke contextelementen<sup>[\[3\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_note-docvqa2021-3)[\[4\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_note-textvqa2019-4)</sup>.

## Context en het op te lossen probleem

Klassieke RAG werkt met tekstpassages en ziet geen visuele structuren (positie van elementen, bijschriften bij figuren, assen van grafieken). MM-RAG sluit deze hiaten: het extraheert gestructureerde elementen (tekst, tabellen, afbeeldingen met coördinaten), indexeert ze in een vectorruimte en combineert bewijzen uit verschillende modaliteiten<sup>[\[5\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_note-layoutlm2020-5)[\[6\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_note-clip2021-6)</sup>.

## Architectuur van MM-RAG

De MM-RAG-pipeline voegt aan klassieke RAG de stappen voor verwerking van visuele gegevens en het uitlijnen van modaliteiten toe: **ingestie → indexering → multimodale retrieval → samenvoeging en herrangschikking → generatie met tracering**.

1.  **Verzameling en preprocessing (Ingestion).** Als invoer dienen PDF's/scans/afbeeldingen. Er worden OCR en **lay-outanalyse** van de pagina uitgevoerd om zones te onderscheiden: alinea's, koppen, tabellen, afbeeldingen en hun coördinaten. Typische hulpmiddelen zijn modellen uit de LayoutLM-familie en de instrumentele bibliotheek LayoutParser; validatie en training steunen vaak op de datasets PubLayNet en DocLayNet<sup>[\[7\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_note-publaynet2019-7)[\[8\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_note-doclaynet2022-8)[\[9\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_note-layoutparser2021-9)</sup>.
2.  **Segmentatie in regio's.** Visuele objecten worden geëxtraheerd (diagrammen, tabellen, illustraties, bijschriften). Voor verhoogde robuustheid worden OCR‑vrije modellen (bijv. Donut) of gecombineerde OCR+VLM-pipelines toegepast<sup>[\[10\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_note-donut2021-10)</sup>.
3.  **Indexering (Vector Index).** Tekstfragmenten en visuele elementen (afbeeldingen of beschrijvingen daarvan) worden omgezet in vectorrepresentaties en opgenomen in een vector-database. Voor een gecombineerde *text↔image*-ruimte worden CLIP of SigLIP gebruikt; voor productieomgevingen zijn multimodale/multivector-indexen handig (één object — meerdere vectoren)<sup>[\[6\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_note-clip2021-6)[\[11\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_note-siglip2023-11)[\[12\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_note-milvus-mv-12)</sup>.
4.  **Multimodale retrieval en herrangschikking.** Er wordt een combinatie van tekstuele en visuele zoekopdrachten uitgevoerd; kandidaten (alinea's, tabellen, afbeeldingen/regio's) worden samengevoegd en herrangschikt door een zwaarder model (cross-encoder/LLM-reranker) om de nauwkeurigheid te verhogen<sup>[\[13\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_note-cohere-rerank-13)</sup>.
5.  **Contextinpakking en generatie.** De geselecteerde fragmenten worden aangeleverd aan de LLM/VLM. Als het model multimodaal is (bijv. GPT‑4V/4o), kunnen afbeeldingen direct worden aangeboden; bij een tekstuele LLM worden afbeeldingen vooraf omgezet in gedetailleerde beschrijvingen<sup>[\[14\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_note-gpt4v-14)[\[15\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_note-gpt4o-15)</sup>.
6.  **Tracering en citering.** Het antwoord wordt vergezeld door klikbare citaten die niet alleen naar het document/de pagina verwijzen, maar ook naar de regio (coördinaten). Dit verhoogt het niveau van *grounding* en het vertrouwen van gebruikers<sup>[\[2\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_note-visrag2024-2)</sup>.

## Kwaliteitsbeoordeling en metrieken

De effectiviteit van MM‑RAG wordt beoordeeld op het niveau van extractie, retrieval en generatie.

- **Kwaliteit van de extractie van visuele gegevens.** OCR-nauwkeurigheid (WER/CER), kwaliteit van de lay-outanalyse (mAP/Precision/Recall) op de datasets DocLayNet/PubLayNet<sup>[\[8\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_note-doclaynet2022-8)[\[7\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_note-publaynet2019-7)</sup>.
- **Kwaliteit van de retrieval.** Standaard IR-metrieken: Recall@K, Precision@K, MRR; voor multimodaliteit — afzonderlijk per modaliteit en in combinatie.
- **Kwaliteit van het antwoord (end‑to‑end).** Automatische metrieken voor *faithfulness*/*groundedness* en menselijke beoordeling. In de praktijk worden de frameworks RAGAS/TruLens/DeepEval gebruikt<sup>[\[16\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_note-ragas-16)</sup>.
- **Benchmarks.**
  - **DocVQA**: vragen bij afbeeldingen van documenten<sup>[\[3\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_note-docvqa2021-3)</sup>.
  - **TextVQA**: vragen waarvoor tekst op afbeeldingen gelezen moet worden<sup>[\[4\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_note-textvqa2019-4)</sup>.
  - **InfographicVQA**: vragen over infographics<sup>[\[17\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_note-infographicvqa2021-17)</sup>.
  - **ChartQA**: vragen over diagrammen met vereiste logische redenering<sup>[\[18\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_note-chartqa2022-18)</sup>.
  - **MMDocRAG**: benchmark voor multimodale RAG voor DocQA (meerpagina-documenten, cross-modale bewijsketens)<sup>[\[19\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_note-mmdocrag2025-19)</sup>.

## Vergelijkingstabel van componenten

| Component                  | Implementatievarianten                                           | Voordelen                                                                                                  | Nadelen / risico's                                                                       | Wanneer te kiezen                                                                          |
|----------------------------|------------------------------------------------------------------|------------------------------------------------------------------------------------------------------------|------------------------------------------------------------------------------------------|--------------------------------------------------------------------------------------------|
| OCR                        | Tesseract / PaddleOCR / cloud-API's                              | Lokaal — privacy en controle; cloud — hoge nauwkeurigheid out-of-the-box.                                  | Fouten bij complexe lay-out; API — kosten en compliance-vereisten.                       | Privégegevens — lokale OCR; maximale nauwkeurigheid — cloud (indien toegestaan).           |
| Lay-outanalyse             | Regels / ML‑model (LayoutLM, LayoutParser)                       | Regels zijn eenvoudig voor eenvormige sjablonen; ML — robuust bij variatie.                                | Regels werken niet bij nieuwe lay-outs; ML vereist middelen/data.                        | Eenvormige formulieren — regels; heterogeen corpus — ML.                                   |
| Vectorisatie (afb.)        | CLIP / SigLIP / OCR‑vrije beschrijvingen (Donut/Pix2Struct)      | Gemeenschappelijke latente ruimte *text↔image* (CLIP/SigLIP); OCR‑vrij elimineert afhankelijkheid van OCR. | CLIP leest geen tekst binnen afbeeldingen; beschrijvingen kunnen de betekenis vervormen. | CLIP/SigLIP — basale multimodale zoekopdracht; OCR‑vrij voor scans van complexe kwaliteit. |
| Samenvoegen van resultaten | Sortering op score / quota per modaliteit / LLM‑reranker         | Reranker verhoogt de nauwkeurigheid van contextselectie aanzienlijk.                                       | Toename van vertraging en kosten.                                                        | High‑precision scenario's; eenvoudige methoden — voor PoC.                                 |
| Opslag/index               | Enkele vector / multivector (text+image) / hybride (BM25+vector) | Multivector dekt verschillende representaties van één object; hybride redt sleutelwoorden/codes.           | Complexere schema's en updates.                                                          | Productiesystemen met gemengde gegevens en strikte SLA's.                                  |

Vergelijking van sleutelcomponenten en benaderingen in MM‑RAG

## Praktische opmerkingen

- **Hybride zoekopdracht** (BM25 + vector) — de facto standaard voor het verbeteren van volledigheid en nauwkeurigheid bij specifieke termen/codes<sup>[\[20\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_note-weaviate-hybrid-20)</sup>.
- **Herrangschikking** met een cross-encoder/LLM bespaart tokens door 'ruis'-kandidaten vóór de generatie te verwijderen<sup>[\[13\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_note-cohere-rerank-13)</sup>.
- **Moderne VLM-retrievers** (bijv. ColPali) tonen voordelen bij visueel rijke documenten dankzij directe indexering van pagina-afbeeldingen<sup>[\[21\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_note-colpali2024-21)</sup>.

## Literatuur

- Lewis, P. et al. (2020). *Retrieval‑Augmented Generation for Knowledge‑Intensive NLP Tasks*. NeurIPS. arXiv:2005.11401.
- Gao, L. et al. (2023). *Precise Zero‑Shot Dense Retrieval without Relevance Labels (HyDE)*. ACL 2023. arXiv:2212.10496.
- Mei, L., Mo, S., Yang, Z., Chen, C. (2025). *A Survey of Multimodal Retrieval‑Augmented Generation*. arXiv:2504.08748.
- Abootorabi, M.M. et al. (2025). *Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval‑Augmented Generation*. Findings of ACL 2025. ACL Anthology.
- Yu, S. et al. (2024). *VisRAG: Vision‑based Retrieval‑augmented Generation on Multi‑modality Documents*. arXiv:2410.10594.
- Cho, J. et al. (2024). *M3DocRAG: Multi‑modal Retrieval is What You Need for Multi‑document QA*. arXiv:2411.04952.
- Tanaka, R. et al. (2025). *VDocRAG: Retrieval‑Augmented Generation over Visually‑Rich Documents*. CVPR 2025. arXiv:2504.09795 • CVF Open Access.
- Dong, K. et al. (2025). *MMDocRAG: Benchmarking Retrieval‑Augmented Multimodal Generation for Document Question Answering*. arXiv:2505.16470.
- Wasserman, N. et al. (2025). *REAL‑MM‑RAG: A Real‑World Multi‑Modal Retrieval Benchmark*. arXiv:2502.12342.
- Faysse, M. et al. (2024). *ColPali: Efficient Document Retrieval with Vision Language Models*. arXiv:2407.01449.
- Radford, A. et al. (2021). *Learning Transferable Visual Models From Natural Language Supervision (CLIP)*. ICML 2021. arXiv:2103.00020.
- Tschannen, M. et al. (2025). *SigLIP 2: Multilingual Vision‑Language Encoders with Improved Semantic Understanding, Localization, and Dense Features*. arXiv:2502.14786.
- Xu, Y. et al. (2020). *LayoutLM: Pre‑training of Text and Layout for Document Image Understanding*. KDD 2020. DOI • arXiv:1912.13318.
- Huang, Y. et al. (2022). *LayoutLMv3: Pre‑training for Document AI with Unified Text and Image Masking*. arXiv:2204.08387.
- Zhong, X., Tang, J., Jimeno‑Yepes, A.J. (2019). *PubLayNet: Largest Dataset Ever for Document Layout Analysis*. arXiv:1908.07836.
- Pfitzmann, B. et al. (2022). *DocLayNet: A Large Human‑Annotated Dataset for Document‑Layout Analysis*. arXiv:2206.01062.
- Kim, G. et al. (2021). *OCR‑free Document Understanding Transformer (Donut)*. arXiv:2111.15664.
- Shen, Z. et al. (2021). *LayoutParser: A Unified Toolkit for Deep Learning Based Document Image Analysis*. arXiv:2103.15348.
- Singh, A. et al. (2019). *Towards VQA Models That Can Read (TextVQA)*. CVPR 2019. arXiv:1904.08920.
- Mathew, M. et al. (2021/2022). *DocVQA / InfographicVQA: Datasets for VQA on Document Images and Infographics*. WACV 2021 / WACV 2022. CVF • arXiv:2104.12756.
- Masry, A. et al. (2022). *ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning*. Findings of ACL 2022. ACL • arXiv:2203.10244.
- Liu, F. et al. (2022). *DePlot: One‑shot Visual Language Reasoning by Plot‑to‑Table Translation*. arXiv:2212.10505.
- Wang, P. et al. (2024). *Qwen2‑VL: Enhancing Vision‑Language Model's Capabilities in OCR and Chart QA*. arXiv:2409.12191.
- Es, S. et al. (2023). *RAGAS: Automated Evaluation of Retrieval‑Augmented Generation*. arXiv:2309.15217.

## Zie ook

- Retrieval-Augmented Generation
- Vectordatabase
- Embedding
- GraphRAG

## Noten

1.  <span id="cite_note-lewis2020-1">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_ref-lewis2020_1-0) Lewis, P., Perez, E., et al. (2020). *Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks*. NeurIPS. <a href="https://arxiv.org/abs/2005.11401" class="external text" rel="nofollow">arXiv:2005.11401</a>.</span>
2.  <span id="cite_note-visrag2024-2">↑ <sup>[2.0](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_ref-visrag2024_2-0)</sup> <sup>[2.1](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_ref-visrag2024_2-1)</sup> Yu, S. et al. (2024). *VisRAG: Vision-based Retrieval-Augmented Generation on Multi-modality Documents*. <a href="https://arxiv.org/abs/2407.06437" class="external text" rel="nofollow">arXiv:2407.06437</a>.</span>
3.  <span id="cite_note-docvqa2021-3">↑ <sup>[3.0](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_ref-docvqa2021_3-0)</sup> <sup>[3.1](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_ref-docvqa2021_3-1)</sup> Mathew, M. et al. (2021). *DocVQA: A Dataset for VQA on Document Images*. WACV. <a href="https://arxiv.org/abs/2007.00398" class="external text" rel="nofollow">arXiv:2007.00398</a>.</span>
4.  <span id="cite_note-textvqa2019-4">↑ <sup>[4.0](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_ref-textvqa2019_4-0)</sup> <sup>[4.1](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_ref-textvqa2019_4-1)</sup> Singh, A. et al. (2019). *TextVQA: Towards VQA Models That Can Read*. CVPR. <a href="https://arxiv.org/abs/1904.08920" class="external text" rel="nofollow">arXiv:1904.08920</a>.</span>
5.  <span id="cite_note-layoutlm2020-5">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_ref-layoutlm2020_5-0) Xu, Y. et al. (2020). *LayoutLM: Pre-training of Text and Layout for Document Image Understanding*. KDD. <a href="https://dl.acm.org/doi/10.1145/3394486.3403172" class="external text" rel="nofollow">DOI</a>.</span>
6.  <span id="cite_note-clip2021-6">↑ <sup>[6.0](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_ref-clip2021_6-0)</sup> <sup>[6.1](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_ref-clip2021_6-1)</sup> Radford, A. et al. (2021). *Learning Transferable Visual Models From Natural Language Supervision (CLIP)*. ICML. <a href="https://arxiv.org/abs/2103.00020" class="external text" rel="nofollow">arXiv:2103.00020</a>.</span>
7.  <span id="cite_note-publaynet2019-7">↑ <sup>[7.0](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_ref-publaynet2019_7-0)</sup> <sup>[7.1](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_ref-publaynet2019_7-1)</sup> Zhong, X., Tang, J., Yepes, A. J. (2019). *PubLayNet: Largest Dataset for Document Layout Analysis*. ICDAR. <a href="https://arxiv.org/abs/1908.07836" class="external text" rel="nofollow">arXiv:1908.07836</a>.</span>
8.  <span id="cite_note-doclaynet2022-8">↑ <sup>[8.0](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_ref-doclaynet2022_8-0)</sup> <sup>[8.1](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_ref-doclaynet2022_8-1)</sup> Pfitzmann, B. et al. (2022). *DocLayNet: A Large Human‑Annotated Dataset for Document‑Layout Analysis*. KDD. <a href="https://dl.acm.org/doi/10.1145/3534678.3539043" class="external text" rel="nofollow">DOI</a> / <a href="https://arxiv.org/abs/2206.01062" class="external text" rel="nofollow">arXiv:2206.01062</a>.</span>
9.  <span id="cite_note-layoutparser2021-9">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_ref-layoutparser2021_9-0) Shen, Z. et al. (2021). *LayoutParser: A Unified Toolkit for DL‑based Document Image Analysis*. <a href="https://arxiv.org/abs/2103.15348" class="external text" rel="nofollow">arXiv:2103.15348</a>.</span>
10. <span id="cite_note-donut2021-10">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_ref-donut2021_10-0) Kim, G. et al. (2021). *Donut: OCR‑free Document Understanding Transformer*. <a href="https://arxiv.org/abs/2111.15664" class="external text" rel="nofollow">arXiv:2111.15664</a>.</span>
11. <span id="cite_note-siglip2023-11">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_ref-siglip2023_11-0) Zhai, X. et al. (2023). *Sigmoid Loss for Language‑Image Pre‑Training (SigLIP)*. ICCV. <a href="https://arxiv.org/abs/2303.15343" class="external text" rel="nofollow">arXiv:2303.15343</a>.</span>
12. <span id="cite_note-milvus-mv-12">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_ref-milvus-mv_12-0) Milvus Docs. *Multi‑Vector Hybrid Search*. <a href="https://milvus.io/docs/multi-vector-search.md" class="external text" rel="nofollow">milvus.io/docs/multi-vector-search.md</a>.</span>
13. <span id="cite_note-cohere-rerank-13">↑ <sup>[13.0](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_ref-cohere-rerank_13-0)</sup> <sup>[13.1](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_ref-cohere-rerank_13-1)</sup> Cohere Docs. *Rerank API*. <a href="https://docs.cohere.com/reference/rerank" class="external text" rel="nofollow">docs.cohere.com/reference/rerank</a>.</span>
14. <span id="cite_note-gpt4v-14">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_ref-gpt4v_14-0) OpenAI. *GPT‑4V(ision) System Card*. (2023). <a href="https://cdn.openai.com/papers/GPTV_System_Card.pdf" class="external text" rel="nofollow">PDF</a>.</span>
15. <span id="cite_note-gpt4o-15">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_ref-gpt4o_15-0) OpenAI. *Hello GPT‑4o*. (2024). <a href="https://openai.com/index/hello-gpt-4o/" class="external text" rel="nofollow">openai.com/index/hello-gpt-4o/</a>.</span>
16. <span id="cite_note-ragas-16">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_ref-ragas_16-0) Es, S. et al. (2023). *RAGAS: Automated Evaluation of Retrieval Augmented Generation*. <a href="https://arxiv.org/abs/2309.15217" class="external text" rel="nofollow">arXiv:2309.15217</a>.</span>
17. <span id="cite_note-infographicvqa2021-17">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_ref-infographicvqa2021_17-0) Mathew, M. et al. (2021). *InfographicVQA: Understanding Infographics via Question Answering*. ICDAR. <a href="https://arxiv.org/abs/2104.12756" class="external text" rel="nofollow">arXiv:2104.12756</a>.</span>
18. <span id="cite_note-chartqa2022-18">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_ref-chartqa2022_18-0) Masry, A. et al. (2022). *ChartQA: A Benchmark for Question Answering about Charts*. ACL (Findings). <a href="https://arxiv.org/abs/2103.16435" class="external text" rel="nofollow">arXiv:2103.16435</a>.</span>
19. <span id="cite_note-mmdocrag2025-19">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_ref-mmdocrag2025_19-0) Dong, K. et al. (2025). *Benchmarking Retrieval‑Augmented Multimodal Generation for Document QA (MMDocRAG)*. <a href="https://arxiv.org/abs/2505.16470" class="external text" rel="nofollow">arXiv:2505.16470</a>.</span>
20. <span id="cite_note-weaviate-hybrid-20">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_ref-weaviate-hybrid_20-0) Weaviate Docs. *Hybrid search*. <a href="https://docs.weaviate.io/weaviate/concepts/search/hybrid-search" class="external text" rel="nofollow">docs.weaviate.io/.../hybrid-search</a>.</span>
21. <span id="cite_note-colpali2024-21">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(NL)#cite_ref-colpali2024_21-0) Faysse, M. et al. (2024). *ColPali: Efficient Document Retrieval with Vision‑Language Models*. <a href="https://arxiv.org/abs/2407.01449" class="external text" rel="nofollow">arXiv:2407.01449</a>.</span>
