---
title: "MM-RAG (Multimodal RAG) (ID)"
source: "https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)"
wiki: "systems-analysis.info/int"
article: "MM-RAG_(Multimodal_RAG)_(ID)"
language: "id"
categories:
  - "Category:Indonesian"
  - "Category:Large language models"
  - "Category:Prompt engineering"
revision_id: 4051
wiki_created_at: 2026-09-06T23:29:47Z
wiki_modified_at: 2026-09-06T23:29:47Z
downloaded_at: 2026-09-07T23:00:35Z
---

# MM-RAG (Multimodal RAG) (ID)

**MM-RAG** (Ingg. *Multimodal Retrieval-Augmented Generation*) — adalah perluasan dari paradigma klasik RAG, di mana LLM menggunakan tidak hanya teks, tetapi juga data visual (gambar, skema, tabel, grafik) untuk menjawab pertanyaan. Retrieval multimodal memungkinkan pencarian dan penghubungan bukti dalam berbagai representasi, mengurangi risiko halusinasi dengan mengandalkan sumber eksternal yang terikat secara tepat pada fragmen halaman dan area (*bounding boxes*)<sup>[\[1\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_note-lewis2020-1)[\[2\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_note-visrag2024-2)</sup>.

MM-RAG sangat berguna untuk dokumen yang sebagian besar maknanya disajikan dalam bentuk non-teks (tata letak halaman, diagram, struktur tabel). Dalam kasus seperti itu, RAG teks klasik sering kehilangan elemen konteks yang penting<sup>[\[3\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_note-docvqa2021-3)[\[4\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_note-textvqa2019-4)</sup>.

## Konteks dan Masalah yang Diselesaikan

RAG klasik beroperasi dengan passages teks dan tidak dapat melihat struktur visual (posisi elemen, keterangan gambar, sumbu grafik). MM-RAG menutup kesenjangan ini: mengekstrak elemen terstruktur (teks, tabel, gambar dengan koordinat), mengindeksnya ke dalam ruang vektor, dan menggabungkan bukti dari berbagai modalitas<sup>[\[5\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_note-layoutlm2020-5)[\[6\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_note-clip2021-6)</sup>.

## Arsitektur MM-RAG

Pipeline MM-RAG menambahkan tahap pemrosesan data visual dan penyelarasan modalitas ke RAG klasik: **ingesti → pengindeksan → retrieval multimodal → penggabungan dan reranking → generasi dengan tracing**.

1.  **Pengumpulan dan Preprocessing (Ingestion).** Input berupa PDF/scan/gambar. Dilakukan OCR dan **analisis tata letak** halaman untuk mengidentifikasi zona: paragraf, judul, tabel, gambar beserta koordinatnya. Alat yang umum digunakan adalah model dari keluarga LayoutLM dan library LayoutParser; validasi dan pelatihan sering bergantung pada dataset PubLayNet dan DocLayNet<sup>[\[7\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_note-publaynet2019-7)[\[8\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_note-doclaynet2022-8)[\[9\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_note-layoutparser2021-9)</sup>.
2.  **Segmentasi Region.** Objek visual diekstrak (diagram, tabel, ilustrasi, keterangan). Untuk ketahanan yang lebih tinggi, digunakan model OCR‑free (misalnya, Donut) atau pipeline gabungan OCR+VLM<sup>[\[10\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_note-donut2021-10)</sup>.
3.  **Pengindeksan (Vector Index).** Chunks teks dan elemen visual (gambar atau deskripsinya) diubah menjadi representasi vektor dan dimasukkan ke dalam basis data vektor. Untuk ruang gabungan *text↔image* digunakan CLIP atau SigLIP; untuk lingkungan produksi, indeks multimodal/multivektor (satu objek — beberapa vektor) lebih praktis<sup>[\[6\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_note-clip2021-6)[\[11\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_note-siglip2023-11)[\[12\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_note-milvus-mv-12)</sup>.
4.  **Retrieval Multimodal dan Reranking.** Dilakukan kombinasi pencarian teks dan visual; kandidat (paragraf, tabel, gambar/region) digabungkan dan di-rerank oleh model yang lebih "berat" (cross‑encoder/LLM‑reranker) untuk meningkatkan akurasi<sup>[\[13\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_note-cohere-rerank-13)</sup>.
5.  **Pengemasan Konteks dan Generasi.** Fragmen yang dipilih dimasukkan ke LLM/VLM. Jika model bersifat multimodal (misalnya, GPT‑4V/4o), gambar dapat dimasukkan secara langsung; pada LLM berbasis teks, gambar terlebih dahulu dikonversi menjadi deskripsi terperinci<sup>[\[14\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_note-gpt4v-14)[\[15\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_note-gpt4o-15)</sup>.
6.  **Tracing dan Sitasi.** Jawaban disertai dengan sitasi yang dapat diklik, yang terhubung tidak hanya ke dokumen/halaman, tetapi juga ke region (koordinat). Hal ini meningkatkan tingkat *grounding* dan kepercayaan pengguna<sup>[\[2\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_note-visrag2024-2)</sup>.

## Evaluasi Kualitas dan Metrik

Efektivitas MM‑RAG dievaluasi pada tingkat ekstraksi, retrieval, dan generasi.

- **Kualitas ekstraksi data visual.** Akurasi OCR (WER/CER), kualitas analisis tata letak (mAP/Precision/Recall) pada dataset DocLayNet/PubLayNet<sup>[\[8\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_note-doclaynet2022-8)[\[7\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_note-publaynet2019-7)</sup>.
- **Kualitas retrieval.** Metrik standar pencarian informasi: Recall@K, Precision@K, MRR; untuk multimodalitas — secara terpisah per modalitas dan secara gabungan.
- **Kualitas jawaban (end‑to‑end).** Metrik otomatis *faithfulness*/*groundedness* dan evaluasi manusia. Dalam praktiknya digunakan framework RAGAS/TruLens/DeepEval<sup>[\[16\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_note-ragas-16)</sup>.
- **Benchmark.**
  - **DocVQA**: pertanyaan berdasarkan gambar dokumen<sup>[\[3\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_note-docvqa2021-3)</sup>.
  - **TextVQA**: pertanyaan yang memerlukan pembacaan teks pada gambar<sup>[\[4\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_note-textvqa2019-4)</sup>.
  - **InfographicVQA**: pertanyaan berdasarkan infografis<sup>[\[17\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_note-infographicvqa2021-17)</sup>.
  - **ChartQA**: pertanyaan berdasarkan diagram dengan persyaratan penalaran logis<sup>[\[18\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_note-chartqa2022-18)</sup>.
  - **MMDocRAG**: benchmark multimodal RAG untuk DocQA (dokumen multi-halaman, rantai bukti lintas modalitas)<sup>[\[19\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_note-mmdocrag2025-19)</sup>.

## Tabel Perbandingan Komponen

| Komponen             | Varian Implementasi                                               | Kelebihan                                                                                        | Kekurangan / Risiko                                                         | Kapan Memilih                                                                       |
|----------------------|-------------------------------------------------------------------|--------------------------------------------------------------------------------------------------|-----------------------------------------------------------------------------|-------------------------------------------------------------------------------------|
| OCR                  | Tesseract / PaddleOCR / API cloud                                 | Lokal — privasi dan kontrol; cloud — akurasi tinggi siap pakai.                                  | Kesalahan pada tata letak kompleks; API — biaya dan persyaratan kepatuhan.  | Data privat — OCR lokal; akurasi maksimal — cloud (jika diizinkan).                 |
| Analisis tata letak  | Aturan / model ML (LayoutLM, LayoutParser)                        | Aturan sederhana untuk template seragam; ML — tahan terhadap keberagaman.                        | Aturan gagal pada tata letak baru; ML membutuhkan sumber daya/data.         | Formulir seragam — aturan; korpus heterogen — ML.                                   |
| Vektorisasi (gambar) | CLIP / SigLIP / deskripsi OCR‑free (Donut/Pix2Struct)             | Ruang laten gabungan *text↔image* (CLIP/SigLIP); OCR‑free menghilangkan ketergantungan pada OCR. | CLIP tidak membaca teks di dalam gambar; deskripsi dapat mendistorsi makna. | CLIP/SigLIP — pencarian multimodal dasar; OCR‑free untuk scan berkualitas kompleks. |
| Penggabungan hasil   | Pengurutan berdasarkan score / kuota per modalitas / LLM‑reranker | Reranker secara signifikan meningkatkan akurasi pemilihan konteks.                               | Peningkatan latensi dan biaya.                                              | Skenario high‑precision; metode sederhana — untuk PoC.                              |
| Penyimpanan/indeks   | Satu vektor / multivektor (text+image) / hibrida (BM25+vektor)    | Multivektor mencakup representasi berbeda dari satu objek; hibrida membantu kata kunci/kode.     | Kompleksitas skema dan pembaruan.                                           | Sistem produksi dengan data campuran dan SLA ketat.                                 |

Perbandingan komponen dan pendekatan utama dalam MM‑RAG

## Catatan Praktis

- **Pencarian hibrida** (BM25 + vektor) — standar de‑facto untuk meningkatkan kelengkapan dan akurasi pada istilah/kode spesifik<sup>[\[20\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_note-weaviate-hybrid-20)</sup>.
- **Reranking** dengan cross‑encoder/LLM menghemat token dengan membuang kandidat "sampah" sebelum generasi<sup>[\[13\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_note-cohere-rerank-13)</sup>.
- **Retriever VLM modern** (misalnya, ColPali) menunjukkan keunggulan pada dokumen kaya visual berkat pengindeksan langsung halaman sebagai gambar<sup>[\[21\]](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_note-colpali2024-21)</sup>.

## Referensi

- Lewis, P. et al. (2020). *Retrieval‑Augmented Generation for Knowledge‑Intensive NLP Tasks*. NeurIPS. arXiv:2005.11401.
- Gao, L. et al. (2023). *Precise Zero‑Shot Dense Retrieval without Relevance Labels (HyDE)*. ACL 2023. arXiv:2212.10496.
- Mei, L., Mo, S., Yang, Z., Chen, C. (2025). *A Survey of Multimodal Retrieval‑Augmented Generation*. arXiv:2504.08748.
- Abootorabi, M.M. et al. (2025). *Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval‑Augmented Generation*. Findings of ACL 2025. ACL Anthology.
- Yu, S. et al. (2024). *VisRAG: Vision‑based Retrieval‑augmented Generation on Multi‑modality Documents*. arXiv:2410.10594.
- Cho, J. et al. (2024). *M3DocRAG: Multi‑modal Retrieval is What You Need for Multi‑document QA*. arXiv:2411.04952.
- Tanaka, R. et al. (2025). *VDocRAG: Retrieval‑Augmented Generation over Visually‑Rich Documents*. CVPR 2025. arXiv:2504.09795 • CVF Open Access.
- Dong, K. et al. (2025). *MMDocRAG: Benchmarking Retrieval‑Augmented Multimodal Generation for Document Question Answering*. arXiv:2505.16470.
- Wasserman, N. et al. (2025). *REAL‑MM‑RAG: A Real‑World Multi‑Modal Retrieval Benchmark*. arXiv:2502.12342.
- Faysse, M. et al. (2024). *ColPali: Efficient Document Retrieval with Vision Language Models*. arXiv:2407.01449.
- Radford, A. et al. (2021). *Learning Transferable Visual Models From Natural Language Supervision (CLIP)*. ICML 2021. arXiv:2103.00020.
- Tschannen, M. et al. (2025). *SigLIP 2: Multilingual Vision‑Language Encoders with Improved Semantic Understanding, Localization, and Dense Features*. arXiv:2502.14786.
- Xu, Y. et al. (2020). *LayoutLM: Pre‑training of Text and Layout for Document Image Understanding*. KDD 2020. DOI • arXiv:1912.13318.
- Huang, Y. et al. (2022). *LayoutLMv3: Pre‑training for Document AI with Unified Text and Image Masking*. arXiv:2204.08387.
- Zhong, X., Tang, J., Jimeno‑Yepes, A.J. (2019). *PubLayNet: Largest Dataset Ever for Document Layout Analysis*. arXiv:1908.07836.
- Pfitzmann, B. et al. (2022). *DocLayNet: A Large Human‑Annotated Dataset for Document‑Layout Analysis*. arXiv:2206.01062.
- Kim, G. et al. (2021). *OCR‑free Document Understanding Transformer (Donut)*. arXiv:2111.15664.
- Shen, Z. et al. (2021). *LayoutParser: A Unified Toolkit for Deep Learning Based Document Image Analysis*. arXiv:2103.15348.
- Singh, A. et al. (2019). *Towards VQA Models That Can Read (TextVQA)*. CVPR 2019. arXiv:1904.08920.
- Mathew, M. et al. (2021/2022). *DocVQA / InfographicVQA: Datasets for VQA on Document Images and Infographics*. WACV 2021 / WACV 2022. CVF • arXiv:2104.12756.
- Masry, A. et al. (2022). *ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning*. Findings of ACL 2022. ACL • arXiv:2203.10244.
- Liu, F. et al. (2022). *DePlot: One‑shot Visual Language Reasoning by Plot‑to‑Table Translation*. arXiv:2212.10505.
- Wang, P. et al. (2024). *Qwen2‑VL: Enhancing Vision‑Language Model's Capabilities in OCR and Chart QA*. arXiv:2409.12191.
- Es, S. et al. (2023). *RAGAS: Automated Evaluation of Retrieval‑Augmented Generation*. arXiv:2309.15217.

## Lihat Juga

- Retrieval-Augmented Generation
- Basis data vektor
- Embedding
- GraphRAG

## Catatan

1.  <span id="cite_note-lewis2020-1">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_ref-lewis2020_1-0) Lewis, P., Perez, E., et al. (2020). *Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks*. NeurIPS. <a href="https://arxiv.org/abs/2005.11401" class="external text" rel="nofollow">arXiv:2005.11401</a>.</span>
2.  <span id="cite_note-visrag2024-2">↑ <sup>[2.0](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_ref-visrag2024_2-0)</sup> <sup>[2.1](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_ref-visrag2024_2-1)</sup> Yu, S. et al. (2024). *VisRAG: Vision-based Retrieval-Augmented Generation on Multi-modality Documents*. <a href="https://arxiv.org/abs/2407.06437" class="external text" rel="nofollow">arXiv:2407.06437</a>.</span>
3.  <span id="cite_note-docvqa2021-3">↑ <sup>[3.0](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_ref-docvqa2021_3-0)</sup> <sup>[3.1](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_ref-docvqa2021_3-1)</sup> Mathew, M. et al. (2021). *DocVQA: A Dataset for VQA on Document Images*. WACV. <a href="https://arxiv.org/abs/2007.00398" class="external text" rel="nofollow">arXiv:2007.00398</a>.</span>
4.  <span id="cite_note-textvqa2019-4">↑ <sup>[4.0](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_ref-textvqa2019_4-0)</sup> <sup>[4.1](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_ref-textvqa2019_4-1)</sup> Singh, A. et al. (2019). *TextVQA: Towards VQA Models That Can Read*. CVPR. <a href="https://arxiv.org/abs/1904.08920" class="external text" rel="nofollow">arXiv:1904.08920</a>.</span>
5.  <span id="cite_note-layoutlm2020-5">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_ref-layoutlm2020_5-0) Xu, Y. et al. (2020). *LayoutLM: Pre-training of Text and Layout for Document Image Understanding*. KDD. <a href="https://dl.acm.org/doi/10.1145/3394486.3403172" class="external text" rel="nofollow">DOI</a>.</span>
6.  <span id="cite_note-clip2021-6">↑ <sup>[6.0](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_ref-clip2021_6-0)</sup> <sup>[6.1](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_ref-clip2021_6-1)</sup> Radford, A. et al. (2021). *Learning Transferable Visual Models From Natural Language Supervision (CLIP)*. ICML. <a href="https://arxiv.org/abs/2103.00020" class="external text" rel="nofollow">arXiv:2103.00020</a>.</span>
7.  <span id="cite_note-publaynet2019-7">↑ <sup>[7.0](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_ref-publaynet2019_7-0)</sup> <sup>[7.1](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_ref-publaynet2019_7-1)</sup> Zhong, X., Tang, J., Yepes, A. J. (2019). *PubLayNet: Largest Dataset for Document Layout Analysis*. ICDAR. <a href="https://arxiv.org/abs/1908.07836" class="external text" rel="nofollow">arXiv:1908.07836</a>.</span>
8.  <span id="cite_note-doclaynet2022-8">↑ <sup>[8.0](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_ref-doclaynet2022_8-0)</sup> <sup>[8.1](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_ref-doclaynet2022_8-1)</sup> Pfitzmann, B. et al. (2022). *DocLayNet: A Large Human‑Annotated Dataset for Document‑Layout Analysis*. KDD. <a href="https://dl.acm.org/doi/10.1145/3534678.3539043" class="external text" rel="nofollow">DOI</a> / <a href="https://arxiv.org/abs/2206.01062" class="external text" rel="nofollow">arXiv:2206.01062</a>.</span>
9.  <span id="cite_note-layoutparser2021-9">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_ref-layoutparser2021_9-0) Shen, Z. et al. (2021). *LayoutParser: A Unified Toolkit for DL‑based Document Image Analysis*. <a href="https://arxiv.org/abs/2103.15348" class="external text" rel="nofollow">arXiv:2103.15348</a>.</span>
10. <span id="cite_note-donut2021-10">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_ref-donut2021_10-0) Kim, G. et al. (2021). *Donut: OCR‑free Document Understanding Transformer*. <a href="https://arxiv.org/abs/2111.15664" class="external text" rel="nofollow">arXiv:2111.15664</a>.</span>
11. <span id="cite_note-siglip2023-11">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_ref-siglip2023_11-0) Zhai, X. et al. (2023). *Sigmoid Loss for Language‑Image Pre‑Training (SigLIP)*. ICCV. <a href="https://arxiv.org/abs/2303.15343" class="external text" rel="nofollow">arXiv:2303.15343</a>.</span>
12. <span id="cite_note-milvus-mv-12">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_ref-milvus-mv_12-0) Milvus Docs. *Multi‑Vector Hybrid Search*. <a href="https://milvus.io/docs/multi-vector-search.md" class="external text" rel="nofollow">milvus.io/docs/multi-vector-search.md</a>.</span>
13. <span id="cite_note-cohere-rerank-13">↑ <sup>[13.0](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_ref-cohere-rerank_13-0)</sup> <sup>[13.1](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_ref-cohere-rerank_13-1)</sup> Cohere Docs. *Rerank API*. <a href="https://docs.cohere.com/reference/rerank" class="external text" rel="nofollow">docs.cohere.com/reference/rerank</a>.</span>
14. <span id="cite_note-gpt4v-14">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_ref-gpt4v_14-0) OpenAI. *GPT‑4V(ision) System Card*. (2023). <a href="https://cdn.openai.com/papers/GPTV_System_Card.pdf" class="external text" rel="nofollow">PDF</a>.</span>
15. <span id="cite_note-gpt4o-15">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_ref-gpt4o_15-0) OpenAI. *Hello GPT‑4o*. (2024). <a href="https://openai.com/index/hello-gpt-4o/" class="external text" rel="nofollow">openai.com/index/hello-gpt-4o/</a>.</span>
16. <span id="cite_note-ragas-16">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_ref-ragas_16-0) Es, S. et al. (2023). *RAGAS: Automated Evaluation of Retrieval Augmented Generation*. <a href="https://arxiv.org/abs/2309.15217" class="external text" rel="nofollow">arXiv:2309.15217</a>.</span>
17. <span id="cite_note-infographicvqa2021-17">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_ref-infographicvqa2021_17-0) Mathew, M. et al. (2021). *InfographicVQA: Understanding Infographics via Question Answering*. ICDAR. <a href="https://arxiv.org/abs/2104.12756" class="external text" rel="nofollow">arXiv:2104.12756</a>.</span>
18. <span id="cite_note-chartqa2022-18">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_ref-chartqa2022_18-0) Masry, A. et al. (2022). *ChartQA: A Benchmark for Question Answering about Charts*. ACL (Findings). <a href="https://arxiv.org/abs/2103.16435" class="external text" rel="nofollow">arXiv:2103.16435</a>.</span>
19. <span id="cite_note-mmdocrag2025-19">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_ref-mmdocrag2025_19-0) Dong, K. et al. (2025). *Benchmarking Retrieval‑Augmented Multimodal Generation for Document QA (MMDocRAG)*. <a href="https://arxiv.org/abs/2505.16470" class="external text" rel="nofollow">arXiv:2505.16470</a>.</span>
20. <span id="cite_note-weaviate-hybrid-20">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_ref-weaviate-hybrid_20-0) Weaviate Docs. *Hybrid search*. <a href="https://docs.weaviate.io/weaviate/concepts/search/hybrid-search" class="external text" rel="nofollow">docs.weaviate.io/.../hybrid-search</a>.</span>
21. <span id="cite_note-colpali2024-21">[↑](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ID)#cite_ref-colpali2024_21-0) Faysse, M. et al. (2024). *ColPali: Efficient Document Retrieval with Vision‑Language Models*. <a href="https://arxiv.org/abs/2407.01449" class="external text" rel="nofollow">arXiv:2407.01449</a>.</span>
