MM-RAG (Multimodal RAG) (TH)
MM-RAG (อังกฤษ: Multimodal Retrieval-Augmented Generation) — คือการขยายกระบวนทัศน์ RAG แบบดั้งเดิม โดยที่ LLM ใช้ไม่เพียงแค่ข้อความในการตอบคำถาม แต่ยังใช้ข้อมูลเชิงภาพด้วย (รูปภาพ แผนภาพ ตาราง กราฟ) การค้นคืนแบบมัลติโมดัลช่วยให้สามารถค้นหาและเชื่อมโยงหลักฐานในรูปแบบต่าง ๆ ได้ ลดความเสี่ยงของการเกิดภาพหลอนโดยอาศัยแหล่งข้อมูลภายนอกที่มีการอ้างอิงแม่นยำถึงส่วนของหน้าและพื้นที่ (bounding boxes)[1][2].
MM-RAG มีประโยชน์เป็นพิเศษสำหรับเอกสารที่ส่วนสำคัญของความหมายถูกนำเสนอในรูปแบบที่ไม่ใช่ข้อความ (โครงร่างหน้า แผนภาพ โครงสร้างตาราง) ในกรณีเช่นนี้ RAG แบบข้อความดั้งเดิมมักสูญเสียองค์ประกอบบริบทที่สำคัญไป[3][4].
บริบทและปัญหาที่แก้ไข
RAG แบบดั้งเดิมทำงานกับข้อความและไม่สามารถมองเห็นโครงสร้างเชิงภาพได้ (การจัดวางองค์ประกอบ คำอธิบายรูป แกนกราฟ) MM-RAG ปิดช่องว่างเหล่านี้: ดึงองค์ประกอบที่มีโครงสร้าง (ข้อความ ตาราง รูปภาพพร้อมพิกัด) จัดทำดัชนีในปริภูมิเวกเตอร์ และรวมหลักฐานจากโมดาลิตีต่าง ๆ[5][6].
สถาปัตยกรรมของ MM-RAG
ไปป์ไลน์ MM-RAG เพิ่มขั้นตอนการประมวลผลข้อมูลเชิงภาพและการจัดแนวโมดาลิตีให้กับ RAG แบบดั้งเดิม ได้แก่: การนำเข้า → การจัดทำดัชนี → การค้นคืนมัลติโมดัล → การรวมและการจัดอันดับใหม่ → การสร้างข้อความพร้อมการติดตาม.
- การรวบรวมและการประมวลผลเบื้องต้น (Ingestion). ข้อมูลนำเข้าคือ PDF/สแกน/รูปภาพ ดำเนินการ OCR และ การวิเคราะห์โครงร่าง หน้าเพื่อแบ่งโซน ได้แก่ ย่อหน้า หัวเรื่อง ตาราง รูปภาพและพิกัดของสิ่งเหล่านี้ เครื่องมือทั่วไปได้แก่โมเดลในตระกูล LayoutLM และไลบรารีเครื่องมือ LayoutParser; การตรวจสอบและการฝึกอบรมมักอาศัย dataset PubLayNet และ DocLayNet[7][8][9].
- การแบ่งส่วนเป็นภูมิภาค. ดึงวัตถุเชิงภาพออกมา (แผนภาพ ตาราง ภาพประกอบ คำอธิบายภาพ) เพื่อความทนทานสูงขึ้นจึงใช้โมเดล OCR-free (เช่น Donut) หรือไปป์ไลน์แบบผสม OCR+VLM[10].
- การจัดทำดัชนี (Vector Index). ข้อความเป็นชิ้น ๆ และองค์ประกอบเชิงภาพ (รูปภาพหรือคำอธิบายของรูปภาพ) ถูกแปลงเป็นเวกเตอร์และบันทึกลงในฐานข้อมูลเวกเตอร์ สำหรับปริภูมิรวม text↔image จะใช้ CLIP หรือ SigLIP; สำหรับการผลิตจริงนิยมใช้ดัชนีมัลติโมดัล/มัลติเวกเตอร์ (หนึ่งวัตถุ — หลายเวกเตอร์)[6][11][12].
- การค้นคืนมัลติโมดัลและการจัดอันดับใหม่. ดำเนินการค้นหาข้อความและการค้นหาเชิงภาพแบบผสม; ผู้สมัคร (ย่อหน้า ตาราง รูปภาพ/ภูมิภาค) ถูกรวมกันและจัดอันดับใหม่ด้วยโมเดลที่ "หนักกว่า" (cross-encoder/LLM-reranker) เพื่อเพิ่มความแม่นยำ[13].
- การบรรจุบริบทและการสร้างข้อความ. ส่วนที่คัดเลือกแล้วถูกป้อนให้กับ LLM/VLM หากโมเดลเป็นมัลติโมดัล (เช่น GPT-4V/4o) รูปภาพสามารถป้อนได้โดยตรง; สำหรับ LLM แบบข้อความรูปภาพจะถูกแปลงเป็นคำอธิบายโดยละเอียดล่วงหน้า[14][15].
- การติดตามและการอ้างอิง. คำตอบมาพร้อมกับการอ้างอิงแบบคลิกได้ที่เชื่อมโยงไม่เพียงแค่กับเอกสาร/หน้า แต่ยังกับภูมิภาค (พิกัด) ด้วย ซึ่งช่วยเพิ่มระดับ grounding และความไว้วางใจของผู้ใช้[2].
การประเมินคุณภาพและเมตริก
ประสิทธิภาพของ MM-RAG ได้รับการประเมินในระดับการดึงข้อมูล การค้นคืน และการสร้างข้อความ
- คุณภาพการดึงข้อมูลเชิงภาพ. ความแม่นยำ OCR (WER/CER) คุณภาพการวิเคราะห์โครงร่าง (mAP/Precision/Recall) บน dataset DocLayNet/PubLayNet[8][7].
- คุณภาพการค้นคืน. เมตริกมาตรฐานของการค้นคืนสารสนเทศ: Recall@K, Precision@K, MRR; สำหรับมัลติโมดัล — แยกตามโมดาลิตีและในการรวมกัน
- คุณภาพคำตอบ (end-to-end). เมตริกอัตโนมัติ faithfulness/groundedness และการประเมินโดยมนุษย์ ในทางปฏิบัติใช้เฟรมเวิร์ก RAGAS/TruLens/DeepEval[16].
- Benchmark.
ตารางเปรียบเทียบส่วนประกอบ
| ส่วนประกอบ | รูปแบบการใช้งาน | ข้อดี | ข้อเสีย / ความเสี่ยง | เมื่อใดควรเลือก |
|---|---|---|---|---|
| OCR | Tesseract / PaddleOCR / API บนคลาวด์ | แบบท้องถิ่น — ความเป็นส่วนตัวและการควบคุม; แบบคลาวด์ — ความแม่นยำสูงพร้อมใช้งาน | ข้อผิดพลาดในโครงร่างที่ซับซ้อน; API — ค่าใช้จ่ายและข้อกำหนดด้านการปฏิบัติตามกฎระเบียบ | ข้อมูลส่วนตัว — OCR ท้องถิ่น; ความแม่นยำสูงสุด — คลาวด์ (หากอนุญาต) |
| การวิเคราะห์โครงร่าง | กฎ / โมเดล ML (LayoutLM, LayoutParser) | กฎเรียบง่ายสำหรับเทมเพลตที่เป็นแบบเดียวกัน; ML — ทนทานต่อความหลากหลาย | กฎพังกับโครงร่างใหม่; ML ต้องการทรัพยากร/ข้อมูล | แบบฟอร์มเดิม — กฎ; คลังข้อมูลหลากหลาย — ML |
| การแปลงเป็นเวกเตอร์ (รูปภาพ) | CLIP / SigLIP / คำอธิบาย OCR-free (Donut/Pix2Struct) | พื้นที่แฝงร่วม text↔image (CLIP/SigLIP); OCR-free ขจัดการพึ่งพา OCR | CLIP ไม่อ่านข้อความภายในรูปภาพ; คำอธิบายอาจบิดเบือนความหมาย | CLIP/SigLIP — การค้นหามัลติโมดัลพื้นฐาน; OCR-free สำหรับสแกนที่มีคุณภาพซับซ้อน |
| การรวมผลลัพธ์ | การเรียงลำดับตามคะแนน / โควตาตามโมดาลิตี / LLM-reranker | Reranker ช่วยเพิ่มความแม่นยำของการคัดเลือกบริบทอย่างเห็นได้ชัด | เพิ่มความหน่วงและค่าใช้จ่าย | สถานการณ์ที่ต้องการความแม่นยำสูง; วิธีการง่าย ๆ — สำหรับ PoC |
| พื้นที่เก็บข้อมูล/ดัชนี | เวกเตอร์เดี่ยว / มัลติเวกเตอร์ (text+image) / ไฮบริด (BM25+vector) | มัลติเวกเตอร์ครอบคลุมการแสดงผลต่าง ๆ ของวัตถุเดียว; ไฮบริดช่วยคำหลัก/รหัส | ความซับซ้อนของสคีมาและการอัปเดต | ระบบการผลิตที่มีข้อมูลผสมและ SLA เข้มงวด |
ข้อสังเกตเชิงปฏิบัติ
- การค้นหาแบบไฮบริด (BM25 + เวกเตอร์) — มาตรฐาน de-facto สำหรับการเพิ่มความครบถ้วนและความแม่นยำในคำศัพท์/รหัสเฉพาะทาง[20].
- การจัดอันดับใหม่ ด้วย cross-encoder/LLM ช่วยประหยัด token โดยคัดทิ้งผู้สมัครที่ "ไม่มีประโยชน์" ก่อนการสร้างข้อความ[13].
- VLM-retriever รุ่นใหม่ (เช่น ColPali) แสดงข้อได้เปรียบในเอกสารที่มีภาพมากโดยการจัดทำดัชนีหน้าในรูปแบบรูปภาพโดยตรง[21].
เอกสารอ้างอิง
- Lewis, P. et al. (2020). Retrieval‑Augmented Generation for Knowledge‑Intensive NLP Tasks. NeurIPS. arXiv:2005.11401.
- Gao, L. et al. (2023). Precise Zero‑Shot Dense Retrieval without Relevance Labels (HyDE). ACL 2023. arXiv:2212.10496.
- Mei, L., Mo, S., Yang, Z., Chen, C. (2025). A Survey of Multimodal Retrieval‑Augmented Generation. arXiv:2504.08748.
- Abootorabi, M.M. et al. (2025). Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval‑Augmented Generation. Findings of ACL 2025. ACL Anthology.
- Yu, S. et al. (2024). VisRAG: Vision‑based Retrieval‑augmented Generation on Multi‑modality Documents. arXiv:2410.10594.
- Cho, J. et al. (2024). M3DocRAG: Multi‑modal Retrieval is What You Need for Multi‑document QA. arXiv:2411.04952.
- Tanaka, R. et al. (2025). VDocRAG: Retrieval‑Augmented Generation over Visually‑Rich Documents. CVPR 2025. arXiv:2504.09795 • CVF Open Access.
- Dong, K. et al. (2025). MMDocRAG: Benchmarking Retrieval‑Augmented Multimodal Generation for Document Question Answering. arXiv:2505.16470.
- Wasserman, N. et al. (2025). REAL‑MM‑RAG: A Real‑World Multi‑Modal Retrieval Benchmark. arXiv:2502.12342.
- Faysse, M. et al. (2024). ColPali: Efficient Document Retrieval with Vision Language Models. arXiv:2407.01449.
- Radford, A. et al. (2021). Learning Transferable Visual Models From Natural Language Supervision (CLIP). ICML 2021. arXiv:2103.00020.
- Tschannen, M. et al. (2025). SigLIP 2: Multilingual Vision‑Language Encoders with Improved Semantic Understanding, Localization, and Dense Features. arXiv:2502.14786.
- Xu, Y. et al. (2020). LayoutLM: Pre‑training of Text and Layout for Document Image Understanding. KDD 2020. DOI • arXiv:1912.13318.
- Huang, Y. et al. (2022). LayoutLMv3: Pre‑training for Document AI with Unified Text and Image Masking. arXiv:2204.08387.
- Zhong, X., Tang, J., Jimeno‑Yepes, A.J. (2019). PubLayNet: Largest Dataset Ever for Document Layout Analysis. arXiv:1908.07836.
- Pfitzmann, B. et al. (2022). DocLayNet: A Large Human‑Annotated Dataset for Document‑Layout Analysis. arXiv:2206.01062.
- Kim, G. et al. (2021). OCR‑free Document Understanding Transformer (Donut). arXiv:2111.15664.
- Shen, Z. et al. (2021). LayoutParser: A Unified Toolkit for Deep Learning Based Document Image Analysis. arXiv:2103.15348.
- Singh, A. et al. (2019). Towards VQA Models That Can Read (TextVQA). CVPR 2019. arXiv:1904.08920.
- Mathew, M. et al. (2021/2022). DocVQA / InfographicVQA: Datasets for VQA on Document Images and Infographics. WACV 2021 / WACV 2022. CVF • arXiv:2104.12756.
- Masry, A. et al. (2022). ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning. Findings of ACL 2022. ACL • arXiv:2203.10244.
- Liu, F. et al. (2022). DePlot: One‑shot Visual Language Reasoning by Plot‑to‑Table Translation. arXiv:2212.10505.
- Wang, P. et al. (2024). Qwen2‑VL: Enhancing Vision‑Language Model's Capabilities in OCR and Chart QA. arXiv:2409.12191.
- Es, S. et al. (2023). RAGAS: Automated Evaluation of Retrieval‑Augmented Generation. arXiv:2309.15217.
ดูเพิ่มเติม
- Retrieval-Augmented Generation
- ฐานข้อมูลเวกเตอร์
- Embedding
- GraphRAG
หมายเหตุ
- ↑ Lewis, P., Perez, E., et al. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS. arXiv:2005.11401.
- ↑ 2.0 2.1 Yu, S. et al. (2024). VisRAG: Vision-based Retrieval-Augmented Generation on Multi-modality Documents. arXiv:2407.06437.
- ↑ 3.0 3.1 Mathew, M. et al. (2021). DocVQA: A Dataset for VQA on Document Images. WACV. arXiv:2007.00398.
- ↑ 4.0 4.1 Singh, A. et al. (2019). TextVQA: Towards VQA Models That Can Read. CVPR. arXiv:1904.08920.
- ↑ Xu, Y. et al. (2020). LayoutLM: Pre-training of Text and Layout for Document Image Understanding. KDD. DOI.
- ↑ 6.0 6.1 Radford, A. et al. (2021). Learning Transferable Visual Models From Natural Language Supervision (CLIP). ICML. arXiv:2103.00020.
- ↑ 7.0 7.1 Zhong, X., Tang, J., Yepes, A. J. (2019). PubLayNet: Largest Dataset for Document Layout Analysis. ICDAR. arXiv:1908.07836.
- ↑ 8.0 8.1 Pfitzmann, B. et al. (2022). DocLayNet: A Large Human‑Annotated Dataset for Document‑Layout Analysis. KDD. DOI / arXiv:2206.01062.
- ↑ Shen, Z. et al. (2021). LayoutParser: A Unified Toolkit for DL‑based Document Image Analysis. arXiv:2103.15348.
- ↑ Kim, G. et al. (2021). Donut: OCR‑free Document Understanding Transformer. arXiv:2111.15664.
- ↑ Zhai, X. et al. (2023). Sigmoid Loss for Language‑Image Pre‑Training (SigLIP). ICCV. arXiv:2303.15343.
- ↑ Milvus Docs. Multi‑Vector Hybrid Search. milvus.io/docs/multi-vector-search.md.
- ↑ 13.0 13.1 Cohere Docs. Rerank API. docs.cohere.com/reference/rerank.
- ↑ OpenAI. GPT‑4V(ision) System Card. (2023). PDF.
- ↑ OpenAI. Hello GPT‑4o. (2024). openai.com/index/hello-gpt-4o/.
- ↑ Es, S. et al. (2023). RAGAS: Automated Evaluation of Retrieval Augmented Generation. arXiv:2309.15217.
- ↑ Mathew, M. et al. (2021). InfographicVQA: Understanding Infographics via Question Answering. ICDAR. arXiv:2104.12756.
- ↑ Masry, A. et al. (2022). ChartQA: A Benchmark for Question Answering about Charts. ACL (Findings). arXiv:2103.16435.
- ↑ Dong, K. et al. (2025). Benchmarking Retrieval‑Augmented Multimodal Generation for Document QA (MMDocRAG). arXiv:2505.16470.
- ↑ Weaviate Docs. Hybrid search. docs.weaviate.io/.../hybrid-search.
- ↑ Faysse, M. et al. (2024). ColPali: Efficient Document Retrieval with Vision‑Language Models. arXiv:2407.01449.