---
title: "RAG patterns — RAG पैटर्न"
source: "https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8"
wiki: "systems-analysis.info/int"
article: "RAG_patterns_—_RAG_पैटर्न"
language: "hi"
categories:
  - "Category:Hindi"
  - "Category:Large language models"
  - "Category:Prompt engineering"
revision_id: 6117
wiki_created_at: 2026-09-06T23:58:34Z
wiki_modified_at: 2026-09-06T23:58:34Z
downloaded_at: 2026-09-07T23:12:08Z
---

# RAG patterns — RAG पैटर्न

**RAG पैटर्न** (अंग्रेज़ी: *RAG Patterns*) — **Retrieval-Augmented Generation** (RAG) प्रणालियों के निर्माण के लिए वास्तुकला और पद्धतिगत दृष्टिकोणों का एक समुच्चय है। ये पैटर्न बड़े भाषा मॉडलों (LLM) की मूलभूत समस्याओं — जैसे कि hallucination, ज्ञान का पुराना पड़ जाना और डोमेन-विशिष्टता की कमी — को हल करने के लिए बनाए गए हैं। इसके लिए LLM को बाहरी, गतिशील रूप से उपलब्ध डेटा स्रोतों के साथ एकीकृत किया जाता है<sup>[\[1\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-lewis2020-1)</sup>। RAG का विकास सरल रैखिक पाइपलाइनों से लेकर जटिल मॉड्यूलर और एजेंट-आधारित प्रणालियों तक हुआ है<sup>[\[2\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-survey2024-2)</sup>।

## RAG के मूलभूत पैटर्न

तकनीक के विकास के साथ अनेक RAG पैटर्न सामने आए हैं, जिनमें से प्रत्येक विशिष्ट समस्याओं को हल करता है और गुणवत्ता, गति तथा लागत के बीच अपने-अपने संतुलन बिंदु रखता है।

- **Classic RAG (क्लासिक RAG)** — मूलभूत दृष्टिकोण, जिसमें उपयोगकर्ता के प्रश्न को vectorize करके vector डेटाबेस में प्रासंगिक अंश (chunk) खोजे जाते हैं; मिले हुए chunk प्रश्न के साथ LLM को दिए जाते हैं जो उत्तर उत्पन्न करता है<sup>[\[1\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-lewis2020-1)</sup>।

<!-- -->

- **Multi‑Query RAG (बहु-प्रश्न RAG)** — LLM मूल प्रश्न के कई पुनः-शब्दांकित/परिष्कृत रूप उत्पन्न करता है; सभी रूपों पर खोज की जाती है और परिणाम एकत्रित किए जाते हैं, जिससे पूर्णता (*recall*) बढ़ती है<sup>[\[3\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-langchain-multiquery-3)</sup>।

<!-- -->

- **HyDE (Hypothetical Document Expansion)** — छोटे प्रश्न और लंबे दस्तावेज़ों के बीच «अर्थात्मक अंतर» को पाटने के लिए। LLM पहले एक «काल्पनिक» उत्तर-दस्तावेज़ उत्पन्न करता है, फिर उसके embedding का उपयोग खोज के लिए किया जाता है, जिससे प्रायः retrieval की गुणवत्ता सुधरती है<sup>[\[4\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-hyde-4)</sup>।

<!-- -->

- **Hybrid Retrieval (हाइब्रिड खोज)** — semantic (vector) और lexical (BM25) खोज का संयोजन। Hybrid योजनाएँ production प्रणालियों के लिए मानक बन चुकी हैं: vector खोज अर्थगत मिलान को कवर करती है, जबकि BM25 सटीक शब्द/ID/acronym खोजता है; परिणाम fusion द्वारा एकत्रित किए जाते हैं<sup>[\[5\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-weaviate-hybrid-5)[\[6\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-qdrant-hybrid-6)[\[7\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-milvus-fulltext-7)</sup>।

<!-- -->

- **Re‑ranking (पुनः-क्रमांकन)** — दो चरणों की प्रक्रिया: एक तीव्र retriever उम्मीदवारों का समुच्चय देता है (जैसे शीर्ष-100), फिर cross-encoder (या अन्य reranker) प्रासंगिकता की पुनर्गणना करके LLM के लिए सर्वोत्तम (जैसे शीर्ष-5) चुनता है<sup>[\[8\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-nogueira2019-8)[\[9\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-cohere-rerank-9)</sup>।

<!-- -->

- **Query Routing (प्रश्न-मार्गनिर्देशन)** — अनेक विविध डेटा स्रोतों वाली प्रणालियों (भिन्न index/DB/API) में प्रश्न को router (LLM-selector या classifier) की सहायता से सर्वोत्तम स्रोत तक भेजा जाता है; इसमें fallback रणनीतियाँ भी शामिल हैं<sup>[\[10\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-llama-router-10)</sup>।

<!-- -->

- **Agentic/Web RAG (एजेंट RAG)** — LLM एक agent की भूमिका में होता है: जटिल प्रश्नों को विघटित करता है, पुनरावृत्तियाँ योजनाबद्ध करता है और प्रतिक्रिया के साथ उपकरणों (vector खोज, वेब खोज) का उपयोग करता है। सामान्य कार्यान्वयन ReAct प्रतिमान है<sup>[\[11\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-react-11)</sup>; वेब-उन्मुख संग्रह और अनिवार्य उद्धरण के लिए WebGPT देखें<sup>[\[12\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-webgpt-12)</sup>।

### संबद्ध और उभरते प्रतिमान

- **GraphRAG (ग्राफ RAG)** — संदर्भ-चयन के स्रोत और तंत्र के रूप में ज्ञान ग्राफ का उपयोग करता है; खोज सत्ताओं के बीच संबंध-संरचना और पाठ पर आधारित होती है, जिससे multi-hop प्रश्नों पर व्याख्यायोग्यता और गुणवत्ता बढ़ती है<sup>[\[13\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-graphrag-13)[\[14\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-graphrag-project-14)</sup>।
- **MM‑RAG (मल्टीमोडल RAG)** — पाठ और दृश्य स्रोतों (स्कैन/आरेख/तालिकाएँ) के साथ कार्य करता है। उदाहरण: VisRAG, मल्टीमोडल दस्तावेज़ों पर VLM-उन्मुख retrieval और generation प्रदर्शित करता है<sup>[\[15\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-visrag-15)</sup>।
- **Packaging & Context Handling (संदर्भ-पैकेजिंग)** — मिले हुए chunk को prompt में एकीकृत करने के तरीके: *Stuff*, *Map‑Reduce*, *Refine*, *Tree‑of‑Chunks (RAPTOR)*<sup>[\[16\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-raptor-16)</sup>।

## पैटर्नों की तुलनात्मक तालिका

| पैटर्न                 | कब लागू करें                                          | गुणवत्ता पर प्रभाव                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                         | लागत / विलंब | जोखिम और सीमाएँ                                   |
|----------------------|----------------------------------------------------|-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------|--------------------------------------------------|
| **Classic RAG**      | सजातीय आधार पर PoC और सरल Q&A                      | आधारभूत स्तर; embedding पर अत्यधिक निर्भर<sup>[\[1\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-lewis2020-1)</sup>                                                                                                                                                                                                                                                                                                                      | कम          | शब्द-चयन के प्रति संवेदनशीलता; अप्रासंगिक संदर्भ का जोखिम |
| **Hybrid Retrieval** | अधिकांश production परिदृश्यों में; अनेक कोड/acronym/ID हों | पूर्णता बढ़ाता है; सटीक शब्दों को कवर करता है<sup>[\[5\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-weaviate-hybrid-5)[\[6\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-qdrant-hybrid-6)[\[7\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-milvus-fulltext-7)</sup> | कम/मध्यम     | fusion भार का समायोजन; दो index                  |
| **Re‑ranking**       | जब उच्च सटीकता महत्त्वपूर्ण हो                          | शीर्ष-k पर precision में उल्लेखनीय वृद्धि<sup>[\[8\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-nogueira2019-8)[\[9\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-cohere-rerank-9)</sup>                                                                                                                                                                | मध्यम/उच्च    | अतिरिक्त विलंब/लागत                                |
| **Multi‑Query**      | संक्षिप्त/बहुआयामी प्रश्न                                | recall बढ़ाता है<sup>[\[3\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-langchain-multiquery-3)</sup>                                                                                                                                                                                                                                                                                                                                  | मध्यम        | अनावश्यक/शोरयुक्त पुनः-शब्दांकन                        |
| **HyDE**             | बड़े «अर्थात्मक अंतर» वाले संक्षिप्त/अस्पष्ट प्रश्न             | *zero‑shot* retrieval की गुणवत्ता सुधारता है<sup>[\[4\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-hyde-4)</sup>                                                                                                                                                                                                                                                                                                                        | मध्यम        | «काल्पनिक» पाठ की गुणवत्ता पर निर्भर                 |
| **Query Routing**    | अनेक स्रोत (दस्तावेज़-आधार, SQL, API, वेब)               | सही स्रोत के कारण प्रासंगिकता बढ़ती है<sup>[\[10\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-llama-router-10)</sup>                                                                                                                                                                                                                                                                                                                      | मध्यम        | मार्ग की त्रुटि = खोज की विफलता                     |
| **Agentic/Web RAG**  | जटिल, शोधपरक, बहु-चरणीय प्रश्न                        | रैखिक पाइपलाइन की सीमाओं से परे समस्याएँ हल करता है<sup>[\[11\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-react-11)[\[12\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-webgpt-12)</sup>                                                                                                                                                                | उच्च         | जटिलता, अनंत-लूप का जोखिम; guardrail आवश्यक         |

मुख्य RAG पैटर्नों की तुलना

## व्यावहारिक कार्यान्वयन और वास्तुकला

### कार्यान्वयन के चरण

1.  **Proof of Concept (PoC):** embedding की गुणवत्ता और आधारभूत retrieval जाँचने के लिए सीमित किंतु प्रतिनिधि डेटासेट पर **Classic RAG** से आरंभ करें<sup>[\[1\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-lewis2020-1)</sup>।
2.  **Minimum Viable Product (MVP):** «प्रयास/प्रभाव» के सर्वोत्तम अनुपात के रूप में **Hybrid Retrieval** और **Re‑ranking** लागू करें<sup>[\[5\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-weaviate-hybrid-5)[\[8\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-nogueira2019-8)</sup>।
3.  **Production:** प्रश्न-रूपांतरण (**HyDE**, **Multi‑Query**) और आवश्यकतानुसार **Query Routing** जोड़ें; observability (retrieval/rerank/उत्तरों का लॉगिंग) और A/B परीक्षण स्थापित करें<sup>[\[3\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-langchain-multiquery-3)[\[10\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-llama-router-10)</sup>।

### मुख्य घटक

- **Chunking (चंकिंग):** गुणवत्ता के सबसे महत्त्वपूर्ण कारकों में से एक। सरल निश्चित-आकार का तरीका प्रायः अर्थ-इकाइयों को तोड़ देता है। संरचना-उन्मुख (markup के अनुसार) या recursive splitter (अनुच्छेद → वाक्य → शब्द) अनुशंसित हैं<sup>[\[17\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-rcsplit-17)[\[18\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-llama-hier-18)</sup>।
- **Embedding और metadata:** प्रत्येक chunk के साथ document_id, पृष्ठ/अनुभाग, शीर्षक, दिनांक संग्रहीत करें; यह फ़िल्टरिंग और स्रोतों के सटीक उद्धरण के लिए आवश्यक है।
- **Hybrid Retrieval और Rerank:** fusion (या RRF) के साथ BM25+vector का उपयोग करें, फिर उम्मीदवारों के छोटे pool पर cross-encoder द्वारा पुनः-क्रमांकन करें<sup>[\[5\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-weaviate-hybrid-5)[\[6\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-qdrant-hybrid-6)[\[8\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-nogueira2019-8)</sup>।
- **संदर्भ-पैकेजिंग:** लंबे corpus के लिए *Map‑Reduce*, *Refine* या *Tree‑of‑Chunks* चुनें<sup>[\[16\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-raptor-16)[\[18\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-llama-hier-18)</sup>।

### सामान्य त्रुटियाँ (एंटीपैटर्न)

- **केवल vector खोज** BM25 के बिना → कोड/ID/acronym पर विफलता<sup>[\[5\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-weaviate-hybrid-5)[\[7\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-milvus-fulltext-7)</sup>।
- **अत्यधिक बड़ा/छोटा chunk** → संदर्भ का नुकसान या embedding का «विसरण»<sup>[\[17\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-rcsplit-17)</sup>।
- **Production में rerank का अभाव** → LLM को शोरयुक्त संदर्भ मिलता है<sup>[\[8\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-nogueira2019-8)</sup>।
- **Observability और स्रोत-tracing का अभाव** → त्रुटियों के कारणों का विश्लेषण असंभव (RAG मूल्यांकन देखें)।

## गुणवत्ता मूल्यांकन और मेट्रिक्स

मूल्यांकन retriever स्तर (ऑफ़लाइन) और end‑to‑end (generation) स्तर पर किया जाता है।

### Retriever मेट्रिक्स

- **Hit Rate, Recall@k, MRR** — प्रासंगिक दस्तावेज़ों का आवरण और स्थिति।
- **Context Precision & Recall** — निकाला गया संदर्भ कितना «शोर-रहित» है और सब आवश्यक को कितना कवर करता है (RAGAS में कार्यान्वित)<sup>[\[19\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-ragas-19)</sup>।

### Generator मेट्रिक्स (end‑to‑end)

- **Faithfulness / Groundedness** — उत्तर का प्रदत्त संदर्भ से अनुरूपता।
- **Answer Relevancy (उत्तर-प्रासंगिकता)** — मूल प्रश्न से अनुरूपता।

मेट्रिक्स को स्वचालित करने के लिए open‑source फ्रेमवर्क उपयोग किए जाते हैं: **RAGAS**, **TruLens** (*RAG triad*: context relevance, groundedness, answer relevance), **DeepEval**<sup>[\[20\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-trulens-20)[\[21\]](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_note-deepeval-21)</sup>।

## यह भी देखें

- Retrieval-Augmented Generation (RAG)
- वेक्टर डेटाबेस
- Embedding
- AI-एजेंट
- GraphRAG
- MM-RAG
- LLM मूल्यांकन और benchmark

## साहित्य

- Lewis, P., Perez, E., et al. (2020). *Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks*. NeurIPS. arXiv:2005.11401.
- Fan, W., Ding, Y., et al. (2024). *A Survey on RAG Meeting LLMs: Towards Retrieval-Augmented Large Language Models*. KDD. DOI:10.1145/3637528.3671470; arXiv:2405.06211.
- Gao, L., Ma, X., Lin, J., Callan, J. (2023). *Precise Zero‑Shot Dense Retrieval without Relevance Labels (HyDE)*. ACL 2023. ACL Anthology; arXiv:2212.10496.
- Nogueira, R., Cho, K. (2019). *Passage Re‑ranking with BERT*. arXiv:1901.04085.
- Weaviate Docs. *Hybrid search (BM25+Vector)*. <a href="https://docs.weaviate.io/weaviate/concepts/search/hybrid-search" class="external autonumber" rel="nofollow">[१]</a>.
- Qdrant Docs. *Hybrid Queries*. <a href="https://qdrant.tech/documentation/concepts/hybrid-queries/" class="external autonumber" rel="nofollow">[२]</a>.
- Milvus Docs. *Full‑Text Search* / *Hybrid Search*. <a href="https://milvus.io/docs/full-text-search.md" class="external autonumber" rel="nofollow">[३]</a> / <a href="https://milvus.io/docs/hybrid_search_with_milvus.md" class="external autonumber" rel="nofollow">[४]</a>.
- LangChain Docs. *MultiQueryRetriever*. <a href="https://python.langchain.com/docs/how_to/MultiQueryRetriever/" class="external autonumber" rel="nofollow">[५]</a>.
- Cohere Docs. *Rerank — best practices*. <a href="https://docs.cohere.com/docs/reranking-best-practices" class="external autonumber" rel="nofollow">[६]</a>.
- LlamaIndex Docs. *Routing (query routers/selectors)*. <a href="https://docs.llamaindex.ai/en/stable/module_guides/querying/router/" class="external autonumber" rel="nofollow">[७]</a>.
- Yao, S., et al. (2023). *ReAct: Synergizing Reasoning and Acting in Language Models*. ICLR. arXiv:2210.03629.
- Nakano, R., et al. (2021). *WebGPT: Browser‑assisted question‑answering with human feedback*. arXiv:2112.09332.
- Microsoft Research Blog. *GraphRAG: Unlocking LLM discovery on narrative private data*. (2024). <a href="https://www.microsoft.com/en-us/research/blog/graphrag-unlocking-llm-discovery-on-narrative-private-data/" class="external autonumber" rel="nofollow">[८]</a>.
- Microsoft Research. *Project GraphRAG*. (2024). <a href="https://www.microsoft.com/en-us/research/project/graphrag/" class="external autonumber" rel="nofollow">[९]</a>.
- Yu, S., et al. (2024). *VisRAG: Vision‑based Retrieval‑augmented Generation on Multi‑modality Documents*. arXiv:2410.10594.
- Sarthi, P., et al. (2024). *RAPTOR: Recursive Abstractive Processing for Tree‑Organized Retrieval*. arXiv:2401.18059.
- Es, S., et al. (2024). *RAGAs: Automated Evaluation of Retrieval Augmented Generation*. EACL (Demo). <a href="https://aclanthology.org/2024.eacl-demo.16/" class="external autonumber" rel="nofollow">[१०]</a>.
- TruLens Docs. *RAG Triad*. <a href="https://www.trulens.org/getting_started/core_concepts/rag_triad/" class="external autonumber" rel="nofollow">[११]</a>.
- DeepEval (GitHub). *The LLM Evaluation Framework*. <a href="https://github.com/confident-ai/deepeval" class="external autonumber" rel="nofollow">[१२]</a>.
- LangChain Docs. *RecursiveCharacterTextSplitter*. <a href="https://python.langchain.com/docs/how_to/recursive_text_splitter/" class="external autonumber" rel="nofollow">[१३]</a>.
- LlamaIndex Docs. *HierarchicalNodeParser*; *Response Synthesis (Tree/Refine)*. <a href="https://docs.llamaindex.ai/en/stable/api/llama_index.core.node_parser.HierarchicalNodeParser.html" class="external autonumber" rel="nofollow">[१४]</a>; <a href="https://docs.llamaindex.ai/en/stable/examples/low_level/response_synthesis/" class="external autonumber" rel="nofollow">[१५]</a>.

## टिप्पणियाँ

1.  <span id="cite_note-lewis2020-1">↑ <sup>[1.0](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-lewis2020_1-0)</sup> <sup>[1.1](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-lewis2020_1-1)</sup> <sup>[1.2](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-lewis2020_1-2)</sup> <sup>[1.3](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-lewis2020_1-3)</sup> Lewis, P., Perez, E., et al. (2020). *Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks*. NeurIPS. arXiv:2005.11401.</span>
2.  <span id="cite_note-survey2024-2">[↑](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-survey2024_2-0) Fan, W., Ding, Y., et al. (2024). *A Survey on RAG Meeting LLMs: Towards Retrieval-Augmented Large Language Models*. KDD. DOI:10.1145/3637528.3671470; arXiv:2405.06211.</span>
3.  <span id="cite_note-langchain-multiquery-3">↑ <sup>[3.0](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-langchain-multiquery_3-0)</sup> <sup>[3.1](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-langchain-multiquery_3-1)</sup> <sup>[3.2](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-langchain-multiquery_3-2)</sup> LangChain Docs. *MultiQueryRetriever*. <a href="https://python.langchain.com/docs/how_to/MultiQueryRetriever/" class="external free" rel="nofollow">https://python.langchain.com/docs/how_to/MultiQueryRetriever/</a></span>
4.  <span id="cite_note-hyde-4">↑ <sup>[4.0](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-hyde_4-0)</sup> <sup>[4.1](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-hyde_4-1)</sup> Gao, L., Ma, X., Lin, J., Callan, J. (2023). *Precise Zero‑Shot Dense Retrieval without Relevance Labels*. ACL 2023. arXiv:2212.10496; ACL Anthology: 2023.acl‑long.99.</span>
5.  <span id="cite_note-weaviate-hybrid-5">↑ <sup>[5.0](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-weaviate-hybrid_5-0)</sup> <sup>[5.1](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-weaviate-hybrid_5-1)</sup> <sup>[5.2](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-weaviate-hybrid_5-2)</sup> <sup>[5.3](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-weaviate-hybrid_5-3)</sup> <sup>[5.4](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-weaviate-hybrid_5-4)</sup> Weaviate Docs. *Hybrid search (BM25+Vector)*. <a href="https://docs.weaviate.io/weaviate/concepts/search/hybrid-search" class="external free" rel="nofollow">https://docs.weaviate.io/weaviate/concepts/search/hybrid-search</a></span>
6.  <span id="cite_note-qdrant-hybrid-6">↑ <sup>[6.0](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-qdrant-hybrid_6-0)</sup> <sup>[6.1](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-qdrant-hybrid_6-1)</sup> <sup>[6.2](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-qdrant-hybrid_6-2)</sup> Qdrant Docs. *Hybrid Queries*. <a href="https://qdrant.tech/documentation/concepts/hybrid-queries/" class="external free" rel="nofollow">https://qdrant.tech/documentation/concepts/hybrid-queries/</a></span>
7.  <span id="cite_note-milvus-fulltext-7">↑ <sup>[7.0](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-milvus-fulltext_7-0)</sup> <sup>[7.1](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-milvus-fulltext_7-1)</sup> <sup>[7.2](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-milvus-fulltext_7-2)</sup> Milvus Docs. *Full‑Text Search* и *Hybrid Search*. <a href="https://milvus.io/docs/full-text-search.md" class="external free" rel="nofollow">https://milvus.io/docs/full-text-search.md</a>; <a href="https://milvus.io/docs/hybrid_search_with_milvus.md" class="external free" rel="nofollow">https://milvus.io/docs/hybrid_search_with_milvus.md</a></span>
8.  <span id="cite_note-nogueira2019-8">↑ <sup>[8.0](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-nogueira2019_8-0)</sup> <sup>[8.1](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-nogueira2019_8-1)</sup> <sup>[8.2](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-nogueira2019_8-2)</sup> <sup>[8.3](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-nogueira2019_8-3)</sup> <sup>[8.4](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-nogueira2019_8-4)</sup> Nogueira, R., Cho, K. (2019). *Passage Re‑ranking with BERT*. arXiv:1901.04085.</span>
9.  <span id="cite_note-cohere-rerank-9">↑ <sup>[9.0](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-cohere-rerank_9-0)</sup> <sup>[9.1](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-cohere-rerank_9-1)</sup> Cohere Docs. *Rerank — best practices*. <a href="https://docs.cohere.com/docs/reranking-best-practices" class="external free" rel="nofollow">https://docs.cohere.com/docs/reranking-best-practices</a></span>
10. <span id="cite_note-llama-router-10">↑ <sup>[10.0](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-llama-router_10-0)</sup> <sup>[10.1](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-llama-router_10-1)</sup> <sup>[10.2](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-llama-router_10-2)</sup> LlamaIndex Docs. *Routing (query routers/selectors)*. <a href="https://docs.llamaindex.ai/en/stable/module_guides/querying/router/" class="external free" rel="nofollow">https://docs.llamaindex.ai/en/stable/module_guides/querying/router/</a></span>
11. <span id="cite_note-react-11">↑ <sup>[11.0](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-react_11-0)</sup> <sup>[11.1](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-react_11-1)</sup> Yao, S., et al. (2023). *ReAct: Synergizing Reasoning and Acting in Language Models*. ICLR 2023. arXiv:2210.03629.</span>
12. <span id="cite_note-webgpt-12">↑ <sup>[12.0](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-webgpt_12-0)</sup> <sup>[12.1](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-webgpt_12-1)</sup> Nakano, R., et al. (2021). *WebGPT: Browser‑assisted question‑answering with human feedback*. arXiv:2112.09332.</span>
13. <span id="cite_note-graphrag-13">[↑](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-graphrag_13-0) Microsoft Research Blog. *GraphRAG: Unlocking LLM discovery on narrative private data*. 2024. <a href="https://www.microsoft.com/en-us/research/blog/graphrag-unlocking-llm-discovery-on-narrative-private-data/" class="external free" rel="nofollow">https://www.microsoft.com/en-us/research/blog/graphrag-unlocking-llm-discovery-on-narrative-private-data/</a></span>
14. <span id="cite_note-graphrag-project-14">[↑](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-graphrag-project_14-0) Microsoft Research. *Project GraphRAG*. <a href="https://www.microsoft.com/en-us/research/project/graphrag/" class="external free" rel="nofollow">https://www.microsoft.com/en-us/research/project/graphrag/</a></span>
15. <span id="cite_note-visrag-15">[↑](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-visrag_15-0) Yu, S., et al. (2024). *VisRAG: Vision‑based Retrieval‑augmented Generation on Multi‑modality Documents*. arXiv:2410.10594; OpenReview: zG459X3Xge.</span>
16. <span id="cite_note-raptor-16">↑ <sup>[16.0](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-raptor_16-0)</sup> <sup>[16.1](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-raptor_16-1)</sup> Sarthi, P., et al. (2024). *RAPTOR: Recursive Abstractive Processing for Tree‑Organized Retrieval*. arXiv:2401.18059.</span>
17. <span id="cite_note-rcsplit-17">↑ <sup>[17.0](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-rcsplit_17-0)</sup> <sup>[17.1](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-rcsplit_17-1)</sup> LangChain Docs. *RecursiveCharacterTextSplitter*. <a href="https://python.langchain.com/docs/how_to/recursive_text_splitter/" class="external free" rel="nofollow">https://python.langchain.com/docs/how_to/recursive_text_splitter/</a></span>
18. <span id="cite_note-llama-hier-18">↑ <sup>[18.0](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-llama-hier_18-0)</sup> <sup>[18.1](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-llama-hier_18-1)</sup> LlamaIndex Docs. *HierarchicalNodeParser* и *Tree Summarization*. <a href="https://docs.llamaindex.ai/en/stable/api/llama_index.core.node_parser.HierarchicalNodeParser.html" class="external free" rel="nofollow">https://docs.llamaindex.ai/en/stable/api/llama_index.core.node_parser.HierarchicalNodeParser.html</a>; <a href="https://docs.llamaindex.ai/en/stable/examples/low_level/response_synthesis/" class="external free" rel="nofollow">https://docs.llamaindex.ai/en/stable/examples/low_level/response_synthesis/</a></span>
19. <span id="cite_note-ragas-19">[↑](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-ragas_19-0) Es, S., et al. (2024). *RAGAs: Automated Evaluation of Retrieval Augmented Generation*. EACL (Demo). <a href="https://aclanthology.org/2024.eacl-demo.16/" class="external free" rel="nofollow">https://aclanthology.org/2024.eacl-demo.16/</a></span>
20. <span id="cite_note-trulens-20">[↑](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-trulens_20-0) TruLens Docs. *RAG Triad*. <a href="https://www.trulens.org/getting_started/core_concepts/rag_triad/" class="external free" rel="nofollow">https://www.trulens.org/getting_started/core_concepts/rag_triad/</a></span>
21. <span id="cite_note-deepeval-21">[↑](https://systems-analysis.info/int/RAG_patterns_%E2%80%94_RAG_%E0%A4%AA%E0%A5%88%E0%A4%9F%E0%A4%B0%E0%A5%8D%E0%A4%A8#cite_ref-deepeval_21-0) DeepEval (GitHub). <a href="https://github.com/confident-ai/deepeval" class="external free" rel="nofollow">https://github.com/confident-ai/deepeval</a></span>
