---
title: "Hypothetical Document Embeddings (HyDE) (BN)"
source: "https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)"
wiki: "systems-analysis.info/int"
article: "Hypothetical_Document_Embeddings_(HyDE)_(BN)"
language: "bn"
categories:
  - "Category:Bengali"
  - "Category:Large language models"
  - "Category:Pages with math errors"
  - "Category:Pages with math render errors"
  - "Category:Prompt engineering"
revision_id: 3106
wiki_created_at: 2026-09-06T23:15:47Z
wiki_modified_at: 2026-09-06T23:15:47Z
downloaded_at: 2026-09-07T22:55:11Z
---

# Hypothetical Document Embeddings (HyDE) (BN)

**Hypothetical Document Expansion (HyDE)** — ভেক্টর retrieval এবং retrieval‑augmented generation (RAG) উন্নত করার একটি পদ্ধতি, যেখানে একটি বৃহৎ ভাষা মডেল (LLM) মূল প্রশ্নের ভিত্তিতে একটি «হাইপোথেটিক্যাল ডকুমেন্ট» তৈরি করে; তারপর এই টেক্সটটি একটি encoder দ্বারা ভেক্টরে রূপান্তরিত হয় এবং প্রাপ্ত ভেক্টরের সাথে নৈকট্যের ভিত্তিতে প্রকৃত ডকুমেন্টগুলির মধ্যে অনুসন্ধান পরিচালিত হয়। এই পদ্ধতি LLM-এ এনকোড করা «প্রাসঙ্গিকতার প্যাটার্ন» ব্যবহার করতে এবং ঘন embedding-এর মাধ্যমে সেগুলিকে corpus-এ «গ্রাউন্ড» করতে সক্ষম করে<sup>[\[1\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-1)</sup>।

## সংজ্ঞা এবং স্বজ্ঞা

HyDE অনুসন্ধানের কাজকে দুটি ধাপে বিভক্ত করে:

\(1\) LLM প্রশ্নের জন্য একটি «প্রাসঙ্গিক উত্তরের উদাহরণ» (*hypothetical document*) তৈরি করে, এর মাধ্যমে প্রাসঙ্গিকতার বৈশিষ্ট্যগুলি মডেল করে;

\(2\) একটি contrastive encoder (যেমন Contriever) এই টেক্সটকে একটি ভেক্টরে রূপান্তরিত করে, যার মাধ্যমে index থেকে প্রকৃত ডকুমেন্টগুলি বের করা হয়। তৈরি করা টেক্সটে তথ্যগত ত্রুটি থাকতে পারে, কিন্তু গুরুত্বপূর্ণ হল encoder দ্বারা ধরা পড়া বিষয়গত ও পরিভাষাগত প্যাটার্নগুলি<sup>[\[2\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-2)</sup>।

## ইতিহাস এবং উৎস

কৃত্রিম টেক্সট দিয়ে অনুসন্ধান প্রসারিত করার ধারণাটি query expansion এবং pseudo-relevant feedback (PRF) সংক্রান্ত গবেষণায় ফিরে যায়: Rocchio অ্যালগরিদম এবং Relevance Language Model<sup>[\[3\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-3)[\[4\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-4)</sup>। ঘন retrieval-এর জন্য contrastively প্রশিক্ষিত encoder (Contriever)<sup>[\[5\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-5)</sup> এবং Dense Passage Retrieval (DPR)<sup>[\[6\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-6)</sup> ব্যবহার করা হয়েছিল। BEIR benchmark zero‑shot মূল্যায়ন মানসম্পন্ন করেছে<sup>[\[7\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-7)</sup>। এই প্রেক্ষাপটে HyDE প্রস্তাব করা হয়েছে encoder-এর fine-tuning ছাড়াই LLM-এর মাধ্যমে শূন্য-মোডে প্রাসঙ্গিকতার জ্ঞান «আনয়ন» করার উপায় হিসেবে<sup>[\[8\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-8)</sup>।

## পদ্ধতি এবং আনুষ্ঠানিকতা

ধরা যাক ডকুমেন্টের corpus হল $\mathcal{D} = \{ d_{1},\ldots,d_{N}\}$, এবং টেক্সট encoder $E:\text{text} \rightarrow {\mathbb{R}}^{n}$ ডকুমেন্টের ভেক্টর উপস্থাপনা $\mathbf{v}_{d} = E(d)$ নির্ধারণ করে। নৈকট্য পরিমাপের জন্য cosine similarity বা dot product ব্যবহার করা হয়; একটি গুরুত্বপূর্ণ মন্তব্য: \*\*dot product কেবলমাত্র তখনই cosine similarity-র সমান হয় যখন উভয় ভেক্টরের L2‑norm একক হয়\*\* ($\|\mathbf{u}\| = \|\mathbf{v}\| = 1$)<sup>[\[9\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-9)</sup>।

HyDE LLM দ্বারা তৈরি «হাইপোথেটিক্যাল ডকুমেন্ট»-এর মাধ্যমে প্রশ্নের উপস্থাপনাকে পুনর্নির্ধারণ করে। আনুষ্ঠানিকভাবে:

$$
\begin{matrix}
 & \text{(1) Генерация гипотетического текста:} & & {\overset{\sim}{d}\; = \; G\!\left( q;\,{inst} \right),} \\
 & \text{(2) Эмбеддинг гипотетического текста:} & & {\mathbf{v}_{h}\; = \; E(\overset{\sim}{d}),} \\
 & \text{(3) Поиск ближайших соседей:} & & {\mathcal{R}_{k}(q)\; = \;{TopK}_{\, d \in \mathcal{D}}\; S\!\left( \mathbf{v}_{h},\mathbf{v}_{d} \right),}
\end{matrix}
$$

যেখানে $G$ — নির্দেশনা $inst$ সহ LLM (যেমন: «প্রশ্নটির উত্তর দিয়ে একটি অনুচ্ছেদ লিখুন …»), $S$ — similarity পরিমাপ (normalization সহ cosine বা IP), এবং $\mathcal{R}_{k}(q)$ — সর্বোচ্চ similarity সহ $k$টি ডকুমেন্টের সেট<sup>[\[10\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-10)[\[11\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-11)</sup>।

ইঞ্জিনিয়ারিং অনুশীলনে প্রায়ই \*\*একাধিক\*\* হাইপোথেটিক্যাল টেক্সট তৈরি করা হয় এবং তাদের উপস্থাপনাগুলি একত্রিত করা হয়, যা স্থিতিশীলতা বাড়ায়:

$$
{\overset{\sim}{d}}^{(j)} = G\!\left( q;\,{inst},\xi_{j} \right),\quad\mathbf{v}_{h}\; = \;\frac{1}{m}\sum\limits_{j = 1}^{m}E\!\left( {\overset{\sim}{d}}^{(j)} \right),
$$

যেখানে $\xi_{j}$ — stochastic decoding প্যারামিটার (যেমন temperature/top‑p)। এই ধরনের ensembling মাঝারি latency বৃদ্ধিতে Recall উন্নত করে<sup>[\[12\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-12)</sup>।

### HyDE-এর মৌলিক pipeline

    # 1) prompt(query) -> hypothetical_doc
    # 2) embed(hypothetical_doc) -> v_h
    # 3) retrieve(index, v_h, k) -> candidates
    # 4) (optional) rerank(query, candidates) -> topN
    # 5) (для RAG) stuff / map-reduce / refine на topN

### অন্যান্য পদ্ধতির সাথে সম্পর্ক (QE, doc2query, PRF)

- **QE (query expansion)** প্রশ্নে term যোগ করে; HyDE পরিবর্তে একটি সম্পূর্ণ «quasi-document» তৈরি করে, যা ঘন encoder-এর সাথে আরও ভালোভাবে সামঞ্জস্যপূর্ণ<sup>[\[13\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-13)</sup>।
- **doc2query / docTTTTTquery** indexing-এর আগে কৃত্রিম প্রশ্ন দিয়ে **ডকুমেন্ট** প্রসারিত করে<sup>[\[14\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-14)[\[15\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-15)</sup>; HyDE তাৎক্ষণিকভাবে **প্রশ্ন** প্রসারিত করে, reindexing ছাড়াই।
- **PRF** (Rocchio, Relevance LM) শীর্ষ ফলাফলের উপর ভিত্তি করে প্রশ্নের ভেক্টর আপডেট করে; HyDE সরাসরি LLM থেকে «প্রাসঙ্গিকতার প্যাটার্ন» বের করে এবং তারপর corpus-এ retrieval দ্বারা সেটি «গ্রাউন্ড» করে<sup>[\[16\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-16)</sup>।

## RAG এবং reranking-এ একীকরণ

RAG-এ HyDE retrieval-এর প্রথম ধাপ হিসেবে প্রয়োগ করা হয়: hypothetical document → embedding → k প্রার্থী। এরপর reranking ব্যবহার করা হয়: BERT-শ্রেণীর cross-encoder<sup>[\[17\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-17)</sup> বা ColBERT late interaction<sup>[\[18\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-18)</sup>। তালিকা একীকরণের জন্য (যেমন BM25+vector হাইব্রিড) সাধারণত RRF (*reciprocal rank fusion*) প্রয়োগ করা হয়: $\operatorname{RRF}(d) = \sum\limits_{r \in \mathcal{R}}\frac{1}{k + \operatorname{rank}_{r}(d)},\qquad k \approx 60.$ RRF পদ্ধতি একত্রিত ranking-এর সামগ্রিক মান ধারাবাহিকভাবে উন্নত করে<sup>[\[19\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-19)</sup>।

## benchmark-এ মূল্যায়ন (BEIR এবং অন্যান্য)

মূল গবেষণাটি TREC DL'19/20 (ওয়েব অনুসন্ধান) এবং BEIR collection-এর একটি উপসেটে (Scifact, ArguAna, TREC‑COVID, FiQA, DBPedia, TREC‑NEWS, Climate‑FEVER) zero-shot মোডে HyDE মূল্যায়ন করে। ফলাফলের একটি অংশ — *2023‑07 তারিখের হিসাবে*:

| পদ্ধতি                     | DL19                   | DL20                   | উৎস                                                                                                              |
|---------------------------|------------------------|------------------------|------------------------------------------------------------------------------------------------------------------|
| BM25                      | 30.1 / 50.6 / 75.0     | 28.6 / 48.0 / 78.6     | <sup>[\[20\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-20)</sup> |
| Contriever (unsup.)       | 24.0 / 44.5 / 74.6     | 24.0 / 42.1 / 75.4     | <sup>[\[21\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-21)</sup> |
| **HyDE** (Contriever+LLM) | **41.8 / 61.3 / 88.0** | **38.2 / 57.9 / 84.4** | <sup>[\[22\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-22)</sup> |
| DPR (ft)                  | 36.5 / 62.2 / 76.9     | 41.8 / 65.3 / 81.4     | <sup>[\[23\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-23)</sup> |
| ANCE (ft)                 | 37.1 / 64.5 / 75.5     | 40.8 / 64.6 / 77.6     | <sup>[\[24\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-24)</sup> |

*TREC DL19/20 (ওয়েব অনুসন্ধান)* — mAP / nDCG@10 / Recall@1k

| পদ্ধতি      | Scifact         | ArguAna         | TREC‑COVID      | FiQA        | DBPedia     | TREC‑NEWS       | Climate‑FEVER   | উৎস                                                                                                              |
|------------|-----------------|-----------------|-----------------|-------------|-------------|-----------------|-----------------|------------------------------------------------------------------------------------------------------------------|
| BM25       | 67.9 / 92.5     | 39.7 / 93.2     | **59.5 / 49.8** | 23.6 / 54.0 | 31.8 / 46.8 | 39.5 / 44.7     | 16.5 / 42.5     | <sup>[\[25\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-25)</sup> |
| Contriever | 64.9 / 92.6     | 37.9 / 90.1     | 27.3 / 17.2     | 24.5 / 56.2 | 29.2 / 45.3 | 34.8 / 42.3     | 15.5 / 44.1     | <sup>[\[26\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-26)</sup> |
| **HyDE**   | **69.1 / 96.4** | **46.6 / 97.9** | 59.3 / 41.4     | 27.3 / 62.1 | 36.8 / 47.2 | **44.0 / 50.9** | **22.3 / 53.0** | <sup>[\[27\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-27)</sup> |

*BEIR (dataset-এর বাছাই)* — nDCG@10 / Recall@100

HyDE বহুভাষিক Mr.TyDi dataset-এ (sw/ko/ja/bn) mContriever-এর তুলনায় MRR@100-ও উন্নত করে<sup>[\[28\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-28)</sup>।

## ব্যবহারিক সুপারিশ

HyDE কখন প্রয়োগ করবেন

- Zero-shot/transfer মোড (কোনো প্রাসঙ্গিক লেবেল নেই; প্রশিক্ষণ corpus থেকে domain «ভিন্নতা»)<sup>[\[29\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-29)</sup>।
- গ্রহণযোগ্য নির্ভুলতায় Recall@k বৃদ্ধি প্রয়োজন — HyDE প্রায়ই ভেক্টর স্থানের প্রাসঙ্গিক অঞ্চলগুলি «উন্মুক্ত» করে<sup>[\[30\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-30)</sup>।

সাধারণ সেটিংস

- **LLM এবং prompt**: «প্রশ্নটির উত্তর দিয়ে একটি অনুচ্ছেদ লিখুন …» নির্দেশনা; মাঝারি stochasticity (যেমন *temperature*≈0.7)<sup>[\[31\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-31)</sup>।
- **হাইপোথেটিক্যাল টেক্সটের সংখ্যা**: ১–৫; embedding-এর গড় করা স্থিতিশীলতা বাড়ায়<sup>[\[32\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-32)</sup>।
- **Embedder**: fine-tuning ছাড়া (m)Contriever; fine-tuned encoder প্রয়োগও সম্ভব (HyDE-এর প্রভাব বজায় থাকে)<sup>[\[33\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-33)</sup>।
- **Embedding normalization**: L2‑norm; dot product cosine-এর সমতুল্য হয়<sup>[\[34\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-34)</sup>।
- **হাইব্রিড retrieval**: পরবর্তী reranking সহ BM25+vector<sup>[\[35\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-35)</sup>।
- **Reranker**: Cross-Encoder (BERT re‑ranker)<sup>[\[36\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-36)</sup> বা ColBERT<sup>[\[37\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-37)</sup>।
- **একীকরণ** বিভিন্ন কৌশলের ফলাফলের: RRF (*k*≈60)<sup>[\[38\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-38)</sup>।

গুণমান/খরচ পর্যবেক্ষণ

- Retrieval: nDCG@k, Recall@k, MRR; end‑to‑end RAG: EM/F1 বা *groundedness* মেট্রিক (RAGAS/TruLens)<sup>[\[39\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-39)[\[40\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-40)</sup>।
- খরচ/latency: LLM generation এবং (যদি থাকে) reranking প্রভাবশালী; হাইপোথেটিক্যাল টেক্সটের সংখ্যা এবং উত্তরের দৈর্ঘ্য দ্বারা অপ্টিমাইজ করা হয়<sup>[\[41\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-41)</sup>।

## সীমাবদ্ধতা এবং উন্মুক্ত প্রশ্ন

- **হাইপোথেটিক্যাল টেক্সটের hallucination**: LLM তথ্যগত ত্রুটি প্রবর্তন করতে পারে; encoder এবং corpus-এর মাধ্যমে «grounding» ঝুঁকি কমায়, কিন্তু সম্পূর্ণ দূর করে না<sup>[\[42\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-42)</sup>।
- **Domain/ভাষার সীমাবদ্ধতা**: অত্যন্ত বিশেষায়িত domain এবং কম-সম্পদ ভাষায় HyDE-এর সুবিধা হ্রাস পায়<sup>[\[43\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-43)</sup>।
- **Latency এবং খরচ**: LLM generation বিলম্ব এবং token খরচ যোগ করে; অনলাইন পরিস্থিতি এবং দীর্ঘ হাইপোথেটিক্যাল টেক্সটের জন্য গুরুত্বপূর্ণ<sup>[\[44\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-44)</sup>।
- **নৈতিকতা এবং পক্ষপাত**: নিরাপদ LLM এবং ফিল্টারিং ব্যবহার করা বাঞ্ছনীয়<sup>[\[45\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-45)</sup>।

## পদ্ধতির তুলনামূলক সারণি

| পদ্ধতি                     | শ্রেণী              | টেক্সট কোথায় তৈরি হয়                  | Encoder/index        | Reranker (২য় ধাপ)            | সাধারণ মেট্রিক (উদাহরণ)                             | খরচ/latency                            | উৎস                                                                                                                                                                                                                   |
|---------------------------|--------------------|--------------------------------------|----------------------|------------------------------|----------------------------------------------------|----------------------------------------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| **HyDE**                  | Query→*hypo‑doc*   | প্রশ্নের পক্ষে (LLM → অনুচ্ছেদ)           | (m)Contriever; ANN   | BERT re‑rank / ColBERT / RRF | DL19 nDCG@10≈61.3; DL20≈57.9; ArguAna nDCG@10≈46.6 | \+ LLM generation; + reranking (ঐচ্ছিক) | <sup>[\[46\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-46)</sup>                                                                                                      |
| BM25                      | লেক্সিক্যাল          | —                                    | Inverted index       | ঐচ্ছিক                        | উপরের সারণি দেখুন                                   | কম (lexical)                           | <sup>[\[47\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-47)</sup>                                                                                                      |
| DPR / ANCE                | ঘন (fine-tuned)    | —                                    | Bi‑encoder; ANN      | ঐচ্ছিক                        | DL19 nDCG@10≈62–65                                 | মাঝারি (LLM ছাড়া)                      | <sup>[\[48\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-48)[\[49\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-49)</sup> |
| doc2query / docTTTTTquery | ডকুমেন্ট expansion   | collection-এর পক্ষে (indexing-এর আগে) | BM25/sparse+expanded | ঐচ্ছিক                        | MS MARCO-তে BM25 উন্নতি                             | উচ্চ অফলাইন generation; দ্রুত অনলাইন      | <sup>[\[50\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-50)[\[51\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-51)</sup> |
| PRF (Rocchio, RLM)        | feedback-ভিত্তিক QE | প্রশ্ন (শীর্ষ ফলাফল অনুযায়ী)             | যেকোনো               | ঐচ্ছিক                        | Recall বৃদ্ধি/drift ঝুঁকি                              | \+ অতিরিক্ত retrieval পাস               | <sup>[\[52\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_note-52)</sup>                                                                                                      |

HyDE এবং সংশ্লিষ্ট পদ্ধতির তুলনা

## আরও দেখুন

- BM25
- ভেক্টর উপস্থাপনা দ্বারা অনুসন্ধান
- RAG
- Pseudo-relevant feedback
- BEIR

## সাহিত্য

- Manning, C. D.; Raghavan, P.; Schütze, H. (2008). *Introduction to Information Retrieval*. Cambridge University Press. ISBN 978‑0521865715.
- Robertson, S.; Zaragoza, H. (2009). *The Probabilistic Relevance Framework: BM25 and Beyond*. Foundations and Trends in IR, 3(4), 333–389. DOI:10.1561/1500000019.

## লিংক

- HyDE রিপোজিটরি: github.com/texttron/hyde।
- ডকুমেন্টেশন: Haystack — HyDE: docs.haystack.deepset.ai।
- ডকুমেন্টেশন: LangChain — HyDE Retriever: docs.langchain.com।

## মন্তব্য

1.  <span id="cite_note-1">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-1) Gao, L.; Ma, X.; Lin, J.; Callan, J. (2023). ‘‘Precise Zero‑Shot Dense Retrieval without Relevance Labels (HyDE)’’. ACL 2023. pp. 1762–1777. DOI:10.18653/v1/2023.acl-long.99. <a href="https://arxiv.org/abs/2212.10496" class="external text" rel="nofollow">arXiv:2212.10496</a></span>
2.  <span id="cite_note-2">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-2) Gao, L. et al. (2023). ACL 2023, §3.2. DOI:10.18653/v1/2023.acl-long.99.</span>
3.  <span id="cite_note-3">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-3) Rocchio, J. (1971). ‘‘Relevance Feedback in Information Retrieval’’. In: Salton, G. (ed.) *The SMART Retrieval System*. Prentice‑Hall, pp. 313–323. ISBN 978‑0138145255.</span>
4.  <span id="cite_note-4">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-4) Lavrenko, V.; Croft, W. B. (2001). ‘‘Relevance‑Based Language Models’’. SIGIR. DOI:10.1145/383952.383972.</span>
5.  <span id="cite_note-5">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-5) Izacard, G. et al. (2021/2022). ‘‘Unsupervised Dense Information Retrieval with Contrastive Learning’’. <a href="https://arxiv.org/abs/2112.09118" class="external text" rel="nofollow">arXiv:2112.09118</a>.</span>
6.  <span id="cite_note-6">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-6) Karpukhin, V. et al. (2020). ‘‘Dense Passage Retrieval for Open‑Domain QA’’. EMNLP. DOI:10.18653/v1/2020.emnlp-main.550.</span>
7.  <span id="cite_note-7">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-7) Thakur, N. et al. (2021). ‘‘BEIR: A Heterogeneous Benchmark for Zero‑shot Evaluation of Information Retrieval Models’’. NeurIPS Datasets Track. <a href="https://arxiv.org/abs/2104.08663" class="external text" rel="nofollow">arXiv:2104.08663</a>.</span>
8.  <span id="cite_note-8">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-8) Gao, L. et al. (2023). DOI:10.18653/v1/2023.acl-long.99.</span>
9.  <span id="cite_note-9">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-9) Milvus Docs. ‘‘Similarity Metrics’’ — при L2‑нормализации векторов внутр. произведение эквивалентно косинусу. URL: <a href="https://milvus.io/docs/v2.2.x/metric.md" class="external free" rel="nofollow">https://milvus.io/docs/v2.2.x/metric.md</a></span>
10. <span id="cite_note-10">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-10) Gao, L.; Ma, X.; Lin, J.; Callan, J. (2023). ‘‘Precise Zero‑Shot Dense Retrieval without Relevance Labels (HyDE)’’. ACL 2023, §3–4. arXiv:2212.10496. DOI:10.18653/v1/2023.acl-long.99.</span>
11. <span id="cite_note-11">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-11) Izacard, G. et al. (2021/2022). ‘‘Unsupervised Dense Information Retrieval with Contrastive Learning (Contriever)’’. arXiv:2112.09118.</span>
12. <span id="cite_note-12">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-12) Gao, L. et al. (2023). Прил. (ablation): влияние числа гипотетических текстов и параметров генерации. arXiv:2212.10496.</span>
13. <span id="cite_note-13">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-13) Gao, L. et al. (2023). DOI:10.18653/v1/2023.acl-long.99.</span>
14. <span id="cite_note-14">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-14) Nogueira, R. et al. (2019). ‘‘Document Expansion by Query Prediction’’ (doc2query). <a href="https://arxiv.org/abs/1904.08375" class="external text" rel="nofollow">arXiv:1904.08375</a>.</span>
15. <span id="cite_note-15">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-15) Nogueira, R.; Lin, J. (2019). ‘‘From doc2query to docTTTTTquery’’ (tech report). <a href="https://cs.uwaterloo.ca/~jimmylin/publications/Nogueira_Lin_2019_docTTTTTquery-v2.pdf" class="external text" rel="nofollow">PDF</a></span>
16. <span id="cite_note-16">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-16) Rocchio, J. (1971); Lavrenko & Croft (2001), см. выше.</span>
17. <span id="cite_note-17">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-17) Nogueira, R.; Cho, K. (2019). ‘‘Passage Re‑ranking with BERT’’. <a href="https://arxiv.org/abs/1901.04085" class="external text" rel="nofollow">arXiv:1901.04085</a>.</span>
18. <span id="cite_note-18">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-18) Khattab, O.; Zaharia, M. (2020). ‘‘ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT’’. SIGIR. DOI:10.1145/3397271.3401075; <a href="https://arxiv.org/abs/2004.12832" class="external text" rel="nofollow">arXiv:2004.12832</a>.</span>
19. <span id="cite_note-19">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-19) Cormack, G. V.; Clarke, C. L. A.; Büttcher, S. (2009). ‘‘Reciprocal Rank Fusion Outperforms Condorcet and Nearly Optimally Combines Rankings’’. SIGIR. DOI:10.1145/1571941.1572114.</span>
20. <span id="cite_note-20">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-20) Gao, L. et al. (2023). Табл. 1. DOI:10.18653/v1/2023.acl-long.99.</span>
21. <span id="cite_note-21">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-21) Izacard, G. et al. (2022); сводные метрики — в Gao et al., 2023, табл. 1. arXiv:2112.09118.</span>
22. <span id="cite_note-22">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-22) Gao, L. et al. (2023). Табл. 1.</span>
23. <span id="cite_note-23">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-23) Karpukhin, V. et al. (2020); сводные — в Gao et al., 2023.</span>
24. <span id="cite_note-24">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-24) Xiong, L. et al. (2021). ICLR. <a href="https://arxiv.org/abs/2007.00808" class="external text" rel="nofollow">arXiv:2007.00808</a>.</span>
25. <span id="cite_note-25">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-25) Thakur, N. et al. (2021); сводные — в Gao et al., 2023, табл. 2. arXiv:2104.08663.</span>
26. <span id="cite_note-26">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-26) Izacard, G. et al. (2022); сводные — в Gao et al., 2023, табл. 2.</span>
27. <span id="cite_note-27">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-27) Gao, L. et al. (2023). Табл. 2.</span>
28. <span id="cite_note-28">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-28) Gao, L. et al. (2023). Табл. 3. DOI:10.18653/v1/2023.acl-long.99.</span>
29. <span id="cite_note-29">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-29) Gao, L. et al. (2023). §4–5.</span>
30. <span id="cite_note-30">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-30) Gao, L. et al. (2023). §4.2–4.3.</span>
31. <span id="cite_note-31">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-31) Gao, L. et al. (2023). §4.1.</span>
32. <span id="cite_note-32">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-32) Haystack Docs. ‘‘Hypothetical Document Embeddings (HyDE)’’ (инженерная справка). <a href="https://docs.haystack.deepset.ai/docs/hypothetical-document-embeddings-hyde" class="external text" rel="nofollow">docs.haystack.deepset.ai</a></span>
33. <span id="cite_note-33">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-33) Gao, L. et al. (2023). Табл. 6.</span>
34. <span id="cite_note-34">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-34) Milvus Docs. ‘‘Similarity Metrics’’.</span>
35. <span id="cite_note-35">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-35) Haystack × Milvus Integration (официальная док.). <a href="https://haystack.deepset.ai/integrations/milvus-document-store" class="external text" rel="nofollow">haystack.deepset.ai</a></span>
36. <span id="cite_note-36">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-36) Nogueira, R.; Cho, K. (2019). arXiv:1901.04085.</span>
37. <span id="cite_note-37">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-37) Khattab, O.; Zaharia, M. (2020). DOI:10.1145/3397271.3401075.</span>
38. <span id="cite_note-38">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-38) Cormack, G. V. et al. (2009). DOI:10.1145/1571941.1572114.</span>
39. <span id="cite_note-39">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-39) Manning, C. D.; Raghavan, P.; Schütze, H. (2008). *Introduction to Information Retrieval*. Cambridge Univ. Press. ISBN 978‑0521865715.</span>
40. <span id="cite_note-40">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-40) Es, S. et al. (2023). ‘‘RAGAS: Automated Evaluation of Retrieval‑Augmented Generation’’. <a href="https://arxiv.org/abs/2309.15217" class="external text" rel="nofollow">arXiv:2309.15217</a>.</span>
41. <span id="cite_note-41">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-41) Gao, L. et al. (2023). §5.</span>
42. <span id="cite_note-42">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-42) Gao, L. et al. (2023). §3.2; §4.1. DOI:10.18653/v1/2023.acl-long.99.</span>
43. <span id="cite_note-43">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-43) Gao, L. et al. (2023). Табл. 3; §4.4.</span>
44. <span id="cite_note-44">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-44) Gao, L. et al. (2023). §4–5.</span>
45. <span id="cite_note-45">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-45) Ouyang, L. et al. (2022). ‘‘Training language models to follow instructions with human feedback (InstructGPT)’’. NeurIPS. <a href="https://arxiv.org/abs/2203.02155" class="external text" rel="nofollow">arXiv:2203.02155</a>.</span>
46. <span id="cite_note-46">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-46) Gao, L. et al. (2023). Табл. 1–2.</span>
47. <span id="cite_note-47">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-47) Robertson, S.; Zaragoza, H. (2009). ‘‘The Probabilistic Relevance Framework: BM25 and Beyond’’. Found. Trends IR. DOI:10.1561/1500000019.</span>
48. <span id="cite_note-48">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-48) Karpukhin, V. et al. (2020). DOI:10.18653/v1/2020.emnlp-main.550.</span>
49. <span id="cite_note-49">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-49) Xiong, L. et al. (2021). <a href="https://arxiv.org/abs/2007.00808" class="external text" rel="nofollow">arXiv:2007.00808</a>.</span>
50. <span id="cite_note-50">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-50) Nogueira, R. et al. (2019). arXiv:1904.08375.</span>
51. <span id="cite_note-51">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-51) Nogueira, R.; Lin, J. (2019). tech report.</span>
52. <span id="cite_note-52">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(BN)#cite_ref-52) Rocchio, J. (1971). SMART; Lavrenko & Croft (2001) SIGIR.</span>
