---
title: "Hypothetical Document Embeddings (HyDE) (UR)"
source: "https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)"
wiki: "systems-analysis.info/int"
article: "Hypothetical_Document_Embeddings_(HyDE)_(UR)"
language: "ur"
categories:
  - "Category:Large language models"
  - "Category:Pages with math errors"
  - "Category:Pages with math render errors"
  - "Category:Prompt engineering"
  - "Category:Urdu"
revision_id: 3126
wiki_created_at: 2026-09-06T23:16:05Z
wiki_modified_at: 2026-09-06T23:16:05Z
downloaded_at: 2026-09-07T22:55:21Z
---

# Hypothetical Document Embeddings (HyDE) (UR)

**Hypothetical Document Expansion (HyDE)** — یہ ویکٹر ریٹریول اور retrieval‑augmented generation (RAG) کو بہتر بنانے کا ایک طریقہ ہے، جس میں بڑا زبانی نمونہ (LLM) اصل سوال کی بنیاد پر ایک «فرضی دستاویز» تیار کرتا ہے؛ پھر اس متن کو encoder کے ذریعے ویکٹر میں تبدیل کیا جاتا ہے، اور حاصل شدہ ویکٹر کے قریب موجود حقیقی دستاویزات میں تلاش کی جاتی ہے۔ یہ طریقہ LLM میں موجود «ربط کے نمونوں» کو استعمال کرنے اور انہیں گھنے embeddings کے ذریعے کارپس سے «جوڑنے» کی اجازت دیتا ہے<sup>[\[1\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-1)</sup>۔

## تعریف اور بنیادی سوچ

HyDE تلاش کے کام کو دو مرحلوں میں تقسیم کرتا ہے:

\(1\) LLM سوال کے جواب میں «متعلقہ جواب کی مثال» (*hypothetical document*) تیار کرتا ہے، اس طرح ربط کی خصوصیات کو ماڈل کرتا ہے؛

\(2\) ایک contrastive encoder (مثلاً Contriever) اس متن کو ایک ایسے ویکٹر میں تبدیل کرتا ہے جس کے ذریعے index سے حقیقی دستاویزات نکالی جاتی ہیں۔ تیار کردہ متن میں حقیقی غلطیاں ہو سکتی ہیں، لیکن encoder کے ذریعے پکڑے جانے والے موضوعاتی اور اصطلاحاتی نمونے اہم ہوتے ہیں<sup>[\[2\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-2)</sup>۔

## تاریخ اور ماخذ

مصنوعی متون سے تلاش کو وسعت دینے کا خیال سوال کی توسیع اور pseudo-relevant feedback (PRF) پر ہونے والے کاموں سے لیا گیا ہے: Rocchio الگورتھم اور ربط کے زبانی نمونے<sup>[\[3\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-3)[\[4\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-4)</sup>۔ گھنے ریٹریول کے لیے contrastively تربیت یافتہ encoders (Contriever)<sup>[\[5\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-5)</sup> اور Dense Passage Retrieval (DPR)<sup>[\[6\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-6)</sup> استعمال کیے گئے۔ benchmark BEIR نے zero‑shot تشخیص کو معیاری بنایا<sup>[\[7\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-7)</sup>۔ اسی پس منظر میں HyDE کو encoder کی دوبارہ تربیت کے بغیر، LLM کے ذریعے نولِ صفر موڈ میں ربط کا علم «لانے» کے ایک طریقے کے طور پر پیش کیا گیا<sup>[\[8\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-8)</sup>۔

## طریقہ اور رسمی بیان

فرض کیجیے دستاویزات کا کارپس $\mathcal{D} = \{ d_{1},\ldots,d_{N}\}$ ہے، اور متن encoder $E:\text{text} \rightarrow {\mathbb{R}}^{n}$ دستاویزات کی ویکٹر نمائندگی $\mathbf{v}_{d} = E(d)$ فراہم کرتا ہے۔ قربت ناپنے کے لیے یا تو cosine similarity یا scalar product استعمال ہوتا ہے؛ ایک اہم نکتہ یہ ہے کہ \*\*scalar product صرف اسی وقت cosine similarity کے برابر ہوتا ہے جب دونوں ویکٹروں کا L2‑norm اکائی ہو\*\* ($\|\mathbf{u}\| = \|\mathbf{v}\| = 1$)<sup>[\[9\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-9)</sup>۔

HyDE سوال کی نمائندگی کو LLM کے ذریعے تیار کردہ «فرضی دستاویز» سے نئے سرے سے متعین کرتا ہے۔ رسمی طور پر:

$$
\begin{matrix}
 & \text{(1) Генерация гипотетического текста:} & & {\overset{\sim}{d}\; = \; G\!\left( q;\,{inst} \right),} \\
 & \text{(2) Эмбеддинг гипотетического текста:} & & {\mathbf{v}_{h}\; = \; E(\overset{\sim}{d}),} \\
 & \text{(3) Поиск ближайших соседей:} & & {\mathcal{R}_{k}(q)\; = \;{TopK}_{\, d \in \mathcal{D}}\; S\!\left( \mathbf{v}_{h},\mathbf{v}_{d} \right),}
\end{matrix}
$$

جہاں $G$ — ہدایت $inst$ کے ساتھ LLM ہے (مثلاً: «ایک پیراگراف لکھیں جو اس سوال کا جواب دے …»)، $S$ — مشابہت کا پیمانہ (cosine یا normalization کے ساتھ IP) ہے، اور $\mathcal{R}_{k}(q)$ — زیادہ سے زیادہ مشابہت والے $k$ دستاویزات کا مجموعہ ہے<sup>[\[10\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-10)[\[11\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-11)</sup>۔

انجینیئری عمل میں اکثر \*\*متعدد\*\* فرضی متون تیار کیے جاتے ہیں اور ان کی نمائندگیوں کو یکجا کیا جاتا ہے، جس سے استحکام بڑھتا ہے:

$$
{\overset{\sim}{d}}^{(j)} = G\!\left( q;\,{inst},\xi_{j} \right),\quad\mathbf{v}_{h}\; = \;\frac{1}{m}\sum\limits_{j = 1}^{m}E\!\left( {\overset{\sim}{d}}^{(j)} \right),
$$

جہاں $\xi_{j}$ — stochastic decoding پیرامیٹرز ہیں (مثلاً temperature/top‑p)۔ اس طرح کی ensembling اعتدال پسند latency اضافے کے ساتھ Recall کو بہتر بناتی ہے<sup>[\[12\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-12)</sup>۔

### HyDE کا بنیادی پائپ لائن

    # 1) prompt(query) -> hypothetical_doc
    # 2) embed(hypothetical_doc) -> v_h
    # 3) retrieve(index, v_h, k) -> candidates
    # 4) (optional) rerank(query, candidates) -> topN
    # 5) (для RAG) stuff / map-reduce / refine на topN

### دیگر طریقوں سے تعلق (QE، doc2query، PRF)

- **QE (سوال کی توسیع)** سوال میں اصطلاحات شامل کرتا ہے؛ HyDE اس کے بجائے ایک مکمل «شبہ دستاویز» تیار کرتا ہے، جو گھنے encoders کے ساتھ بہتر ہم آہنگ ہے<sup>[\[13\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-13)</sup>۔
- **doc2query / docTTTTTquery** indexing سے پہلے مصنوعی سوالات سے **دستاویزات** کو وسعت دیتا ہے<sup>[\[14\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-14)[\[15\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-15)</sup>؛ HyDE **سوال** کو بروقت وسعت دیتا ہے، re-indexing کی ضرورت نہیں۔
- **PRF** (Rocchio، Relevance LM) ٹاپ نتائج کی بنیاد پر سوال کے ویکٹر کو اپ ڈیٹ کرتا ہے؛ HyDE «ربط کا نمونہ» براہِ راست LLM سے نکالتا ہے اور پھر اسے کارپس پر ریٹریول کے ذریعے «جوڑتا» ہے<sup>[\[16\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-16)</sup>۔

## RAG اور دوبارہ درجہ بندی میں انضمام

RAG میں HyDE ریٹریول کے پہلے مرحلے کے طور پر استعمال ہوتا ہے: فرضی دستاویز ← embedding ← k امیدوار۔ اس کے بعد دوبارہ درجہ بندی کی جاتی ہے: BERT کلاس کے cross‑encoders<sup>[\[17\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-17)</sup> یا ColBERT کا late interaction<sup>[\[18\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-18)</sup>۔ فہرستوں کو ضم کرنے کے لیے (مثلاً BM25+vector کا ہائبرڈ) عام طور پر RRF (*reciprocal rank fusion*) استعمال ہوتا ہے: $\operatorname{RRF}(d) = \sum\limits_{r \in \mathcal{R}}\frac{1}{k + \operatorname{rank}_{r}(d)},\qquad k \approx 60.$ RRF طریقہ مجتمع درجہ بندیوں کے مجموعی معیار کو مستقل طور پر بہتر بناتا ہے<sup>[\[19\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-19)</sup>۔

## benchmarks پر تشخیص (BEIR وغیرہ)

اصل تحقیق TREC DL'19/20 (ویب تلاش) اور BEIR مجموعوں کے ذیلی حصے (Scifact، ArguAna، TREC‑COVID، FiQA، DBPedia، TREC‑NEWS، Climate‑FEVER) پر نولِ صفر موڈ میں HyDE کا جائزہ لیتی ہے۔ نتائج کا ایک حصہ — *جولائی 2023 کی صورتِ حال کے مطابق*:

| طریقہ                     | DL19                   | DL20                   | ماخذ                                                                                                             |
|---------------------------|------------------------|------------------------|------------------------------------------------------------------------------------------------------------------|
| BM25                      | 30.1 / 50.6 / 75.0     | 28.6 / 48.0 / 78.6     | <sup>[\[20\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-20)</sup> |
| Contriever (unsup.)       | 24.0 / 44.5 / 74.6     | 24.0 / 42.1 / 75.4     | <sup>[\[21\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-21)</sup> |
| **HyDE** (Contriever+LLM) | **41.8 / 61.3 / 88.0** | **38.2 / 57.9 / 84.4** | <sup>[\[22\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-22)</sup> |
| DPR (ft)                  | 36.5 / 62.2 / 76.9     | 41.8 / 65.3 / 81.4     | <sup>[\[23\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-23)</sup> |
| ANCE (ft)                 | 37.1 / 64.5 / 75.5     | 40.8 / 64.6 / 77.6     | <sup>[\[24\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-24)</sup> |

*TREC DL19/20 (ویب تلاش)* — mAP / nDCG@10 / Recall@1k

| طریقہ      | Scifact         | ArguAna         | TREC‑COVID      | FiQA        | DBPedia     | TREC‑NEWS       | Climate‑FEVER   | ماخذ                                                                                                             |
|------------|-----------------|-----------------|-----------------|-------------|-------------|-----------------|-----------------|------------------------------------------------------------------------------------------------------------------|
| BM25       | 67.9 / 92.5     | 39.7 / 93.2     | **59.5 / 49.8** | 23.6 / 54.0 | 31.8 / 46.8 | 39.5 / 44.7     | 16.5 / 42.5     | <sup>[\[25\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-25)</sup> |
| Contriever | 64.9 / 92.6     | 37.9 / 90.1     | 27.3 / 17.2     | 24.5 / 56.2 | 29.2 / 45.3 | 34.8 / 42.3     | 15.5 / 44.1     | <sup>[\[26\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-26)</sup> |
| **HyDE**   | **69.1 / 96.4** | **46.6 / 97.9** | 59.3 / 41.4     | 27.3 / 62.1 | 36.8 / 47.2 | **44.0 / 50.9** | **22.3 / 53.0** | <sup>[\[27\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-27)</sup> |

*BEIR (منتخب مجموعے)* — nDCG@10 / Recall@100

HyDE کثیر لسانی مجموعوں Mr.TyDi (sw/ko/ja/bn) پر mContriever کے مقابلے میں MRR@100 کو بھی بہتر بناتا ہے<sup>[\[28\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-28)</sup>۔

## عملی سفارشات

HyDE کب استعمال کریں

- نولِ صفر / منتقلی موڈ (کوئی متعلقہ لیبل نہ ہو؛ تربیتی کارپس سے ڈومین کا «فرق»)<sup>[\[29\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-29)</sup>۔
- قابلِ قبول درستگی کے ساتھ Recall@k میں اضافہ درکار ہو — HyDE اکثر ویکٹر اسپیس کے متعلقہ علاقوں کو «کھولتا» ہے<sup>[\[30\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-30)</sup>۔

عام ترتیبات

- **LLM اور prompt**: ہدایت «ایک پیراگراف لکھیں جو اس سوال کا جواب دے …»؛ اعتدال پسند stochasticity (مثلاً *temperature*≈0.7)<sup>[\[31\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-31)</sup>۔
- **فرضی متون کی تعداد**: 1–5؛ embeddings کا اوسط نکالنا استحکام بڑھاتا ہے<sup>[\[32\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-32)</sup>۔
- **Embedder**: دوبارہ تربیت کے بغیر (m)Contriever؛ fine-tuned encoders کا استعمال بھی ممکن ہے (HyDE کا اثر برقرار رہتا ہے)<sup>[\[33\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-33)</sup>۔
- **Embeddings کا معمول بنانا**: L2‑norm؛ اندرونی ضرب cosine کے مساوی ہے<sup>[\[34\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-34)</sup>۔
- **ہائبرڈ ریٹریول**: BM25+vector اور بعد میں دوبارہ درجہ بندی<sup>[\[35\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-35)</sup>۔
- **Re-ranker**: Cross-Encoder (BERT re‑ranker)<sup>[\[36\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-36)</sup> یا ColBERT<sup>[\[37\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-37)</sup>۔
- مختلف حکمتِ عملیوں کے نتائج کا **ضم**: RRF (*k*≈60)<sup>[\[38\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-38)</sup>۔

معیار/لاگت کی نگرانی

- ریٹریول: nDCG@k، Recall@k، MRR؛ end‑to‑end RAG: EM/F1 یا *groundedness* کے پیمانے (RAGAS/TruLens)<sup>[\[39\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-39)[\[40\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-40)</sup>۔
- لاگت/latency: LLM کی تیاری اور (اگر ہو تو) دوبارہ درجہ بندی غالب ہے؛ «فرضی» متون کی تعداد اور جواب کی لمبائی سے بہتر بنایا جاتا ہے<sup>[\[41\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-41)</sup>۔

## حدود اور کھلے سوالات

- فرضی متن میں **ہذیان**: LLM حقیقی غلطیاں داخل کر سکتا ہے؛ encoder اور کارپس کے ذریعے «جوڑنا» خطرہ کم کرتا ہے لیکن مکمل طور پر ختم نہیں کرتا<sup>[\[42\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-42)</sup>۔
- **ڈومین/زبان کی حدود**: انتہائی مخصوص ڈومینز اور کم وسائل والی زبانوں میں HyDE کا فائدہ کم ہو جاتا ہے<sup>[\[43\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-43)</sup>۔
- **Latency اور لاگت**: LLM کی تیاری تاخیر اور token لاگت بڑھاتی ہے؛ آن لائن منظرناموں اور لمبے «فرضی» متون کے لیے یہ اہم ہے<sup>[\[44\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-44)</sup>۔
- **اخلاقیات اور جھکاؤ**: محفوظ LLMs اور فلٹرنگ کا استعمال ترجیحی ہے<sup>[\[45\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-45)</sup>۔

## طریقوں کا تقابلی جدول

| طریقہ                     | طبقہ             | متن کہاں تیار ہوتا ہے             | Encoder/Index        | Re-ranker (دوسرا مرحلہ)      | عام پیمانے (مثال)                                  | لاگت/Latency                                  | ماخذ                                                                                                                                                                                                                  |
|---------------------------|------------------|-----------------------------------|----------------------|------------------------------|----------------------------------------------------|-----------------------------------------------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| **HyDE**                  | Query→*hypo‑doc* | سوال کی جانب (LLM → پیراگراف)     | (m)Contriever; ANN   | BERT re‑rank / ColBERT / RRF | DL19 nDCG@10≈61.3; DL20≈57.9; ArguAna nDCG@10≈46.6 | \+ LLM کی تیاری؛ + دوبارہ درجہ بندی (اختیاری) | <sup>[\[46\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-46)</sup>                                                                                                      |
| BM25                      | لغوی             | —                                 | الٹا Index           | اختیاری                      | مذکورہ جدول دیکھیں                                 | کم (lexical)                                  | <sup>[\[47\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-47)</sup>                                                                                                      |
| DPR / ANCE                | گھنا (ft)        | —                                 | Bi‑encoder; ANN      | اختیاری                      | DL19 nDCG@10≈62–65                                 | درمیانی (بغیر LLM)                            | <sup>[\[48\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-48)[\[49\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-49)</sup> |
| doc2query / docTTTTTquery | دستاویز کی توسیع | مجموعے کی جانب (indexing سے پہلے) | BM25/sparse+expanded | اختیاری                      | MS MARCO پر BM25 میں بہتری                         | آف لائن تیاری زیادہ؛ آن لائن تیز              | <sup>[\[50\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-50)[\[51\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-51)</sup> |
| PRF (Rocchio, RLM)        | feedback پر QE   | سوال (ٹاپ نتائج کی بنیاد پر)      | کوئی بھی             | اختیاری                      | Recall میں اضافہ/بہاؤ کے خطرات                     | \+ ریٹریول کا اضافی دور                       | <sup>[\[52\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_note-52)</sup>                                                                                                      |

HyDE اور متعلقہ طریقوں کا موازنہ

## یہ بھی دیکھیں

- BM25
- ویکٹر نمائندگی پر مبنی تلاش،
- RAG
- Pseudo-relevant feedback
- BEIR

## کتابیات

- Manning, C. D.; Raghavan, P.; Schütze, H. (2008). *Introduction to Information Retrieval*. Cambridge University Press. ISBN 978‑0521865715.
- Robertson, S.; Zaragoza, H. (2009). *The Probabilistic Relevance Framework: BM25 and Beyond*. Foundations and Trends in IR, 3(4), 333–389. DOI:10.1561/1500000019.

## حوالہ جات

- HyDE ریپوزیٹری: github.com/texttron/hyde۔
- دستاویز: Haystack — HyDE: docs.haystack.deepset.ai۔
- دستاویز: LangChain — HyDE Retriever: docs.langchain.com۔

## نوٹس

1.  <span id="cite_note-1">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-1) Gao, L.; Ma, X.; Lin, J.; Callan, J. (2023). ‘‘Precise Zero‑Shot Dense Retrieval without Relevance Labels (HyDE)’’. ACL 2023. pp. 1762–1777. DOI:10.18653/v1/2023.acl-long.99. <a href="https://arxiv.org/abs/2212.10496" class="external text" rel="nofollow">arXiv:2212.10496</a></span>
2.  <span id="cite_note-2">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-2) Gao, L. et al. (2023). ACL 2023, §3.2. DOI:10.18653/v1/2023.acl-long.99.</span>
3.  <span id="cite_note-3">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-3) Rocchio, J. (1971). ‘‘Relevance Feedback in Information Retrieval’’. In: Salton, G. (ed.) *The SMART Retrieval System*. Prentice‑Hall, pp. 313–323. ISBN 978‑0138145255.</span>
4.  <span id="cite_note-4">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-4) Lavrenko, V.; Croft, W. B. (2001). ‘‘Relevance‑Based Language Models’’. SIGIR. DOI:10.1145/383952.383972.</span>
5.  <span id="cite_note-5">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-5) Izacard, G. et al. (2021/2022). ‘‘Unsupervised Dense Information Retrieval with Contrastive Learning’’. <a href="https://arxiv.org/abs/2112.09118" class="external text" rel="nofollow">arXiv:2112.09118</a>.</span>
6.  <span id="cite_note-6">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-6) Karpukhin, V. et al. (2020). ‘‘Dense Passage Retrieval for Open‑Domain QA’’. EMNLP. DOI:10.18653/v1/2020.emnlp-main.550.</span>
7.  <span id="cite_note-7">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-7) Thakur, N. et al. (2021). ‘‘BEIR: A Heterogeneous Benchmark for Zero‑shot Evaluation of Information Retrieval Models’’. NeurIPS Datasets Track. <a href="https://arxiv.org/abs/2104.08663" class="external text" rel="nofollow">arXiv:2104.08663</a>.</span>
8.  <span id="cite_note-8">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-8) Gao, L. et al. (2023). DOI:10.18653/v1/2023.acl-long.99.</span>
9.  <span id="cite_note-9">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-9) Milvus Docs. ‘‘Similarity Metrics’’ — при L2‑нормализации векторов внутр. произведение эквивалентно косинусу. URL: <a href="https://milvus.io/docs/v2.2.x/metric.md" class="external free" rel="nofollow">https://milvus.io/docs/v2.2.x/metric.md</a></span>
10. <span id="cite_note-10">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-10) Gao, L.; Ma, X.; Lin, J.; Callan, J. (2023). ‘‘Precise Zero‑Shot Dense Retrieval without Relevance Labels (HyDE)’’. ACL 2023, §3–4. arXiv:2212.10496. DOI:10.18653/v1/2023.acl-long.99.</span>
11. <span id="cite_note-11">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-11) Izacard, G. et al. (2021/2022). ‘‘Unsupervised Dense Information Retrieval with Contrastive Learning (Contriever)’’. arXiv:2112.09118.</span>
12. <span id="cite_note-12">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-12) Gao, L. et al. (2023). Прил. (ablation): влияние числа гипотетических текстов и параметров генерации. arXiv:2212.10496.</span>
13. <span id="cite_note-13">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-13) Gao, L. et al. (2023). DOI:10.18653/v1/2023.acl-long.99.</span>
14. <span id="cite_note-14">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-14) Nogueira, R. et al. (2019). ‘‘Document Expansion by Query Prediction’’ (doc2query). <a href="https://arxiv.org/abs/1904.08375" class="external text" rel="nofollow">arXiv:1904.08375</a>.</span>
15. <span id="cite_note-15">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-15) Nogueira, R.; Lin, J. (2019). ‘‘From doc2query to docTTTTTquery’’ (tech report). <a href="https://cs.uwaterloo.ca/~jimmylin/publications/Nogueira_Lin_2019_docTTTTTquery-v2.pdf" class="external text" rel="nofollow">PDF</a></span>
16. <span id="cite_note-16">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-16) Rocchio, J. (1971); Lavrenko & Croft (2001), см. выше.</span>
17. <span id="cite_note-17">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-17) Nogueira, R.; Cho, K. (2019). ‘‘Passage Re‑ranking with BERT’’. <a href="https://arxiv.org/abs/1901.04085" class="external text" rel="nofollow">arXiv:1901.04085</a>.</span>
18. <span id="cite_note-18">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-18) Khattab, O.; Zaharia, M. (2020). ‘‘ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT’’. SIGIR. DOI:10.1145/3397271.3401075; <a href="https://arxiv.org/abs/2004.12832" class="external text" rel="nofollow">arXiv:2004.12832</a>.</span>
19. <span id="cite_note-19">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-19) Cormack, G. V.; Clarke, C. L. A.; Büttcher, S. (2009). ‘‘Reciprocal Rank Fusion Outperforms Condorcet and Nearly Optimally Combines Rankings’’. SIGIR. DOI:10.1145/1571941.1572114.</span>
20. <span id="cite_note-20">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-20) Gao, L. et al. (2023). Табл. 1. DOI:10.18653/v1/2023.acl-long.99.</span>
21. <span id="cite_note-21">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-21) Izacard, G. et al. (2022); сводные метрики — в Gao et al., 2023, табл. 1. arXiv:2112.09118.</span>
22. <span id="cite_note-22">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-22) Gao, L. et al. (2023). Табл. 1.</span>
23. <span id="cite_note-23">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-23) Karpukhin, V. et al. (2020); сводные — в Gao et al., 2023.</span>
24. <span id="cite_note-24">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-24) Xiong, L. et al. (2021). ICLR. <a href="https://arxiv.org/abs/2007.00808" class="external text" rel="nofollow">arXiv:2007.00808</a>.</span>
25. <span id="cite_note-25">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-25) Thakur, N. et al. (2021); сводные — в Gao et al., 2023, табл. 2. arXiv:2104.08663.</span>
26. <span id="cite_note-26">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-26) Izacard, G. et al. (2022); сводные — в Gao et al., 2023, табл. 2.</span>
27. <span id="cite_note-27">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-27) Gao, L. et al. (2023). Табл. 2.</span>
28. <span id="cite_note-28">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-28) Gao, L. et al. (2023). Табл. 3. DOI:10.18653/v1/2023.acl-long.99.</span>
29. <span id="cite_note-29">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-29) Gao, L. et al. (2023). §4–5.</span>
30. <span id="cite_note-30">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-30) Gao, L. et al. (2023). §4.2–4.3.</span>
31. <span id="cite_note-31">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-31) Gao, L. et al. (2023). §4.1.</span>
32. <span id="cite_note-32">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-32) Haystack Docs. ‘‘Hypothetical Document Embeddings (HyDE)’’ (инженерная справка). <a href="https://docs.haystack.deepset.ai/docs/hypothetical-document-embeddings-hyde" class="external text" rel="nofollow">docs.haystack.deepset.ai</a></span>
33. <span id="cite_note-33">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-33) Gao, L. et al. (2023). Табл. 6.</span>
34. <span id="cite_note-34">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-34) Milvus Docs. ‘‘Similarity Metrics’’.</span>
35. <span id="cite_note-35">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-35) Haystack × Milvus Integration (официальная док.). <a href="https://haystack.deepset.ai/integrations/milvus-document-store" class="external text" rel="nofollow">haystack.deepset.ai</a></span>
36. <span id="cite_note-36">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-36) Nogueira, R.; Cho, K. (2019). arXiv:1901.04085.</span>
37. <span id="cite_note-37">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-37) Khattab, O.; Zaharia, M. (2020). DOI:10.1145/3397271.3401075.</span>
38. <span id="cite_note-38">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-38) Cormack, G. V. et al. (2009). DOI:10.1145/1571941.1572114.</span>
39. <span id="cite_note-39">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-39) Manning, C. D.; Raghavan, P.; Schütze, H. (2008). *Introduction to Information Retrieval*. Cambridge Univ. Press. ISBN 978‑0521865715.</span>
40. <span id="cite_note-40">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-40) Es, S. et al. (2023). ‘‘RAGAS: Automated Evaluation of Retrieval‑Augmented Generation’’. <a href="https://arxiv.org/abs/2309.15217" class="external text" rel="nofollow">arXiv:2309.15217</a>.</span>
41. <span id="cite_note-41">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-41) Gao, L. et al. (2023). §5.</span>
42. <span id="cite_note-42">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-42) Gao, L. et al. (2023). §3.2; §4.1. DOI:10.18653/v1/2023.acl-long.99.</span>
43. <span id="cite_note-43">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-43) Gao, L. et al. (2023). Табл. 3; §4.4.</span>
44. <span id="cite_note-44">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-44) Gao, L. et al. (2023). §4–5.</span>
45. <span id="cite_note-45">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-45) Ouyang, L. et al. (2022). ‘‘Training language models to follow instructions with human feedback (InstructGPT)’’. NeurIPS. <a href="https://arxiv.org/abs/2203.02155" class="external text" rel="nofollow">arXiv:2203.02155</a>.</span>
46. <span id="cite_note-46">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-46) Gao, L. et al. (2023). Табл. 1–2.</span>
47. <span id="cite_note-47">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-47) Robertson, S.; Zaragoza, H. (2009). ‘‘The Probabilistic Relevance Framework: BM25 and Beyond’’. Found. Trends IR. DOI:10.1561/1500000019.</span>
48. <span id="cite_note-48">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-48) Karpukhin, V. et al. (2020). DOI:10.18653/v1/2020.emnlp-main.550.</span>
49. <span id="cite_note-49">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-49) Xiong, L. et al. (2021). <a href="https://arxiv.org/abs/2007.00808" class="external text" rel="nofollow">arXiv:2007.00808</a>.</span>
50. <span id="cite_note-50">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-50) Nogueira, R. et al. (2019). arXiv:1904.08375.</span>
51. <span id="cite_note-51">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-51) Nogueira, R.; Lin, J. (2019). tech report.</span>
52. <span id="cite_note-52">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(UR)#cite_ref-52) Rocchio, J. (1971). SMART; Lavrenko & Croft (2001) SIGIR.</span>
