---
title: "Hypothetical Document Embeddings (HyDE) (FA)"
source: "https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)"
wiki: "systems-analysis.info/int"
article: "Hypothetical_Document_Embeddings_(HyDE)_(FA)"
language: "fa"
categories:
  - "Category:Large language models"
  - "Category:Persian"
  - "Category:Prompt engineering"
revision_id: 3111
wiki_created_at: 2026-09-06T23:15:51Z
wiki_modified_at: 2026-09-06T23:15:51Z
downloaded_at: 2026-09-07T22:55:13Z
---

# Hypothetical Document Embeddings (HyDE) (FA)

**Hypothetical Document Expansion (HyDE)** — روشی برای بهبود بازیابی برداری و retrieval‑augmented generation (RAG) است که در آن یک مدل زبانی بزرگ (LLM) بر اساس پرسش اولیه یک «سند فرضی» تولید می‌کند؛ سپس این متن توسط یک encoder به بردار تبدیل می‌شود و جستجو میان اسناد واقعی بر اساس نزدیکی به بردار به‌دست‌آمده انجام می‌گیرد. این رویکرد امکان بهره‌گیری از «الگوهای ارتباط» رمزگذاری‌شده توسط LLM را فراهم می‌کند و آن‌ها را از طریق embedding‌های متراکم بر روی مجموعه اسناد «متکی» می‌سازد<sup>[\[1\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-1)</sup>.

## تعریف و شهود

HyDE وظیفه جستجو را به دو مرحله تجزیه می‌کند:

\(1\) LLM یک «نمونه پاسخ مرتبط» (*hypothetical document*) برای پرسش تولید می‌کند و بدین ترتیب ویژگی‌های ارتباط را مدل‌سازی می‌نماید;

\(2\) یک encoder تقابلی (مثلاً Contriever) این متن را به برداری تبدیل می‌کند که اسناد واقعی از فهرست بر اساس آن بازیابی می‌شوند. متن تولیدشده ممکن است حاوی خطاهای واقعی باشد، اما آنچه اهمیت دارد الگوهای موضوعی و اصطلاحی است که توسط encoder دریافت می‌شوند<sup>[\[2\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-2)</sup>.

## تاریخچه و منابع

ایده گسترش جستجو با متن‌های مصنوعی ریشه در پژوهش‌های مربوط به گسترش پرسش و بازخورد شبه‌مرتبط (PRF) دارد: الگوریتم Rocchio و مدل‌های زبانی ارتباط<sup>[\[3\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-3)[\[4\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-4)</sup>. برای بازیابی متراکم از encoder‌های آموزش‌دیده با روش تقابلی (Contriever)<sup>[\[5\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-5)</sup> و Dense Passage Retrieval (DPR)<sup>[\[6\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-6)</sup> استفاده شد. بنچمارک BEIR ارزیابی zero‑shot را استانداردسازی کرد<sup>[\[7\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-7)</sup>. در این زمینه، HyDE به عنوان روشی برای «وارد کردن» دانش ارتباط از طریق LLM در حالت نقطه صفر، بدون نیاز به fine-tuning encoder، پیشنهاد شد<sup>[\[8\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-8)</sup>.

## روش و صورت‌بندی

فرض کنید مجموعه اسناد $\mathcal{D} = \{ d_{1},\ldots,d_{N}\}$ و encoder متن $E:\text{text} \rightarrow {\mathbb{R}}^{n}$ بازنمایی‌های برداری اسناد $\mathbf{v}_{d} = E(d)$ را تعریف می‌کند. برای اندازه‌گیری نزدیکی از شباهت کسینوسی یا حاصل‌ضرب داخلی استفاده می‌شود؛ نکته مهم: \*\*حاصل‌ضرب داخلی تنها زمانی با شباهت کسینوسی برابر است که هر دو بردار دارای نُرم L2 واحد باشند\*\* ($\|\mathbf{u}\| = \|\mathbf{v}\| = 1$)<sup>[\[9\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-9)</sup>.

HyDE بازنمایی پرسش را از طریق «سند فرضی» تولیدشده توسط LLM بازتعریف می‌کند. به صورت رسمی:

$$
\begin{matrix}
 & \text{(1) Генерация гипотетического текста:} & & {\overset{\sim}{d}\; = \; G\!\left( q;\,{inst} \right),} \\
 & \text{(2) Эмбеддинг гипотетического текста:} & & {\mathbf{v}_{h}\; = \; E(\overset{\sim}{d}),} \\
 & \text{(3) Поиск ближайших соседей:} & & {\mathcal{R}_{k}(q)\; = \;{TopK}_{\, d \in \mathcal{D}}\; S\!\left( \mathbf{v}_{h},\mathbf{v}_{d} \right),}
\end{matrix}
$$

که در آن $G$ — LLM با دستورالعمل $inst$ (برای مثال: «یک پاراگراف بنویس که به سوال … پاسخ دهد»)، $S$ — معیار شباهت (کسینوس یا IP با نرمال‌سازی)، و $\mathcal{R}_{k}(q)$ — مجموعه‌ای از $k$ سند با بیشترین شباهت است<sup>[\[10\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-10)[\[11\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-11)</sup>.

در عمل مهندسی، اغلب \*\*چندین\*\* متن فرضی تولید می‌شود و بازنمایی‌های آن‌ها تجمیع می‌گردد که پایداری را افزایش می‌دهد:

$$
{\overset{\sim}{d}}^{(j)} = G\!\left( q;\,{inst},\xi_{j} \right),\quad\mathbf{v}_{h}\; = \;\frac{1}{m}\sum\limits_{j = 1}^{m}E\!\left( {\overset{\sim}{d}}^{(j)} \right),
$$

که در آن $\xi_{j}$ — پارامترهای تصادفی رمزگشایی (مثلاً temperature/top‑p) هستند. این ensemble‌سازی Recall را با افزایش متوسط تأخیر بهبود می‌بخشد<sup>[\[12\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-12)</sup>.

### پایپ‌لاین پایه HyDE

    # 1) prompt(query) -> hypothetical_doc
    # 2) embed(hypothetical_doc) -> v_h
    # 3) retrieve(index, v_h, k) -> candidates
    # 4) (optional) rerank(query, candidates) -> topN
    # 5) (для RAG) stuff / map-reduce / refine на topN

### ارتباط با سایر روش‌ها (QE، doc2query، PRF)

- **QE (گسترش پرسش)** اصطلاحات را به پرسش اضافه می‌کند؛ HyDE به جای آن یک «شبه‌سند» کامل تولید می‌کند که با encoder‌های متراکم هم‌راستایی بهتری دارد<sup>[\[13\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-13)</sup>.
- **doc2query / docTTTTTquery** **اسناد** را با پرسش‌های مصنوعی قبل از ایندکس‌گذاری گسترش می‌دهند<sup>[\[14\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-14)[\[15\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-15)</sup>؛ HyDE **پرسش** را به صورت آنی گسترش می‌دهد و نیازی به ایندکس‌گذاری مجدد ندارد.
- **PRF** (Rocchio، Relevance LM) بردار پرسش را بر اساس نتایج برتر به‌روزرسانی می‌کند؛ HyDE «الگوی ارتباط» را مستقیماً از LLM استخراج کرده و سپس آن را از طریق بازیابی از روی مجموعه اسناد «متکی» می‌سازد<sup>[\[16\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-16)</sup>.

## یکپارچه‌سازی در RAG و بازرتبه‌بندی

در RAG، HyDE به عنوان مرحله اول بازیابی به کار می‌رود: سند فرضی → embedding → k کاندیدا. سپس بازرتبه‌بندی انجام می‌شود: cross-encoder‌های کلاس BERT<sup>[\[17\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-17)</sup> یا تعامل دیرهنگام ColBERT<sup>[\[18\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-18)</sup>. برای ادغام فهرست‌ها (مثلاً ترکیب BM25+vector) معمولاً از RRF (*reciprocal rank fusion*) استفاده می‌شود: $\operatorname{RRF}(d) = \sum\limits_{r \in \mathcal{R}}\frac{1}{k + \operatorname{rank}_{r}(d)},\qquad k \approx 60.$ روش RRF به طور پایدار کیفیت کلی رتبه‌بندی‌های ادغام‌شده را بهبود می‌بخشد<sup>[\[19\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-19)</sup>.

## ارزیابی روی بنچمارک‌ها (BEIR و غیره)

مقاله اصلی، HyDE را در حالت نقطه صفر روی TREC DL'19/20 (جستجوی وب) و زیرمجموعه‌ای از مجموعه‌های BEIR (Scifact، ArguAna، TREC‑COVID، FiQA، DBPedia، TREC‑NEWS، Climate‑FEVER) ارزیابی می‌کند. بخشی از نتایج — *به‌روز تا 2023‑07*:

| روش                       | DL19                   | DL20                   | منبع                                                                                                             |
|---------------------------|------------------------|------------------------|------------------------------------------------------------------------------------------------------------------|
| BM25                      | 30.1 / 50.6 / 75.0     | 28.6 / 48.0 / 78.6     | <sup>[\[20\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-20)</sup> |
| Contriever (unsup.)       | 24.0 / 44.5 / 74.6     | 24.0 / 42.1 / 75.4     | <sup>[\[21\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-21)</sup> |
| **HyDE** (Contriever+LLM) | **41.8 / 61.3 / 88.0** | **38.2 / 57.9 / 84.4** | <sup>[\[22\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-22)</sup> |
| DPR (ft)                  | 36.5 / 62.2 / 76.9     | 41.8 / 65.3 / 81.4     | <sup>[\[23\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-23)</sup> |
| ANCE (ft)                 | 37.1 / 64.5 / 75.5     | 40.8 / 64.6 / 77.6     | <sup>[\[24\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-24)</sup> |

*TREC DL19/20 (جستجوی وب)* — mAP / nDCG@10 / Recall@1k

| روش        | Scifact         | ArguAna         | TREC‑COVID      | FiQA        | DBPedia     | TREC‑NEWS       | Climate‑FEVER   | منبع                                                                                                             |
|------------|-----------------|-----------------|-----------------|-------------|-------------|-----------------|-----------------|------------------------------------------------------------------------------------------------------------------|
| BM25       | 67.9 / 92.5     | 39.7 / 93.2     | **59.5 / 49.8** | 23.6 / 54.0 | 31.8 / 46.8 | 39.5 / 44.7     | 16.5 / 42.5     | <sup>[\[25\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-25)</sup> |
| Contriever | 64.9 / 92.6     | 37.9 / 90.1     | 27.3 / 17.2     | 24.5 / 56.2 | 29.2 / 45.3 | 34.8 / 42.3     | 15.5 / 44.1     | <sup>[\[26\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-26)</sup> |
| **HyDE**   | **69.1 / 96.4** | **46.6 / 97.9** | 59.3 / 41.4     | 27.3 / 62.1 | 36.8 / 47.2 | **44.0 / 50.9** | **22.3 / 53.0** | <sup>[\[27\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-27)</sup> |

*BEIR (گزینشی از مجموعه‌ها)* — nDCG@10 / Recall@100

HyDE همچنین MRR@100 را روی مجموعه‌های چندزبانه Mr.TyDi (sw/ko/ja/bn) نسبت به mContriever بهبود می‌بخشد<sup>[\[28\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-28)</sup>.

## توصیه‌های عملی

چه زمانی HyDE را به‌کار ببریم

- حالت‌های نقطه صفر/انتقال‌پذیر (بدون برچسب‌های مرتبط؛ «ناهمشکلی» دامنه با مجموعه‌های آموزشی)<sup>[\[29\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-29)</sup>.
- نیاز به افزایش Recall@k با دقت قابل‌قبول — HyDE اغلب نواحی مرتبط فضای برداری را «باز می‌کند»<sup>[\[30\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-30)</sup>.

تنظیمات معمول

- **LLM و prompt**: دستورالعمل «یک پاراگراف بنویس که به سوال … پاسخ دهد»؛ تصادفی‌بودن متوسط (مثلاً *temperature*≈0.7)<sup>[\[31\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-31)</sup>.
- **تعداد متن‌های فرضی**: ۱ تا ۵؛ میانگین‌گیری از embedding‌ها پایداری را افزایش می‌دهد<sup>[\[32\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-32)</sup>.
- **Embedder**: (m)Contriever بدون fine-tuning؛ امکان استفاده از encoder‌های fine-tune‌شده نیز وجود دارد (اثر HyDE حفظ می‌شود)<sup>[\[33\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-33)</sup>.
- **نرمال‌سازی embedding‌ها**: نُرم L2؛ حاصل‌ضرب داخلی معادل کسینوس است<sup>[\[34\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-34)</sup>.
- **بازیابی ترکیبی**: BM25+vector با بازرتبه‌بندی بعدی<sup>[\[35\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-35)</sup>.
- **بازرتبه‌بند**: Cross-Encoder (BERT re‑ranker)<sup>[\[36\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-36)</sup> یا ColBERT<sup>[\[37\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-37)</sup>.
- **ادغام** نتایج استراتژی‌های مختلف: RRF (*k*≈60)<sup>[\[38\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-38)</sup>.

پایش کیفیت/هزینه

- بازیابی: nDCG@k، Recall@k، MRR؛ end‑to‑end RAG: EM/F1 یا معیارهای *groundedness* (RAGAS/TruLens)<sup>[\[39\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-39)[\[40\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-40)</sup>.
- هزینه/تأخیر: تولید LLM و (در صورت وجود) بازرتبه‌بندی غالب هستند؛ با تنظیم تعداد «فرضیه‌ها» و طول پاسخ بهینه می‌شوند<sup>[\[41\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-41)</sup>.

## محدودیت‌ها و سوالات باز

- **توهم‌زایی** متن فرضی: LLM ممکن است خطاهای واقعی وارد کند؛ «متکی‌سازی» از طریق encoder و مجموعه اسناد خطر را کاهش می‌دهد اما کاملاً از بین نمی‌برد<sup>[\[42\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-42)</sup>.
- **محدودیت‌های دامنه‌ای/زبانی**: مزیت HyDE در دامنه‌های بسیار تخصصی و زبان‌های کم‌منبع کاهش می‌یابد<sup>[\[43\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-43)</sup>.
- **تأخیر و هزینه**: تولید LLM تأخیر و هزینه token اضافه می‌کند؛ برای سناریوهای آنلاین و «فرضیه‌»های طولانی حیاتی است<sup>[\[44\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-44)</sup>.
- **اخلاق و سوگیری**: استفاده از LLM‌های ایمن و فیلترگذاری ترجیح داده می‌شود<sup>[\[45\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-45)</sup>.

## جدول مقایسه‌ای روش‌ها

| روش                       | رده                | محل تولید متن                   | Encoder/فهرست        | بازرتبه‌بند (مرحله ۲)         | معیارهای معمول (نمونه)                             | هزینه/تأخیر                           | منابع                                                                                                                                                                                                                 |
|---------------------------|--------------------|---------------------------------|----------------------|------------------------------|----------------------------------------------------|---------------------------------------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| **HyDE**                  | Query→*hypo‑doc*   | سمت پرسش (LLM → پاراگراف)       | (m)Contriever؛ ANN   | BERT re‑rank / ColBERT / RRF | DL19 nDCG@10≈61.3؛ DL20≈57.9؛ ArguAna nDCG@10≈46.6 | \+ تولید LLM؛ + بازرتبه‌بندی (اختیاری) | <sup>[\[46\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-46)</sup>                                                                                                      |
| BM25                      | واژگانی            | —                               | فهرست معکوس          | اختیاری                      | ر.ک جدول (بالا)                                    | پایین (واژگانی)                       | <sup>[\[47\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-47)</sup>                                                                                                      |
| DPR / ANCE                | متراکم (ft)        | —                               | Bi‑encoder؛ ANN      | اختیاری                      | DL19 nDCG@10≈62–65                                 | متوسط (بدون LLM)                      | <sup>[\[48\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-48)[\[49\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-49)</sup> |
| doc2query / docTTTTTquery | گسترش سند          | سمت مجموعه (قبل از ایندکس‌گذاری) | BM25/sparse+expanded | اختیاری                      | بهبودهای BM25 روی MS MARCO                         | تولید آفلاین بالا؛ آنلاین سریع        | <sup>[\[50\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-50)[\[51\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-51)</sup> |
| PRF (Rocchio, RLM)        | QE بر اساس بازخورد | پرسش (بر اساس نتایج برتر)       | هر نوع               | اختیاری                      | افزایش Recall/خطرات انحراف                         | \+ پاس اضافی بازیابی                  | <sup>[\[52\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_note-52)</sup>                                                                                                      |

مقایسه HyDE و رویکردهای مرتبط

## همچنین ببینید

- BM25
- جستجو بر اساس بازنمایی‌های برداری،
- RAG
- بازخورد شبه‌مرتبط
- BEIR

## پیوندها

- مخزن HyDE: github.com/texttron/hyde.
- مستندات: Haystack — HyDE: docs.haystack.deepset.ai.
- مستندات: LangChain — HyDE Retriever: docs.langchain.com.

## منابع

- Manning, C. D.; Raghavan, P.; Schütze, H. (2008). *Introduction to Information Retrieval*. Cambridge University Press. ISBN 978‑0521865715.
- Robertson, S.; Zaragoza, H. (2009). *The Probabilistic Relevance Framework: BM25 and Beyond*. Foundations and Trends in IR, 3(4), 333–389. DOI:10.1561/1500000019.

## یادداشت‌ها

1.  <span id="cite_note-1">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-1) Gao, L.; Ma, X.; Lin, J.; Callan, J. (2023). ‘‘Precise Zero‑Shot Dense Retrieval without Relevance Labels (HyDE)’’. ACL 2023. pp. 1762–1777. DOI:10.18653/v1/2023.acl-long.99. <a href="https://arxiv.org/abs/2212.10496" class="external text" rel="nofollow">arXiv:2212.10496</a></span>
2.  <span id="cite_note-2">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-2) Gao, L. et al. (2023). ACL 2023, §3.2. DOI:10.18653/v1/2023.acl-long.99.</span>
3.  <span id="cite_note-3">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-3) Rocchio, J. (1971). ‘‘Relevance Feedback in Information Retrieval’’. In: Salton, G. (ed.) *The SMART Retrieval System*. Prentice‑Hall, pp. 313–323. ISBN 978‑0138145255.</span>
4.  <span id="cite_note-4">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-4) Lavrenko, V.; Croft, W. B. (2001). ‘‘Relevance‑Based Language Models’’. SIGIR. DOI:10.1145/383952.383972.</span>
5.  <span id="cite_note-5">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-5) Izacard, G. et al. (2021/2022). ‘‘Unsupervised Dense Information Retrieval with Contrastive Learning’’. <a href="https://arxiv.org/abs/2112.09118" class="external text" rel="nofollow">arXiv:2112.09118</a>.</span>
6.  <span id="cite_note-6">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-6) Karpukhin, V. et al. (2020). ‘‘Dense Passage Retrieval for Open‑Domain QA’’. EMNLP. DOI:10.18653/v1/2020.emnlp-main.550.</span>
7.  <span id="cite_note-7">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-7) Thakur, N. et al. (2021). ‘‘BEIR: A Heterogeneous Benchmark for Zero‑shot Evaluation of Information Retrieval Models’’. NeurIPS Datasets Track. <a href="https://arxiv.org/abs/2104.08663" class="external text" rel="nofollow">arXiv:2104.08663</a>.</span>
8.  <span id="cite_note-8">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-8) Gao, L. et al. (2023). DOI:10.18653/v1/2023.acl-long.99.</span>
9.  <span id="cite_note-9">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-9) Milvus Docs. ‘‘Similarity Metrics’’ — при L2‑нормализации векторов внутр. произведение эквивалентно косинусу. URL: <a href="https://milvus.io/docs/v2.2.x/metric.md" class="external free" rel="nofollow">https://milvus.io/docs/v2.2.x/metric.md</a></span>
10. <span id="cite_note-10">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-10) Gao, L.; Ma, X.; Lin, J.; Callan, J. (2023). ‘‘Precise Zero‑Shot Dense Retrieval without Relevance Labels (HyDE)’’. ACL 2023, §3–4. arXiv:2212.10496. DOI:10.18653/v1/2023.acl-long.99.</span>
11. <span id="cite_note-11">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-11) Izacard, G. et al. (2021/2022). ‘‘Unsupervised Dense Information Retrieval with Contrastive Learning (Contriever)’’. arXiv:2112.09118.</span>
12. <span id="cite_note-12">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-12) Gao, L. et al. (2023). Прил. (ablation): влияние числа гипотетических текстов и параметров генерации. arXiv:2212.10496.</span>
13. <span id="cite_note-13">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-13) Gao, L. et al. (2023). DOI:10.18653/v1/2023.acl-long.99.</span>
14. <span id="cite_note-14">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-14) Nogueira, R. et al. (2019). ‘‘Document Expansion by Query Prediction’’ (doc2query). <a href="https://arxiv.org/abs/1904.08375" class="external text" rel="nofollow">arXiv:1904.08375</a>.</span>
15. <span id="cite_note-15">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-15) Nogueira, R.; Lin, J. (2019). ‘‘From doc2query to docTTTTTquery’’ (tech report). <a href="https://cs.uwaterloo.ca/~jimmylin/publications/Nogueira_Lin_2019_docTTTTTquery-v2.pdf" class="external text" rel="nofollow">PDF</a></span>
16. <span id="cite_note-16">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-16) Rocchio, J. (1971); Lavrenko & Croft (2001), см. выше.</span>
17. <span id="cite_note-17">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-17) Nogueira, R.; Cho, K. (2019). ‘‘Passage Re‑ranking with BERT’’. <a href="https://arxiv.org/abs/1901.04085" class="external text" rel="nofollow">arXiv:1901.04085</a>.</span>
18. <span id="cite_note-18">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-18) Khattab, O.; Zaharia, M. (2020). ‘‘ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT’’. SIGIR. DOI:10.1145/3397271.3401075; <a href="https://arxiv.org/abs/2004.12832" class="external text" rel="nofollow">arXiv:2004.12832</a>.</span>
19. <span id="cite_note-19">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-19) Cormack, G. V.; Clarke, C. L. A.; Büttcher, S. (2009). ‘‘Reciprocal Rank Fusion Outperforms Condorcet and Nearly Optimally Combines Rankings’’. SIGIR. DOI:10.1145/1571941.1572114.</span>
20. <span id="cite_note-20">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-20) Gao, L. et al. (2023). Табл. 1. DOI:10.18653/v1/2023.acl-long.99.</span>
21. <span id="cite_note-21">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-21) Izacard, G. et al. (2022); сводные метрики — в Gao et al., 2023, табл. 1. arXiv:2112.09118.</span>
22. <span id="cite_note-22">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-22) Gao, L. et al. (2023). Табл. 1.</span>
23. <span id="cite_note-23">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-23) Karpukhin, V. et al. (2020); сводные — в Gao et al., 2023.</span>
24. <span id="cite_note-24">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-24) Xiong, L. et al. (2021). ICLR. <a href="https://arxiv.org/abs/2007.00808" class="external text" rel="nofollow">arXiv:2007.00808</a>.</span>
25. <span id="cite_note-25">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-25) Thakur, N. et al. (2021); сводные — в Gao et al., 2023, табл. 2. arXiv:2104.08663.</span>
26. <span id="cite_note-26">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-26) Izacard, G. et al. (2022); сводные — в Gao et al., 2023, табл. 2.</span>
27. <span id="cite_note-27">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-27) Gao, L. et al. (2023). Табл. 2.</span>
28. <span id="cite_note-28">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-28) Gao, L. et al. (2023). Табл. 3. DOI:10.18653/v1/2023.acl-long.99.</span>
29. <span id="cite_note-29">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-29) Gao, L. et al. (2023). §4–5.</span>
30. <span id="cite_note-30">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-30) Gao, L. et al. (2023). §4.2–4.3.</span>
31. <span id="cite_note-31">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-31) Gao, L. et al. (2023). §4.1.</span>
32. <span id="cite_note-32">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-32) Haystack Docs. ‘‘Hypothetical Document Embeddings (HyDE)’’ (инженерная справка). <a href="https://docs.haystack.deepset.ai/docs/hypothetical-document-embeddings-hyde" class="external text" rel="nofollow">docs.haystack.deepset.ai</a></span>
33. <span id="cite_note-33">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-33) Gao, L. et al. (2023). Табл. 6.</span>
34. <span id="cite_note-34">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-34) Milvus Docs. ‘‘Similarity Metrics’’.</span>
35. <span id="cite_note-35">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-35) Haystack × Milvus Integration (официальная док.). <a href="https://haystack.deepset.ai/integrations/milvus-document-store" class="external text" rel="nofollow">haystack.deepset.ai</a></span>
36. <span id="cite_note-36">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-36) Nogueira, R.; Cho, K. (2019). arXiv:1901.04085.</span>
37. <span id="cite_note-37">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-37) Khattab, O.; Zaharia, M. (2020). DOI:10.1145/3397271.3401075.</span>
38. <span id="cite_note-38">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-38) Cormack, G. V. et al. (2009). DOI:10.1145/1571941.1572114.</span>
39. <span id="cite_note-39">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-39) Manning, C. D.; Raghavan, P.; Schütze, H. (2008). *Introduction to Information Retrieval*. Cambridge Univ. Press. ISBN 978‑0521865715.</span>
40. <span id="cite_note-40">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-40) Es, S. et al. (2023). ‘‘RAGAS: Automated Evaluation of Retrieval‑Augmented Generation’’. <a href="https://arxiv.org/abs/2309.15217" class="external text" rel="nofollow">arXiv:2309.15217</a>.</span>
41. <span id="cite_note-41">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-41) Gao, L. et al. (2023). §5.</span>
42. <span id="cite_note-42">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-42) Gao, L. et al. (2023). §3.2; §4.1. DOI:10.18653/v1/2023.acl-long.99.</span>
43. <span id="cite_note-43">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-43) Gao, L. et al. (2023). Табл. 3; §4.4.</span>
44. <span id="cite_note-44">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-44) Gao, L. et al. (2023). §4–5.</span>
45. <span id="cite_note-45">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-45) Ouyang, L. et al. (2022). ‘‘Training language models to follow instructions with human feedback (InstructGPT)’’. NeurIPS. <a href="https://arxiv.org/abs/2203.02155" class="external text" rel="nofollow">arXiv:2203.02155</a>.</span>
46. <span id="cite_note-46">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-46) Gao, L. et al. (2023). Табл. 1–2.</span>
47. <span id="cite_note-47">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-47) Robertson, S.; Zaragoza, H. (2009). ‘‘The Probabilistic Relevance Framework: BM25 and Beyond’’. Found. Trends IR. DOI:10.1561/1500000019.</span>
48. <span id="cite_note-48">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-48) Karpukhin, V. et al. (2020). DOI:10.18653/v1/2020.emnlp-main.550.</span>
49. <span id="cite_note-49">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-49) Xiong, L. et al. (2021). <a href="https://arxiv.org/abs/2007.00808" class="external text" rel="nofollow">arXiv:2007.00808</a>.</span>
50. <span id="cite_note-50">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-50) Nogueira, R. et al. (2019). arXiv:1904.08375.</span>
51. <span id="cite_note-51">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-51) Nogueira, R.; Lin, J. (2019). tech report.</span>
52. <span id="cite_note-52">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(FA)#cite_ref-52) Rocchio, J. (1971). SMART; Lavrenko & Croft (2001) SIGIR.</span>
