---
title: "Hypothetical Document Embeddings (HyDE)"
source: "https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)"
wiki: "systems-analysis.info/eng"
article: "Hypothetical_Document_Embeddings_(HyDE)"
language: "en"
categories:
  - "Category:English"
  - "Category:Large language models"
  - "Category:Prompt engineering"
  - "Category:Technology"
revision_id: 187
wiki_created_at: 2026-09-06T22:18:38Z
wiki_modified_at: 2026-09-06T22:18:38Z
downloaded_at: 2026-09-07T22:21:41Z
---

# Hypothetical Document Embeddings (HyDE)

**Hypothetical Document Expansion (HyDE)** is a method for improving vector retrieval and retrieval-augmented generation (RAG), in which a large language model (LLM) generates a "hypothetical document" based on an initial query; this text is then vectorized by an encoder, and the search for real documents is performed based on proximity to the resulting vector. The approach allows leveraging "relevance patterns" encoded by the LLM and "grounding" them in a corpus using dense embeddings<sup>[\[1\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-1)</sup>.

## Definition and Intuition

HyDE decomposes the search task into two stages:

\(1\) The LLM creates an "example of a relevant answer" (*hypothetical document*) for the query, thereby modeling the features of relevance;

\(2\) A contrastive encoder (e.g., Contriever) translates this text into a vector, which is then used to retrieve real documents from the index. The generated text may contain factual errors, but what is important are the thematic and terminological patterns captured by the encoder<sup>[\[2\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-2)</sup>.

## History and Sources

The idea of enhancing search with synthetic texts dates back to work on query expansion and pseudo-relevance feedback (PRF): the Rocchio algorithm and relevance language models<sup>[\[3\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-3)[\[4\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-4)</sup>. For dense retrieval, contrastively trained encoders like Contriever<sup>[\[5\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-5)</sup> and Dense Passage Retrieval (DPR)<sup>[\[6\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-6)</sup> were used. The BEIR benchmark standardized zero-shot evaluation<sup>[\[7\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-7)</sup>. Against this backdrop, HyDE was proposed as a way to "inject" relevance knowledge from an LLM into the zero-shot setting without fine-tuning the encoder<sup>[\[8\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-8)</sup>.

## Method and Formalization

Let the document corpus be $\mathcal{D} = \{ d_{1},\ldots,d_{N}\}$, and a text encoder $E:\text{text} \rightarrow {\mathbb{R}}^{n}$ define the vector representations of documents $\mathbf{v}_{d} = E(d)$. To measure proximity, either cosine similarity or the dot product is used; an important note: **the dot product is equivalent to cosine similarity only when both vectors have a unit L2-norm** ($\|\mathbf{u}\| = \|\mathbf{v}\| = 1$)<sup>[\[9\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-9)</sup>.

HyDE redefines the query representation through a "hypothetical document" generated by an LLM. Formally:

$$
\begin{matrix}
 & \text{(1) Generation of the hypothetical document:} & & {\overset{\sim}{d}\; = \; G\!\left( q;\,{inst} \right),} \\
 & \text{(2) Embedding of the hypothetical document:} & & {\mathbf{v}_{h}\; = \; E(\overset{\sim}{d}),} \\
 & \text{(3) Nearest neighbor search:} & & {\mathcal{R}_{k}(q)\; = \;{TopK}_{d \in \mathcal{D}}\; S\!\left( \mathbf{v}_{h},\mathbf{v}_{d} \right),}
\end{matrix}
$$

where $G$ is an LLM with an instruction $inst$ (e.g., "Write a paragraph that answers the question..."), $S$ is the similarity measure (cosine or IP with normalization), and $\mathcal{R}_{k}(q)$ is the set of $k$ documents with the highest similarity<sup>[\[10\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-10)[\[11\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-11)</sup>.

In engineering practice, it is common to generate **multiple** hypothetical documents and aggregate their representations to increase robustness:

$$
{\overset{\sim}{d}}^{(j)} = G\!\left( q;\,{inst},\xi_{j} \right),\quad\mathbf{v}_{h}\; = \;\frac{1}{m}\sum\limits_{j = 1}^{m}E\!\left( {\overset{\sim}{d}}^{(j)} \right),
$$

where $\xi_{j}$ are stochastic decoding parameters (e.g., temperature/top-p). This ensembling improves Recall with a moderate increase in latency<sup>[\[12\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-12)</sup>.

### Basic HyDE Pipeline

    # 1) prompt(query) -> hypothetical_doc
    # 2) embed(hypothetical_doc) -> v_h
    # 3) retrieve(index, v_h, k) -> candidates
    # 4) (optional) rerank(query, candidates) -> topN
    # 5) (for RAG) stuff / map-reduce / refine on topN

### Relation to Other Methods (QE, doc2query, PRF)

- **QE (Query Expansion)** adds terms to the query; HyDE, instead, generates an entire "quasi-document," which aligns better with dense encoders<sup>[\[13\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-13)</sup>.
- **doc2query / docTTTTTquery** expand **documents** with synthetic queries before indexing<sup>[\[14\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-14)[\[15\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-15)</sup>; HyDE expands the **query** on the fly, without requiring re-indexing.
- **PRF** (Rocchio, Relevance LM) updates the query vector based on the top results; HyDE extracts the "relevance pattern" directly from the LLM and then "grounds" it via retrieval from the corpus<sup>[\[16\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-16)</sup>.

## Integration into RAG and Reranking

In RAG, HyDE is applied as the first retrieval stage: hypothetical document → embedding → k candidates. This is followed by reranking using BERT-class cross-encoders<sup>[\[17\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-17)</sup> or a late interaction model like ColBERT<sup>[\[18\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-18)</sup>. For merging result lists (e.g., a BM25+vector hybrid), RRF (*reciprocal rank fusion*) is typically used: $\operatorname{RRF}(d) = \sum\limits_{r \in \mathcal{R}}\frac{1}{k + \operatorname{rank}_{r}(d)},\qquad k \approx 60.$ The RRF method consistently improves the aggregate quality of combined rankings<sup>[\[19\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-19)</sup>.

## Evaluation on Benchmarks (BEIR et al.)

The original paper evaluates HyDE in a zero-shot setting on TREC DL’19/20 (web search) and on a subset of the BEIR collections (Scifact, ArguAna, TREC-COVID, FiQA, DBPedia, TREC-NEWS, Climate-FEVER). A fragment of the results—*as of July 2023*:

| Method                    | DL19                   | DL20                   | Source                                                                                                      |
|---------------------------|------------------------|------------------------|-------------------------------------------------------------------------------------------------------------|
| BM25                      | 30.1 / 50.6 / 75.0     | 28.6 / 48.0 / 78.6     | <sup>[\[20\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-20)</sup> |
| Contriever (unsup.)       | 24.0 / 44.5 / 74.6     | 24.0 / 42.1 / 75.4     | <sup>[\[21\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-21)</sup> |
| **HyDE** (Contriever+LLM) | **41.8 / 61.3 / 88.0** | **38.2 / 57.9 / 84.4** | <sup>[\[22\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-22)</sup> |
| DPR (ft)                  | 36.5 / 62.2 / 76.9     | 41.8 / 65.3 / 81.4     | <sup>[\[23\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-23)</sup> |
| ANCE (ft)                 | 37.1 / 64.5 / 75.5     | 40.8 / 64.6 / 77.6     | <sup>[\[24\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-24)</sup> |

*TREC DL19/20 (Web Search)* — mAP / nDCG@10 / Recall@1k

| Method     | Scifact         | ArguAna         | TREC‑COVID      | FiQA        | DBPedia     | TREC‑NEWS       | Climate‑FEVER   | Source                                                                                                      |
|------------|-----------------|-----------------|-----------------|-------------|-------------|-----------------|-----------------|-------------------------------------------------------------------------------------------------------------|
| BM25       | 67.9 / 92.5     | 39.7 / 93.2     | **59.5 / 49.8** | 23.6 / 54.0 | 31.8 / 46.8 | 39.5 / 44.7     | 16.5 / 42.5     | <sup>[\[25\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-25)</sup> |
| Contriever | 64.9 / 92.6     | 37.9 / 90.1     | 27.3 / 17.2     | 24.5 / 56.2 | 29.2 / 45.3 | 34.8 / 42.3     | 15.5 / 44.1     | <sup>[\[26\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-26)</sup> |
| **HyDE**   | **69.1 / 96.4** | **46.6 / 97.9** | 59.3 / 41.4     | 27.3 / 62.1 | 36.8 / 47.2 | **44.0 / 50.9** | **22.3 / 53.0** | <sup>[\[27\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-27)</sup> |

*BEIR (Selection of Datasets)* — nDCG@10 / Recall@100

HyDE also improves MRR@100 on the multilingual Mr.TyDi datasets (sw/ko/ja/bn) relative to mContriever<sup>[\[28\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-28)</sup>.

## Practical Recommendations

When to use HyDE

- Zero-shot/transfer settings (no relevance labels; domain dissimilarity from training corpora)<sup>[\[29\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-29)</sup>.
- When higher Recall@k is needed with acceptable precision—HyDE often "unlocks" relevant areas of the vector space<sup>[\[30\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-30)</sup>.

Typical Settings

- **LLM and prompt**: An instruction like "Write a paragraph that answers the question..."; moderate stochasticity (e.g., *temperature*≈0.7)<sup>[\[31\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-31)</sup>.
- **Number of hypothetical documents**: 1–5; averaging embeddings improves robustness<sup>[\[32\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-32)</sup>.
- **Embedder**: (m)Contriever without fine-tuning; fine-tuned encoders can also be used (the HyDE effect persists)<sup>[\[33\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-33)</sup>.
- **Embedding normalization**: L2-norm; the inner product is equivalent to cosine similarity<sup>[\[34\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-34)</sup>.
- **Hybrid retrieval**: BM25+vector followed by reranking<sup>[\[35\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-35)</sup>.
- **Reranker**: Cross-Encoder (BERT re-ranker)<sup>[\[36\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-36)</sup> or ColBERT<sup>[\[37\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-37)</sup>.
- **Merging** results from different strategies: RRF (*k*≈60)<sup>[\[38\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-38)</sup>.

Quality/Cost Monitoring

- Retrieval: nDCG@k, Recall@k, MRR; end-to-end RAG: EM/F1 or *groundedness* metrics (RAGAS/TruLens)<sup>[\[39\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-39)[\[40\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-40)</sup>.
- Cost/latency: Dominated by LLM generation and (if applicable) reranking; optimized by the number of hypothetical documents and response length<sup>[\[41\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-41)</sup>.

## Limitations and Open Questions

- **Hallucinations** in the hypothetical document: The LLM can introduce factual errors; "grounding" via the encoder and corpus reduces the risk but does not eliminate it entirely<sup>[\[42\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-42)</sup>.
- **Domain/Language limitations**: The benefit of HyDE diminishes in highly specialized domains and for low-resource languages<sup>[\[43\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-43)</sup>.
- **Latency and cost**: LLM generation adds delay and token costs; critical for online scenarios and long hypothetical documents<sup>[\[44\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-44)</sup>.
- **Ethics and biases**: It is preferable to use safe LLMs and content filtering<sup>[\[45\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-45)</sup>.

## Comparative Table of Methods

| Method                    | Class            | Where text is generated                  | Encoder/Index        | Reranker (2nd stage)         | Typical Metrics (Example)                          | Cost/Latency                              | Sources                                                                                                                                                                                                     |
|---------------------------|------------------|------------------------------------------|----------------------|------------------------------|----------------------------------------------------|-------------------------------------------|-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| **HyDE**                  | Query→*hypo‑doc* | On the query side (LLM → paragraph)      | (m)Contriever; ANN   | BERT re‑rank / ColBERT / RRF | DL19 nDCG@10≈61.3; DL20≈57.9; ArguAna nDCG@10≈46.6 | \+ LLM generation; + reranking (opt.)     | <sup>[\[46\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-46)</sup>                                                                                                 |
| BM25                      | Lexical          | —                                        | Inverted index       | Optional                     | see table (above)                                  | Low (lexical)                             | <sup>[\[47\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-47)</sup>                                                                                                 |
| DPR / ANCE                | Dense (ft)       | —                                        | Bi‑encoder; ANN      | Optional                     | DL19 nDCG@10≈62–65                                 | Medium (no LLM)                           | <sup>[\[48\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-48)[\[49\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-49)</sup> |
| doc2query / docTTTTTquery | Doc. expansion   | On the collection side (before indexing) | BM25/sparse+expanded | Optional                     | Improvements over BM25 on MS MARCO                 | High offline generation cost; fast online | <sup>[\[50\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-50)[\[51\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-51)</sup> |
| PRF (Rocchio, RLM)        | QE via feedback  | Query (from top results)                 | Any                  | Optional                     | Increased Recall / risk of drift                   | \+ additional retrieval pass              | <sup>[\[52\]](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_note-52)</sup>                                                                                                 |

Comparison of HyDE and Related Approaches

## See also

- [Vector database](https://systems-analysis.info/eng/Vector_database "Vector database")
- [Retrieval-augmented generation (RAG)](https://systems-analysis.info/eng/Retrieval-augmented_generation_(RAG) "Retrieval-augmented generation (RAG)")

## External links

- HyDE Repository: <a href="https://github.com/texttron/hyde" class="external text" rel="nofollow">github.com/texttron/hyde</a>.
- Documentation: Haystack — HyDE: <a href="https://docs.haystack.deepset.ai/docs/hypothetical-document-embeddings-hyde" class="external text" rel="nofollow">docs.haystack.deepset.ai</a>.
- Documentation: LangChain — HyDE Retriever: <a href="https://docs.langchain.com/oss/javascript/integrations/retrievers/hyde" class="external text" rel="nofollow">docs.langchain.com</a>.

## Literature

- Manning, C. D.; Raghavan, P.; Schütze, H. (2008). *Introduction to Information Retrieval*. Cambridge University Press. ISBN 978‑0521865715.
- Robertson, S.; Zaragoza, H. (2009). *The Probabilistic Relevance Framework: BM25 and Beyond*. Foundations and Trends in IR, 3(4), 333–389. DOI:10.1561/1500000019.

## References

1.  <span id="cite_note-1">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-1) Gao, L.; Ma, X.; Lin, J.; Callan, J. (2023). ‘‘Precise Zero‑Shot Dense Retrieval without Relevance Labels (HyDE)’’. ACL 2023. pp. 1762–1777. DOI:10.18653/v1/2023.acl-long.99. <a href="https://arxiv.org/abs/2212.10496" class="external text" rel="nofollow">arXiv:2212.10496</a></span>
2.  <span id="cite_note-2">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-2) Gao, L. et al. (2023). ACL 2023, §3.2. DOI:10.18653/v1/2023.acl-long.99.</span>
3.  <span id="cite_note-3">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-3) Rocchio, J. (1971). ‘‘Relevance Feedback in Information Retrieval’’. In: Salton, G. (ed.) *The SMART Retrieval System*. Prentice‑Hall, pp. 313–323. ISBN 978‑0138145255.</span>
4.  <span id="cite_note-4">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-4) Lavrenko, V.; Croft, W. B. (2001). ‘‘Relevance‑Based Language Models’’. SIGIR. DOI:10.1145/383952.383972.</span>
5.  <span id="cite_note-5">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-5) Izacard, G. et al. (2021/2022). ‘‘Unsupervised Dense Information Retrieval with Contrastive Learning’’. <a href="https://arxiv.org/abs/2112.09118" class="external text" rel="nofollow">arXiv:2112.09118</a>.</span>
6.  <span id="cite_note-6">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-6) Karpukhin, V. et al. (2020). ‘‘Dense Passage Retrieval for Open‑Domain QA’’. EMNLP. DOI:10.18653/v1/2020.emnlp-main.550.</span>
7.  <span id="cite_note-7">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-7) Thakur, N. et al. (2021). ‘‘BEIR: A Heterogeneous Benchmark for Zero‑shot Evaluation of Information Retrieval Models’’. NeurIPS Datasets Track. <a href="https://arxiv.org/abs/2104.08663" class="external text" rel="nofollow">arXiv:2104.08663</a>.</span>
8.  <span id="cite_note-8">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-8) Gao, L. et al. (2023). DOI:10.18653/v1/2023.acl-long.99.</span>
9.  <span id="cite_note-9">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-9) Milvus Docs. ‘‘Similarity Metrics’’ — With L2-normalized vectors, the inner product is equivalent to cosine similarity. URL: <a href="https://milvus.io/docs/v2.2.x/metric.md" class="external free" rel="nofollow">https://milvus.io/docs/v2.2.x/metric.md</a></span>
10. <span id="cite_note-10">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-10) Gao, L.; Ma, X.; Lin, J.; Callan, J. (2023). ‘‘Precise Zero‑Shot Dense Retrieval without Relevance Labels (HyDE)’’. ACL 2023, §3–4. arXiv:2212.10496. DOI:10.18653/v1/2023.acl-long.99.</span>
11. <span id="cite_note-11">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-11) Izacard, G. et al. (2021/2022). ‘‘Unsupervised Dense Information Retrieval with Contrastive Learning (Contriever)’’. arXiv:2112.09118.</span>
12. <span id="cite_note-12">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-12) Gao, L. et al. (2023). Appendix (ablation): impact of the number of hypothetical documents and generation parameters. arXiv:2212.10496.</span>
13. <span id="cite_note-13">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-13) Gao, L. et al. (2023). DOI:10.18653/v1/2023.acl-long.99.</span>
14. <span id="cite_note-14">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-14) Nogueira, R. et al. (2019). ‘‘Document Expansion by Query Prediction’’ (doc2query). <a href="https://arxiv.org/abs/1904.08375" class="external text" rel="nofollow">arXiv:1904.08375</a>.</span>
15. <span id="cite_note-15">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-15) Nogueira, R.; Lin, J. (2019). ‘‘From doc2query to docTTTTTquery’’ (tech report). <a href="https://cs.uwaterloo.ca/~jimmylin/publications/Nogueira_Lin_2019_docTTTTTquery-v2.pdf" class="external text" rel="nofollow">PDF</a></span>
16. <span id="cite_note-16">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-16) Rocchio, J. (1971); Lavrenko & Croft (2001), see above.</span>
17. <span id="cite_note-17">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-17) Nogueira, R.; Cho, K. (2019). ‘‘Passage Re‑ranking with BERT’’. <a href="https://arxiv.org/abs/1901.04085" class="external text" rel="nofollow">arXiv:1901.04085</a>.</span>
18. <span id="cite_note-18">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-18) Khattab, O.; Zaharia, M. (2020). ‘‘ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT’’. SIGIR. DOI:10.1145/3397271.3401075; <a href="https://arxiv.org/abs/2004.12832" class="external text" rel="nofollow">arXiv:2004.12832</a>.</span>
19. <span id="cite_note-19">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-19) Cormack, G. V.; Clarke, C. L. A.; Büttcher, S. (2009). ‘‘Reciprocal Rank Fusion Outperforms Condorcet and Nearly Optimally Combines Rankings’’. SIGIR. DOI:10.1145/1571941.1572114.</span>
20. <span id="cite_note-20">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-20) Gao, L. et al. (2023). Table 1. DOI:10.18653/v1/2023.acl-long.99.</span>
21. <span id="cite_note-21">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-21) Izacard, G. et al. (2022); summary metrics in Gao et al., 2023, Table 1. arXiv:2112.09118.</span>
22. <span id="cite_note-22">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-22) Gao, L. et al. (2023). Table 1.</span>
23. <span id="cite_note-23">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-23) Karpukhin, V. et al. (2020); summary in Gao et al., 2023.</span>
24. <span id="cite_note-24">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-24) Xiong, L. et al. (2021). ICLR. <a href="https://arxiv.org/abs/2007.00808" class="external text" rel="nofollow">arXiv:2007.00808</a>.</span>
25. <span id="cite_note-25">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-25) Thakur, N. et al. (2021); summary in Gao et al., 2023, Table 2. arXiv:2104.08663.</span>
26. <span id="cite_note-26">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-26) Izacard, G. et al. (2022); summary in Gao et al., 2023, Table 2.</span>
27. <span id="cite_note-27">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-27) Gao, L. et al. (2023). Table 2.</span>
28. <span id="cite_note-28">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-28) Gao, L. et al. (2023). Table 3. DOI:10.18653/v1/2023.acl-long.99.</span>
29. <span id="cite_note-29">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-29) Gao, L. et al. (2023). §4–5.</span>
30. <span id="cite_note-30">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-30) Gao, L. et al. (2023). §4.2–4.3.</span>
31. <span id="cite_note-31">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-31) Gao, L. et al. (2023). §4.1.</span>
32. <span id="cite_note-32">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-32) Haystack Docs. ‘‘Hypothetical Document Embeddings (HyDE)’’ (engineering reference). <a href="https://docs.haystack.deepset.ai/docs/hypothetical-document-embeddings-hyde" class="external text" rel="nofollow">docs.haystack.deepset.ai</a></span>
33. <span id="cite_note-33">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-33) Gao, L. et al. (2023). Table 6.</span>
34. <span id="cite_note-34">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-34) Milvus Docs. ‘‘Similarity Metrics’’.</span>
35. <span id="cite_note-35">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-35) Haystack × Milvus Integration (official docs). <a href="https://haystack.deepset.ai/integrations/milvus-document-store" class="external text" rel="nofollow">haystack.deepset.ai</a></span>
36. <span id="cite_note-36">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-36) Nogueira, R.; Cho, K. (2019). arXiv:1901.04085.</span>
37. <span id="cite_note-37">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-37) Khattab, O.; Zaharia, M. (2020). DOI:10.1145/3397271.3401075.</span>
38. <span id="cite_note-38">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-38) Cormack, G. V. et al. (2009). DOI:10.1145/1571941.1572114.</span>
39. <span id="cite_note-39">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-39) Manning, C. D.; Raghavan, P.; Schütze, H. (2008). *Introduction to Information Retrieval*. Cambridge Univ. Press. ISBN 978‑0521865715.</span>
40. <span id="cite_note-40">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-40) Es, S. et al. (2023). ‘‘RAGAS: Automated Evaluation of Retrieval‑Augmented Generation’’. <a href="https://arxiv.org/abs/2309.15217" class="external text" rel="nofollow">arXiv:2309.15217</a>.</span>
41. <span id="cite_note-41">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-41) Gao, L. et al. (2023). §5.</span>
42. <span id="cite_note-42">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-42) Gao, L. et al. (2023). §3.2; §4.1. DOI:10.18653/v1/2023.acl-long.99.</span>
43. <span id="cite_note-43">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-43) Gao, L. et al. (2023). Table 3; §4.4.</span>
44. <span id="cite_note-44">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-44) Gao, L. et al. (2023). §4–5.</span>
45. <span id="cite_note-45">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-45) Ouyang, L. et al. (2022). ‘‘Training language models to follow instructions with human feedback (InstructGPT)’’. NeurIPS. <a href="https://arxiv.org/abs/2203.02155" class="external text" rel="nofollow">arXiv:2203.02155</a>.</span>
46. <span id="cite_note-46">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-46) Gao, L. et al. (2023). Tables 1–2.</span>
47. <span id="cite_note-47">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-47) Robertson, S.; Zaragoza, H. (2009). ‘‘The Probabilistic Relevance Framework: BM25 and Beyond’’. Found. Trends IR. DOI:10.1561/1500000019.</span>
48. <span id="cite_note-48">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-48) Karpukhin, V. et al. (2020). DOI:10.18653/v1/2020.emnlp-main.550.</span>
49. <span id="cite_note-49">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-49) Xiong, L. et al. (2021). <a href="https://arxiv.org/abs/2007.00808" class="external text" rel="nofollow">arXiv:2007.00808</a>.</span>
50. <span id="cite_note-50">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-50) Nogueira, R. et al. (2019). arXiv:1904.08375.</span>
51. <span id="cite_note-51">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-51) Nogueira, R.; Lin, J. (2019). tech report.</span>
52. <span id="cite_note-52">[↑](https://systems-analysis.info/eng/Hypothetical_Document_Embeddings_(HyDE)#cite_ref-52) Rocchio, J. (1971). SMART; Lavrenko & Croft (2001) SIGIR.</span>
