---
title: "Hypothetical Document Embeddings (HyDE) (DE)"
source: "https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)"
wiki: "systems-analysis.info/int"
article: "Hypothetical_Document_Embeddings_(HyDE)_(DE)"
language: "de"
categories:
  - "Category:German"
  - "Category:Large language models"
  - "Category:Prompt engineering"
revision_id: 3108
wiki_created_at: 2026-09-06T23:15:49Z
wiki_modified_at: 2026-09-06T23:15:49Z
downloaded_at: 2026-09-07T22:55:12Z
---

# Hypothetical Document Embeddings (HyDE) (DE)

**Hypothetical Document Expansion (HyDE)** ist eine Methode zur Verbesserung des Vektor-Retrievals und der Retrieval-Augmented Generation (RAG), bei der ein großes Sprachmodell (LLM) aus einer ursprünglichen Anfrage ein „hypothetisches Dokument“ generiert. Dieser Text wird anschließend von einem Encoder vektorisiert, und die Suche nach realen Dokumenten erfolgt auf der Grundlage der Nähe zum resultierenden Vektor. Der Ansatz ermöglicht es, die von einem LLM kodierten „Relevanzmuster“ zu nutzen und sie mithilfe von dichten Embeddings im Korpus zu „verankern“ (to ground)<sup>[\[1\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-1)</sup>.

## Definition und Intuition

HyDE zerlegt die Suchaufgabe in zwei Schritte:

\(1\) Ein LLM erstellt ein „Beispiel für eine relevante Antwort“ (*hypothetisches Dokument*) auf die Anfrage und modelliert dadurch Relevanzmerkmale.

\(2\) Ein kontrastiver Encoder (z. B. Contriever) wandelt diesen Text in einen Vektor um, anhand dessen reale Dokumente aus dem Index abgerufen werden. Der generierte Text kann sachliche Fehler enthalten, entscheidend sind jedoch die thematischen und terminologischen Muster, die der Encoder erfasst<sup>[\[2\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-2)</sup>.

## Geschichte und Quellen

Die Idee, die Suche durch synthetische Texte zu erweitern, geht auf Arbeiten zur Query Expansion und zum Pseudo-Relevanz-Feedback (PRF) zurück: den Rocchio-Algorithmus und Relevanz-Sprachmodelle<sup>[\[3\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-3)[\[4\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-4)</sup>. Für das Dense Retrieval wurden kontrastiv trainierte Encoder wie Contriever<sup>[\[5\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-5)</sup> und Dense Passage Retrieval (DPR)<sup>[\[6\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-6)</sup> eingesetzt. Der BEIR-Benchmark standardisierte die Zero-Shot-Evaluation<sup>[\[7\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-7)</sup>. Vor diesem Hintergrund wurde HyDE als eine Methode vorgeschlagen, um Relevanzwissen aus einem LLM in ein Zero-Shot-Szenario „einzubringen“, ohne den Encoder nachtrainieren zu müssen<sup>[\[8\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-8)</sup>.

## Methode und Formalisierung

Sei $\mathcal{D} = \{ d_{1},\ldots,d_{N}\}$ ein Korpus von Dokumenten und $E:\text{text} \rightarrow {\mathbb{R}}^{n}$ ein Text-Encoder, der Vektorrepräsentationen von Dokumenten $\mathbf{v}_{d} = E(d)$ erzeugt. Zur Messung der Ähnlichkeit wird entweder die Kosinus-Ähnlichkeit oder das Skalarprodukt verwendet. Ein wichtiger Hinweis: \*\*Das Skalarprodukt entspricht der Kosinus-Ähnlichkeit nur, wenn beide Vektoren eine L2-Norm von eins haben\*\* ($\|\mathbf{u}\| = \|\mathbf{v}\| = 1$)<sup>[\[9\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-9)</sup>.

HyDE definiert die Repräsentation einer Anfrage über ein vom LLM generiertes „hypothetisches Dokument“. Formal ausgedrückt:

$$
\begin{matrix}
 & \text{(1) Generierung des hypothetischen Textes:} & & {\overset{\sim}{d}\; = \; G\!\left( q;\,{inst} \right),} \\
 & \text{(2) Einbettung des hypothetischen Textes:} & & {\mathbf{v}_{h}\; = \; E(\overset{\sim}{d}),} \\
 & \text{(3) Suche der nächsten Nachbarn:} & & {\mathcal{R}_{k}(q)\; = \;{TopK}_{\, d \in \mathcal{D}}\; S\!\left( \mathbf{v}_{h},\mathbf{v}_{d} \right),}
\end{matrix}
$$

wo $G$ das LLM mit der Instruktion $inst$ ist (z. B. „Schreibe einen Absatz, der die folgende Frage beantwortet …“), $S$ das Ähnlichkeitsmaß (Kosinus oder IP mit Normalisierung) und $\mathcal{R}_{k}(q)$ die Menge der $k$ Dokumente mit der höchsten Ähnlichkeit ist<sup>[\[10\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-10)[\[11\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-11)</sup>.

In der Praxis werden oft \*\*mehrere\*\* hypothetische Texte generiert und ihre Repräsentationen aggregiert, was die Robustheit erhöht:

$$
{\overset{\sim}{d}}^{(j)} = G\!\left( q;\,{inst},\xi_{j} \right),\quad\mathbf{v}_{h}\; = \;\frac{1}{m}\sum\limits_{j = 1}^{m}E\!\left( {\overset{\sim}{d}}^{(j)} \right),
$$

wo $\xi_{j}$ stochastische Dekodierungsparameter sind (z. B. Temperature/Top-p). Ein solches Ensembling verbessert den Recall bei einem moderaten Anstieg der Latenz<sup>[\[12\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-12)</sup>.

### Grundlegende HyDE-Pipeline

    # 1) prompt(query) -> hypothetical_doc
    # 2) embed(hypothetical_doc) -> v_h
    # 3) retrieve(index, v_h, k) -> candidates
    # 4) (optional) rerank(query, candidates) -> topN
    # 5) (for RAG) stuff / map-reduce / refine on topN

### Beziehung zu anderen Methoden (QE, doc2query, PRF)

- **QE (Query Expansion)** fügt der Anfrage Terme hinzu; HyDE generiert stattdessen ein ganzes „Quasi-Dokument“, was besser mit dichten Encodern harmoniert<sup>[\[13\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-13)</sup>.
- **doc2query / docTTTTTquery** erweitern **Dokumente** vor der Indizierung mit synthetischen Anfragen<sup>[\[14\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-14)[\[15\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-15)</sup>; HyDE erweitert die **Anfrage** on-the-fly, ohne eine Neuindizierung zu erfordern.
- **PRF** (Rocchio, Relevance LM) aktualisiert den Anfragevektor basierend auf den Top-Ergebnissen; HyDE extrahiert das „Relevanzmuster“ direkt aus dem LLM und „verankert“ es dann durch Retrieval im Korpus<sup>[\[16\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-16)</sup>.

## Integration in RAG und Re-Ranking

In RAG wird HyDE als erste Retrieval-Stufe eingesetzt: hypothetisches Dokument → Embedding → k Kandidaten. Anschließend wird ein Re-Ranking durchgeführt, z. B. mit BERT-basierten Cross-Encodern<sup>[\[17\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-17)</sup> oder durch späte Interaktion mit ColBERT<sup>[\[18\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-18)</sup>. Zur Fusion von Ranglisten (z. B. bei hybrider Suche mit BM25+Vektor) wird typischerweise RRF (*Reciprocal Rank Fusion*) verwendet: $\operatorname{RRF}(d) = \sum\limits_{r \in \mathcal{R}}\frac{1}{k + \operatorname{rank}_{r}(d)},\qquad k \approx 60.$ Die RRF-Methode verbessert konsistent die Gesamtqualität der fusionierten Ranglisten<sup>[\[19\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-19)</sup>.

## Evaluation auf Benchmarks (BEIR etc.)

Die Originalarbeit evaluiert HyDE im Zero-Shot-Szenario auf TREC DL’19/20 (Websuche) und auf einer Teilmenge der BEIR-Sammlungen (Scifact, ArguAna, TREC‑COVID, FiQA, DBPedia, TREC‑NEWS, Climate‑FEVER). Auszug der Ergebnisse – *Stand Juli 2023*:

| Methode                   | DL19                   | DL20                   | Quelle                                                                                                           |
|---------------------------|------------------------|------------------------|------------------------------------------------------------------------------------------------------------------|
| BM25                      | 30.1 / 50.6 / 75.0     | 28.6 / 48.0 / 78.6     | <sup>[\[20\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-20)</sup> |
| Contriever (unsup.)       | 24.0 / 44.5 / 74.6     | 24.0 / 42.1 / 75.4     | <sup>[\[21\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-21)</sup> |
| **HyDE** (Contriever+LLM) | **41.8 / 61.3 / 88.0** | **38.2 / 57.9 / 84.4** | <sup>[\[22\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-22)</sup> |
| DPR (ft)                  | 36.5 / 62.2 / 76.9     | 41.8 / 65.3 / 81.4     | <sup>[\[23\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-23)</sup> |
| ANCE (ft)                 | 37.1 / 64.5 / 75.5     | 40.8 / 64.6 / 77.6     | <sup>[\[24\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-24)</sup> |

*TREC DL19/20 (Websuche)* – mAP / nDCG@10 / Recall@1k

| Methode    | Scifact         | ArguAna         | TREC‑COVID      | FiQA        | DBPedia     | TREC‑NEWS       | Climate‑FEVER   | Quelle                                                                                                           |
|------------|-----------------|-----------------|-----------------|-------------|-------------|-----------------|-----------------|------------------------------------------------------------------------------------------------------------------|
| BM25       | 67.9 / 92.5     | 39.7 / 93.2     | **59.5 / 49.8** | 23.6 / 54.0 | 31.8 / 46.8 | 39.5 / 44.7     | 16.5 / 42.5     | <sup>[\[25\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-25)</sup> |
| Contriever | 64.9 / 92.6     | 37.9 / 90.1     | 27.3 / 17.2     | 24.5 / 56.2 | 29.2 / 45.3 | 34.8 / 42.3     | 15.5 / 44.1     | <sup>[\[26\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-26)</sup> |
| **HyDE**   | **69.1 / 96.4** | **46.6 / 97.9** | 59.3 / 41.4     | 27.3 / 62.1 | 36.8 / 47.2 | **44.0 / 50.9** | **22.3 / 53.0** | <sup>[\[27\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-27)</sup> |

*BEIR (Auswahl von Datensätzen)* – nDCG@10 / Recall@100

HyDE verbessert auch den MRR@100 auf den mehrsprachigen Mr.TyDi-Datensätzen (sw/ko/ja/bn) im Vergleich zu mContriever<sup>[\[28\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-28)</sup>.

## Praktische Empfehlungen

Wann HyDE einsetzen

- Zero-Shot- oder Transfer-Szenarien (keine Relevanz-Labels vorhanden; Domänenunterschied zu den Trainingskorpora)<sup>[\[29\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-29)</sup>.
- Wenn eine Erhöhung des Recall@k bei akzeptabler Präzision erforderlich ist – HyDE „erschließt“ oft relevante Bereiche des Vektorraums<sup>[\[30\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-30)</sup>.

Typische Konfigurationen

- **LLM und Prompt**: Eine Anweisung wie „Schreibe einen Absatz, der die folgende Frage beantwortet …“; moderate Stochastizität (z. B. *temperature*≈0.7)<sup>[\[31\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-31)</sup>.
- **Anzahl der hypothetischen Texte**: 1–5; die Mittelung der Embeddings erhöht die Robustheit<sup>[\[32\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-32)</sup>.
- **Embedder**: (m)Contriever ohne Fine-Tuning; es können auch feinabgestimmte Encoder verwendet werden (der HyDE-Effekt bleibt bestehen)<sup>[\[33\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-33)</sup>.
- **Normalisierung der Embeddings**: L2-Norm; das innere Produkt ist äquivalent zur Kosinus-Ähnlichkeit<sup>[\[34\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-34)</sup>.
- **Hybrides Retrieval**: BM25+Vektor mit anschließendem Re-Ranking<sup>[\[35\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-35)</sup>.
- **Re-Ranker**: Cross-Encoder (BERT Re-Ranker)<sup>[\[36\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-36)</sup> oder ColBERT<sup>[\[37\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-37)</sup>.
- **Fusion** von Ergebnissen verschiedener Strategien: RRF (*k*≈60)<sup>[\[38\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-38)</sup>.

Überwachung von Qualität und Kosten

- Retrieval: nDCG@k, Recall@k, MRR; End-to-End-RAG: EM/F1 oder *Groundedness*-Metriken (RAGAS/TruLens)<sup>[\[39\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-39)[\[40\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-40)</sup>.
- Kosten/Latenz: Dominiert durch die LLM-Generierung und (falls vorhanden) das Re-Ranking; optimiert durch die Anzahl der „Hypotheticals“ und die Länge der Antwort<sup>[\[41\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-41)</sup>.

## Einschränkungen und offene Fragen

- **Halluzinationen** im hypothetischen Text: Das LLM kann sachliche Fehler einbringen; die „Verankerung“ durch den Encoder und den Korpus reduziert das Risiko, eliminiert es aber nicht vollständig<sup>[\[42\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-42)</sup>.
- **Domänen- und Sprachbeschränkungen**: Der Vorteil von HyDE verringert sich in hochspezialisierten Domänen und bei unterversorgten Sprachen<sup>[\[43\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-43)</sup>.
- **Latenz und Kosten**: Die LLM-Generierung führt zu zusätzlicher Latenz und Token-Kosten; dies ist kritisch für Online-Szenarien und lange „Hypotheticals“<sup>[\[44\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-44)</sup>.
- **Ethik und Bias**: Es sollten vorzugsweise sichere LLMs und Filtermechanismen verwendet werden<sup>[\[45\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-45)</sup>.

## Vergleichstabelle der Methoden

| Methode                   | Klasse               | Wo wird Text generiert                    | Encoder/Index        | Re-Ranker (2. Stufe)           | Typische Metriken (Beispiel)                       | Kosten/Latenz                            | Quellen                                                                                                                                                                                                               |
|---------------------------|----------------------|-------------------------------------------|----------------------|--------------------------------|----------------------------------------------------|------------------------------------------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| **HyDE**                  | Query→*hypo‑doc*     | Auf der Anfrageseite (LLM → Absatz)       | (m)Contriever; ANN   | BERT Re‑Ranker / ColBERT / RRF | DL19 nDCG@10≈61.3; DL20≈57.9; ArguAna nDCG@10≈46.6 | \+ LLM-Generierung; + Re-Ranking (opt.)  | <sup>[\[46\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-46)</sup>                                                                                                      |
| BM25                      | Lexikalisch          | —                                         | Invertierter Index   | Optional                       | Siehe Tabellen (oben)                              | Niedrig (lexikalisch)                    | <sup>[\[47\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-47)</sup>                                                                                                      |
| DPR / ANCE                | Dense (ft)           | —                                         | Bi‑Encoder; ANN      | Optional                       | DL19 nDCG@10≈62–65                                 | Mittel (ohne LLM)                        | <sup>[\[48\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-48)[\[49\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-49)</sup> |
| doc2query / docTTTTTquery | Dokumenten-Expansion | Auf der Korpusseite (vor der Indizierung) | BM25/sparse+expanded | Optional                       | Verbesserungen von BM25 auf MS MARCO               | Hohe Offline-Generierung; schnell online | <sup>[\[50\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-50)[\[51\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-51)</sup> |
| PRF (Rocchio, RLM)        | QE durch Feedback    | Anfrage (basierend auf Top-Ergebnissen)   | Beliebig             | Optional                       | Recall-Steigerung/Risiko von Drift                 | \+ zusätzlicher Retrieval-Durchlauf      | <sup>[\[52\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_note-52)</sup>                                                                                                      |

Vergleich von HyDE und verwandten Ansätzen

## Siehe auch

- BM25
- Vektorsuche
- RAG
- Pseudo-Relevanz-Feedback
- BEIR

## Weblinks

- HyDE-Repository: <a href="https://github.com/texttron/hyde" class="external text" rel="nofollow">github.com/texttron/hyde</a>.
- Dokumentation: Haystack – HyDE: <a href="https://docs.haystack.deepset.ai/docs/hypothetical-document-embeddings-hyde" class="external text" rel="nofollow">docs.haystack.deepset.ai</a>.
- Dokumentation: LangChain – HyDE Retriever: <a href="https://docs.langchain.com/oss/javascript/integrations/retrievers/hyde" class="external text" rel="nofollow">docs.langchain.com</a>.

## Literatur

- Manning, C. D.; Raghavan, P.; Schütze, H. (2008). *Introduction to Information Retrieval*. Cambridge University Press. ISBN 978‑0521865715.
- Robertson, S.; Zaragoza, H. (2009). *The Probabilistic Relevance Framework: BM25 and Beyond*. Foundations and Trends in IR, 3(4), 333–389. DOI:10.1561/1500000019.

## Einzelnachweise

1.  <span id="cite_note-1">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-1) Gao, L.; Ma, X.; Lin, J.; Callan, J. (2023). ‘‘Precise Zero‑Shot Dense Retrieval without Relevance Labels (HyDE)’’. ACL 2023. pp. 1762–1777. DOI:10.18653/v1/2023.acl-long.99. <a href="https://arxiv.org/abs/2212.10496" class="external text" rel="nofollow">arXiv:2212.10496</a></span>
2.  <span id="cite_note-2">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-2) Gao, L. et al. (2023). ACL 2023, §3.2. DOI:10.18653/v1/2023.acl-long.99.</span>
3.  <span id="cite_note-3">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-3) Rocchio, J. (1971). ‘‘Relevance Feedback in Information Retrieval’’. In: Salton, G. (ed.) *The SMART Retrieval System*. Prentice‑Hall, pp. 313–323. ISBN 978‑0138145255.</span>
4.  <span id="cite_note-4">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-4) Lavrenko, V.; Croft, W. B. (2001). ‘‘Relevance‑Based Language Models’’. SIGIR. DOI:10.1145/383952.383972.</span>
5.  <span id="cite_note-5">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-5) Izacard, G. et al. (2021/2022). ‘‘Unsupervised Dense Information Retrieval with Contrastive Learning’’. <a href="https://arxiv.org/abs/2112.09118" class="external text" rel="nofollow">arXiv:2112.09118</a>.</span>
6.  <span id="cite_note-6">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-6) Karpukhin, V. et al. (2020). ‘‘Dense Passage Retrieval for Open‑Domain QA’’. EMNLP. DOI:10.18653/v1/2020.emnlp-main.550.</span>
7.  <span id="cite_note-7">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-7) Thakur, N. et al. (2021). ‘‘BEIR: A Heterogeneous Benchmark for Zero‑shot Evaluation of Information Retrieval Models’’. NeurIPS Datasets Track. <a href="https://arxiv.org/abs/2104.08663" class="external text" rel="nofollow">arXiv:2104.08663</a>.</span>
8.  <span id="cite_note-8">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-8) Gao, L. et al. (2023). DOI:10.18653/v1/2023.acl-long.99.</span>
9.  <span id="cite_note-9">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-9) Milvus Docs. ‘‘Similarity Metrics’’ — bei L2-Normalisierung von Vektoren ist das innere Produkt äquivalent zum Kosinus. URL: <a href="https://milvus.io/docs/v2.2.x/metric.md" class="external free" rel="nofollow">https://milvus.io/docs/v2.2.x/metric.md</a></span>
10. <span id="cite_note-10">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-10) Gao, L.; Ma, X.; Lin, J.; Callan, J. (2023). ‘‘Precise Zero‑Shot Dense Retrieval without Relevance Labels (HyDE)’’. ACL 2023, §3–4. arXiv:2212.10496. DOI:10.18653/v1/2023.acl-long.99.</span>
11. <span id="cite_note-11">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-11) Izacard, G. et al. (2021/2022). ‘‘Unsupervised Dense Information Retrieval with Contrastive Learning (Contriever)’’. arXiv:2112.09118.</span>
12. <span id="cite_note-12">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-12) Gao, L. et al. (2023). Anhang (Ablation): Einfluss der Anzahl hypothetischer Texte und Generierungsparameter. arXiv:2212.10496.</span>
13. <span id="cite_note-13">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-13) Gao, L. et al. (2023). DOI:10.18653/v1/2023.acl-long.99.</span>
14. <span id="cite_note-14">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-14) Nogueira, R. et al. (2019). ‘‘Document Expansion by Query Prediction’’ (doc2query). <a href="https://arxiv.org/abs/1904.08375" class="external text" rel="nofollow">arXiv:1904.08375</a>.</span>
15. <span id="cite_note-15">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-15) Nogueira, R.; Lin, J. (2019). ‘‘From doc2query to docTTTTTquery’’ (tech report). <a href="https://cs.uwaterloo.ca/~jimmylin/publications/Nogueira_Lin_2019_docTTTTTquery-v2.pdf" class="external text" rel="nofollow">PDF</a></span>
16. <span id="cite_note-16">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-16) Rocchio, J. (1971); Lavrenko & Croft (2001), siehe oben.</span>
17. <span id="cite_note-17">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-17) Nogueira, R.; Cho, K. (2019). ‘‘Passage Re‑ranking with BERT’’. <a href="https://arxiv.org/abs/1901.04085" class="external text" rel="nofollow">arXiv:1901.04085</a>.</span>
18. <span id="cite_note-18">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-18) Khattab, O.; Zaharia, M. (2020). ‘‘ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT’’. SIGIR. DOI:10.1145/3397271.3401075; <a href="https://arxiv.org/abs/2004.12832" class="external text" rel="nofollow">arXiv:2004.12832</a>.</span>
19. <span id="cite_note-19">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-19) Cormack, G. V.; Clarke, C. L. A.; Büttcher, S. (2009). ‘‘Reciprocal Rank Fusion Outperforms Condorcet and Nearly Optimally Combines Rankings’’. SIGIR. DOI:10.1145/1571941.1572114.</span>
20. <span id="cite_note-20">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-20) Gao, L. et al. (2023). Tab. 1. DOI:10.18653/v1/2023.acl-long.99.</span>
21. <span id="cite_note-21">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-21) Izacard, G. et al. (2022); zusammenfassende Metriken – in Gao et al., 2023, Tab. 1. arXiv:2112.09118.</span>
22. <span id="cite_note-22">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-22) Gao, L. et al. (2023). Tab. 1.</span>
23. <span id="cite_note-23">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-23) Karpukhin, V. et al. (2020); zusammenfassende – in Gao et al., 2023.</span>
24. <span id="cite_note-24">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-24) Xiong, L. et al. (2021). ICLR. <a href="https://arxiv.org/abs/2007.00808" class="external text" rel="nofollow">arXiv:2007.00808</a>.</span>
25. <span id="cite_note-25">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-25) Thakur, N. et al. (2021); zusammenfassende – in Gao et al., 2023, Tab. 2. arXiv:2104.08663.</span>
26. <span id="cite_note-26">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-26) Izacard, G. et al. (2022); zusammenfassende – in Gao et al., 2023, Tab. 2.</span>
27. <span id="cite_note-27">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-27) Gao, L. et al. (2023). Tab. 2.</span>
28. <span id="cite_note-28">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-28) Gao, L. et al. (2023). Tab. 3. DOI:10.18653/v1/2023.acl-long.99.</span>
29. <span id="cite_note-29">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-29) Gao, L. et al. (2023). §4–5.</span>
30. <span id="cite_note-30">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-30) Gao, L. et al. (2023). §4.2–4.3.</span>
31. <span id="cite_note-31">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-31) Gao, L. et al. (2023). §4.1.</span>
32. <span id="cite_note-32">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-32) Haystack Docs. ‘‘Hypothetical Document Embeddings (HyDE)’’ (Engineering-Referenz). <a href="https://docs.haystack.deepset.ai/docs/hypothetical-document-embeddings-hyde" class="external text" rel="nofollow">docs.haystack.deepset.ai</a></span>
33. <span id="cite_note-33">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-33) Gao, L. et al. (2023). Tab. 6.</span>
34. <span id="cite_note-34">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-34) Milvus Docs. ‘‘Similarity Metrics’’.</span>
35. <span id="cite_note-35">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-35) Haystack × Milvus Integration (offizielle Dok.). <a href="https://haystack.deepset.ai/integrations/milvus-document-store" class="external text" rel="nofollow">haystack.deepset.ai</a></span>
36. <span id="cite_note-36">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-36) Nogueira, R.; Cho, K. (2019). arXiv:1901.04085.</span>
37. <span id="cite_note-37">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-37) Khattab, O.; Zaharia, M. (2020). DOI:10.1145/3397271.3401075.</span>
38. <span id="cite_note-38">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-38) Cormack, G. V. et al. (2009). DOI:10.1145/1571941.1572114.</span>
39. <span id="cite_note-39">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-39) Manning, C. D.; Raghavan, P.; Schütze, H. (2008). *Introduction to Information Retrieval*. Cambridge Univ. Press. ISBN 978‑0521865715.</span>
40. <span id="cite_note-40">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-40) Es, S. et al. (2023). ‘‘RAGAS: Automated Evaluation of Retrieval‑Augmented Generation’’. <a href="https://arxiv.org/abs/2309.15217" class="external text" rel="nofollow">arXiv:2309.15217</a>.</span>
41. <span id="cite_note-41">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-41) Gao, L. et al. (2023). §5.</span>
42. <span id="cite_note-42">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-42) Gao, L. et al. (2023). §3.2; §4.1. DOI:10.18653/v1/2023.acl-long.99.</span>
43. <span id="cite_note-43">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-43) Gao, L. et al. (2023). Tab. 3; §4.4.</span>
44. <span id="cite_note-44">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-44) Gao, L. et al. (2023). §4–5.</span>
45. <span id="cite_note-45">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-45) Ouyang, L. et al. (2022). ‘‘Training language models to follow instructions with human feedback (InstructGPT)’’. NeurIPS. <a href="https://arxiv.org/abs/2203.02155" class="external text" rel="nofollow">arXiv:2203.02155</a>.</span>
46. <span id="cite_note-46">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-46) Gao, L. et al. (2023). Tab. 1–2.</span>
47. <span id="cite_note-47">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-47) Robertson, S.; Zaragoza, H. (2009). ‘‘The Probabilistic Relevance Framework: BM25 and Beyond’’. Found. Trends IR. DOI:10.1561/1500000019.</span>
48. <span id="cite_note-48">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-48) Karpukhin, V. et al. (2020). DOI:10.18653/v1/2020.emnlp-main.550.</span>
49. <span id="cite_note-49">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-49) Xiong, L. et al. (2021). <a href="https://arxiv.org/abs/2007.00808" class="external text" rel="nofollow">arXiv:2007.00808</a>.</span>
50. <span id="cite_note-50">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-50) Nogueira, R. et al. (2019). arXiv:1904.08375.</span>
51. <span id="cite_note-51">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-51) Nogueira, R.; Lin, J. (2019). tech report.</span>
52. <span id="cite_note-52">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(DE)#cite_ref-52) Rocchio, J. (1971). SMART; Lavrenko & Croft (2001) SIGIR.</span>
