---
title: "Hypothetical Document Embeddings (HyDE) — 假设性文档扩展"
source: "https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95"
wiki: "systems-analysis.info/int"
article: "Hypothetical_Document_Embeddings_(HyDE)_—_假设性文档扩展"
language: "zh"
categories:
  - "Category:Chinese"
  - "Category:Large language models"
  - "Category:Prompt engineering"
revision_id: 3130
wiki_created_at: 2026-09-06T23:16:09Z
wiki_modified_at: 2026-09-06T23:16:09Z
downloaded_at: 2026-09-07T22:55:23Z
---

# Hypothetical Document Embeddings (HyDE) — 假设性文档扩展

**Hypothetical Document Expansion (HyDE)** 是一种改进向量检索和检索增强生成 (retrieval‑augmented generation, RAG) 的方法，其中大语言模型 (LLM) 根据原始查询生成一个“假设性文档”；然后，该文本由编码器进行向量化，并通过与得到的向量进行相似度比较，在真实文档中进行搜索。该方法利用了 LLM 中编码的“相关性模式”，并通过密集嵌入将其“锚定”到语料库中<sup>[\[1\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-1)</sup>。

## 定义与直觉

HyDE 将搜索任务分解为两个阶段：

\(1\) LLM 针对查询创建一个“相关答案示例”（*hypothetical document*），从而模拟相关性特征；

\(2\) 对比编码器（例如 Contriever）将该文本转换为向量，并据此从索引中检索真实文档。生成的文本可能包含事实错误，但编码器捕捉到的主题和术语模式才是关键<sup>[\[2\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-2)</sup>。

## 历史与渊源

使用合成文本扩展搜索的思想可追溯至查询扩展和伪相关反馈 (PRF) 的研究工作：例如罗奇奥算法和相关性语言模型<sup>[\[3\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-3)[\[4\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-4)</sup>。对于密集检索，则使用了对比学习训练的编码器（Contriever）<sup>[\[5\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-5)</sup> 和密集段落检索（Dense Passage Retrieval, DPR）<sup>[\[6\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-6)</sup>。BEIR 基准测试标准化了零样本评估<sup>[\[7\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-7)</sup>。在此背景下，HyDE 被提出，旨在通过 LLM 在无需微调编码器的情况下，将相关性知识“注入”到零样本场景中<sup>[\[8\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-8)</sup>。

## 方法与形式化

假设文档语料库为 $\mathcal{D} = \{ d_{1},\ldots,d_{N}\}$，文本编码器 $E:\text{text} \rightarrow {\mathbb{R}}^{n}$ 用于生成文档的向量表示 $\mathbf{v}_{d} = E(d)$。相似度度量可使用余弦相似度或点积；需要注意的是：\*\*只有当两个向量的 L2 范数均为单位长度时（$\|\mathbf{u}\| = \|\mathbf{v}\| = 1$），点积才等同于余弦相似度\*\*<sup>[\[9\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-9)</sup>。

HyDE 通过 LLM 生成的“假设性文档”来重新定义查询的表示。形式化如下：

$$
\begin{matrix}
 & \text{(1) 生成假设性文本：} & & {\overset{\sim}{d}\; = \; G\!\left( q;\,{inst} \right),} \\
 & \text{(2) 嵌入假设性文本：} & & {\mathbf{v}_{h}\; = \; E(\overset{\sim}{d}),} \\
 & \text{(3) 搜索最近邻：} & & {\mathcal{R}_{k}(q)\; = \;{TopK}_{\, d \in \mathcal{D}}\; S\!\left( \mathbf{v}_{h},\mathbf{v}_{d} \right),}
\end{matrix}
$$

其中 $G$ 是带有指令 $inst$ 的 LLM（例如：“写一段回答……问题的段落”），$S$ 是相似度度量（余弦或经过归一化的内积），$\mathcal{R}_{k}(q)$ 是相似度最高的 $k$ 个文档集合<sup>[\[10\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-10)[\[11\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-11)</sup>。

在工程实践中，通常会生成\*\*多个\*\*假设性文本并聚合其表示，以提高鲁棒性：

$$
{\overset{\sim}{d}}^{(j)} = G\!\left( q;\,{inst},\xi_{j} \right),\quad\mathbf{v}_{h}\; = \;\frac{1}{m}\sum\limits_{j = 1}^{m}E\!\left( {\overset{\sim}{d}}^{(j)} \right),
$$

其中 $\xi_{j}$ 是随机解码参数（例如 temperature/top‑p）。这种集成方法能在延迟适度增加的情况下提高召回率（Recall）<sup>[\[12\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-12)</sup>。

### HyDE 的基本流程

    # 1) 提示词(查询) -> 假设性文档
    # 2) 嵌入(假设性文档) -> v_h
    # 3) 检索(索引, v_h, k) -> 候选文档
    # 4) (可选) 重排序(查询, 候选文档) -> topN
    # 5) (用于 RAG) 对 topN 进行 stuff / map-reduce / refine 操作

### 与其他方法（QE、doc2query、PRF）的关联

- **QE (查询扩展)** 向查询中添加术语；而 HyDE 则是生成一个完整的“准文档”，这与密集编码器更契合<sup>[\[13\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-13)</sup>。
- **doc2query / docTTTTTquery** 在索引前用合成的查询来扩展**文档**<sup>[\[14\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-14)[\[15\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-15)</sup>；而 HyDE 是实时扩展**查询**，无需重新索引。
- **PRF** (罗奇奥算法, 相关性语言模型) 根据排名靠前的结果更新查询向量；而 HyDE 直接从 LLM 中提取“相关性模式”，然后通过在语料库中检索来将其“锚定”<sup>[\[16\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-16)</sup>。

## 在 RAG 中的集成与重排序

在 RAG 中，HyDE 作为检索的第一阶段：假设性文档 → 嵌入 → k 个候选文档。接下来使用重排序技术：BERT 类交叉编码器<sup>[\[17\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-17)</sup> 或 ColBERT 的后期交互<sup>[\[18\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-18)</sup>。为了融合不同检索策略的列表（例如 BM25+向量的混合检索），通常采用倒数排名融合 (*reciprocal rank fusion*, RRF)： $\operatorname{RRF}(d) = \sum\limits_{r \in \mathcal{R}}\frac{1}{k + \operatorname{rank}_{r}(d)},\qquad k \approx 60.$

    RRF 方法能稳定地提升合并后排序列表的综合质量[19]。

## 在基准测试上的评估（BEIR 等）

原始论文在零样本模式下对 HyDE 进行了评估，使用了 TREC DL’19/20（网页搜索）以及 BEIR 数据集集合的一个子集（Scifact, ArguAna, TREC‑COVID, FiQA, DBPedia, TREC‑NEWS, Climate‑FEVER）。以下为部分结果 — *截至 2023 年 7 月*：

| 方法                      | DL19                   | DL20                   | 来源                                                                                                                                                                                  |
|---------------------------|------------------------|------------------------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| BM25                      | 30.1 / 50.6 / 75.0     | 28.6 / 48.0 / 78.6     | <sup>[\[20\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-20)</sup> |
| Contriever (unsup.)       | 24.0 / 44.5 / 74.6     | 24.0 / 42.1 / 75.4     | <sup>[\[21\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-21)</sup> |
| **HyDE** (Contriever+LLM) | **41.8 / 61.3 / 88.0** | **38.2 / 57.9 / 84.4** | <sup>[\[22\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-22)</sup> |
| DPR (ft)                  | 36.5 / 62.2 / 76.9     | 41.8 / 65.3 / 81.4     | <sup>[\[23\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-23)</sup> |
| ANCE (ft)                 | 37.1 / 64.5 / 75.5     | 40.8 / 64.6 / 77.6     | <sup>[\[24\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-24)</sup> |

*TREC DL19/20 (网页搜索)* — mAP / nDCG@10 / Recall@1k

| 方法       | Scifact         | ArguAna         | TREC‑COVID      | FiQA        | DBPedia     | TREC‑NEWS       | Climate‑FEVER   | 来源                                                                                                                                                                                  |
|------------|-----------------|-----------------|-----------------|-------------|-------------|-----------------|-----------------|---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| BM25       | 67.9 / 92.5     | 39.7 / 93.2     | **59.5 / 49.8** | 23.6 / 54.0 | 31.8 / 46.8 | 39.5 / 44.7     | 16.5 / 42.5     | <sup>[\[25\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-25)</sup> |
| Contriever | 64.9 / 92.6     | 37.9 / 90.1     | 27.3 / 17.2     | 24.5 / 56.2 | 29.2 / 45.3 | 34.8 / 42.3     | 15.5 / 44.1     | <sup>[\[26\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-26)</sup> |
| **HyDE**   | **69.1 / 96.4** | **46.6 / 97.9** | 59.3 / 41.4     | 27.3 / 62.1 | 36.8 / 47.2 | **44.0 / 50.9** | **22.3 / 53.0** | <sup>[\[27\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-27)</sup> |

*BEIR (数据集选集)* — nDCG@10 / Recall@100

相对于 mContriever，HyDE 也在多语言数据集 Mr.TyDi (sw/ko/ja/bn) 上提升了 MRR@100<sup>[\[28\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-28)</sup>。

## 实践建议

何时使用 HyDE

- 零样本/迁移场景（没有相关性标签；领域与训练语料库“不相似”）<sup>[\[29\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-29)</sup>。
- 需要在可接受的精度下提高 Recall@k — HyDE 通常能“解锁”向量空间中的相关区域<sup>[\[30\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-30)</sup>。

典型配置

- **LLM 与提示词**：指令“写一段回答……问题的段落”；适度的随机性（例如 *temperature*≈0.7）<sup>[\[31\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-31)</sup>。
- **假设性文本数量**：1–5 个；对嵌入进行平均可以提高鲁棒性<sup>[\[32\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-32)</sup>。
- **嵌入器**：未经微调的 (m)Contriever；也可以使用经过微调的编码器（HyDE 的效果依然存在）<sup>[\[33\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-33)</sup>。
- **嵌入归一化**：L2 归一化；此时内积等价于余弦相似度<sup>[\[34\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-34)</sup>。
- **混合检索**：BM25+向量，之后进行重排序<sup>[\[35\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-35)</sup>。
- **重排序器**：交叉编码器 (BERT re‑ranker)<sup>[\[36\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-36)</sup> 或 ColBERT<sup>[\[37\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-37)</sup>。
- **结果融合**：RRF (*k*≈60)<sup>[\[38\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-38)</sup>。

质量/成本监控

- 检索：nDCG@k, Recall@k, MRR；端到端 RAG：EM/F1 或 *groundedness* 指标 (RAGAS/TruLens)<sup>[\[39\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-39)[\[40\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-40)</sup>。
- 成本/延迟：主要由 LLM 生成和（如果使用）重排序主导；可通过控制“假设性文档”的数量和长度进行优化<sup>[\[41\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-41)</sup>。

## 局限性与开放问题

- **假设性文本的幻觉**：LLM 可能引入事实错误；通过编码器和语料库进行“锚定”可以降低风险，但无法完全消除<sup>[\[42\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-42)</sup>。
- **领域/语言限制**：在高度专业化的领域和资源匮乏的语言上，HyDE 的优势会减弱<sup>[\[43\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-43)</sup>。
- **延迟与成本**：LLM 生成会增加延迟和 token 成本；这对于在线场景和较长的“假设性文档”至关重要<sup>[\[44\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-44)</sup>。
- **伦理与偏见**：建议使用安全的 LLM 并进行过滤<sup>[\[45\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-45)</sup>。

## 方法对比表

| 方法                      | 类别               | 文本生成位置               | 编码器/索引                | 重排序器（第二阶段）         | 典型指标（示例）                                   | 成本/延迟                     | 来源                                                                                                                                                                                                                                                                                                                                                            |
|---------------------------|--------------------|----------------------------|----------------------------|------------------------------|----------------------------------------------------|-------------------------------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| **HyDE**                  | 查询→*假设性文档*  | 在查询侧（LLM → 段落）     | (m)Contriever; ANN         | BERT re‑rank / ColBERT / RRF | DL19 nDCG@10≈61.3; DL20≈57.9; ArguAna nDCG@10≈46.6 | \+ LLM 生成；+ 重排序（可选） | <sup>[\[46\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-46)</sup>                                                                                                                                                                           |
| BM25                      | 词法               | —                          | 倒排索引                   | 可选                         | 见上表                                             | 低（词法）                    | <sup>[\[47\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-47)</sup>                                                                                                                                                                           |
| DPR / ANCE                | 密集 (ft)          | —                          | 双编码器 (Bi‑encoder); ANN | 可选                         | DL19 nDCG@10≈62–65                                 | 中（无 LLM）                  | <sup>[\[48\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-48)[\[49\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-49)</sup> |
| doc2query / docTTTTTquery | 文档扩展           | 在集合侧（索引前）         | BM25/稀疏+扩展             | 可选                         | 提升了 BM25 在 MS MARCO 上的表现                   | 高离线生成成本；快速在线检索  | <sup>[\[50\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-50)[\[51\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-51)</sup> |
| PRF (罗奇奥算法, RLM)     | 基于反馈的查询扩展 | 查询（根据排名靠前的结果） | 任何                       | 可选                         | 提高召回率/存在漂移风险                            | \+ 额外的检索步骤             | <sup>[\[52\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_note-52)</sup>                                                                                                                                                                           |

HyDE 与相关方法的比较

## 参见

- BM25
- 向量搜索
- RAG
- 伪相关反馈
- BEIR

## 外部链接

- HyDE 代码库: <a href="https://github.com/texttron/hyde" class="external text" rel="nofollow">github.com/texttron/hyde</a>.
- 文档：Haystack — HyDE: <a href="https://docs.haystack.deepset.ai/docs/hypothetical-document-embeddings-hyde" class="external text" rel="nofollow">docs.haystack.deepset.ai</a>.
- 文档：LangChain — HyDE Retriever: <a href="https://docs.langchain.com/oss/javascript/integrations/retrievers/hyde" class="external text" rel="nofollow">docs.langchain.com</a>.

## 参考文献

- Manning, C. D.; Raghavan, P.; Schütze, H. (2008). *Introduction to Information Retrieval*. Cambridge University Press. ISBN 978‑0521865715.
- Robertson, S.; Zaragoza, H. (2009). *The Probabilistic Relevance Framework: BM25 and Beyond*. Foundations and Trends in IR, 3(4), 333–389. DOI:10.1561/1500000019.

## 注释

1.  <span id="cite_note-1">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-1) Gao, L.; Ma, X.; Lin, J.; Callan, J. (2023). ‘‘Precise Zero‑Shot Dense Retrieval without Relevance Labels (HyDE)’’. ACL 2023. pp. 1762–1777. DOI:10.18653/v1/2023.acl-long.99. <a href="https://arxiv.org/abs/2212.10496" class="external text" rel="nofollow">arXiv:2212.10496</a></span>
2.  <span id="cite_note-2">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-2) Gao, L. et al. (2023). ACL 2023, §3.2. DOI:10.18653/v1/2023.acl-long.99.</span>
3.  <span id="cite_note-3">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-3) Rocchio, J. (1971). ‘‘Relevance Feedback in Information Retrieval’’. In: Salton, G. (ed.) *The SMART Retrieval System*. Prentice‑Hall, pp. 313–323. ISBN 978‑0138145255.</span>
4.  <span id="cite_note-4">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-4) Lavrenko, V.; Croft, W. B. (2001). ‘‘Relevance‑Based Language Models’’. SIGIR. DOI:10.1145/383952.383972.</span>
5.  <span id="cite_note-5">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-5) Izacard, G. et al. (2021/2022). ‘‘Unsupervised Dense Information Retrieval with Contrastive Learning’’. <a href="https://arxiv.org/abs/2112.09118" class="external text" rel="nofollow">arXiv:2112.09118</a>.</span>
6.  <span id="cite_note-6">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-6) Karpukhin, V. et al. (2020). ‘‘Dense Passage Retrieval for Open‑Domain QA’’. EMNLP. DOI:10.18653/v1/2020.emnlp-main.550.</span>
7.  <span id="cite_note-7">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-7) Thakur, N. et al. (2021). ‘‘BEIR: A Heterogeneous Benchmark for Zero‑shot Evaluation of Information Retrieval Models’’. NeurIPS Datasets Track. <a href="https://arxiv.org/abs/2104.08663" class="external text" rel="nofollow">arXiv:2104.08663</a>.</span>
8.  <span id="cite_note-8">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-8) Gao, L. et al. (2023). DOI:10.18653/v1/2023.acl-long.99.</span>
9.  <span id="cite_note-9">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-9) Milvus Docs. ‘‘Similarity Metrics’’ — при L2‑нормализации векторов внутр. произведение эквивалентно косинусу. URL: <a href="https://milvus.io/docs/v2.2.x/metric.md" class="external free" rel="nofollow">https://milvus.io/docs/v2.2.x/metric.md</a></span>
10. <span id="cite_note-10">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-10) Gao, L.; Ma, X.; Lin, J.; Callan, J. (2023). ‘‘Precise Zero‑Shot Dense Retrieval without Relevance Labels (HyDE)’’. ACL 2023, §3–4. arXiv:2212.10496. DOI:10.18653/v1/2023.acl-long.99.</span>
11. <span id="cite_note-11">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-11) Izacard, G. et al. (2021/2022). ‘‘Unsupervised Dense Information Retrieval with Contrastive Learning (Contriever)’’. arXiv:2112.09118.</span>
12. <span id="cite_note-12">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-12) Gao, L. et al. (2023). Прил. (ablation): влияние числа гипотетических текстов и параметров генерации. arXiv:2212.10496.</span>
13. <span id="cite_note-13">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-13) Gao, L. et al. (2023). DOI:10.18653/v1/2023.acl-long.99.</span>
14. <span id="cite_note-14">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-14) Nogueira, R. et al. (2019). ‘‘Document Expansion by Query Prediction’’ (doc2query). <a href="https://arxiv.org/abs/1904.08375" class="external text" rel="nofollow">arXiv:1904.08375</a>.</span>
15. <span id="cite_note-15">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-15) Nogueira, R.; Lin, J. (2019). ‘‘From doc2query to docTTTTTquery’’ (tech report). <a href="https://cs.uwaterloo.ca/~jimmylin/publications/Nogueira_Lin_2019_docTTTTTquery-v2.pdf" class="external text" rel="nofollow">PDF</a></span>
16. <span id="cite_note-16">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-16) Rocchio, J. (1971); Lavrenko & Croft (2001), см. выше.</span>
17. <span id="cite_note-17">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-17) Nogueira, R.; Cho, K. (2019). ‘‘Passage Re‑ranking with BERT’’. <a href="https://arxiv.org/abs/1901.04085" class="external text" rel="nofollow">arXiv:1901.04085</a>.</span>
18. <span id="cite_note-18">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-18) Khattab, O.; Zaharia, M. (2020). ‘‘ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT’’. SIGIR. DOI:10.1145/3397271.3401075; <a href="https://arxiv.org/abs/2004.12832" class="external text" rel="nofollow">arXiv:2004.12832</a>.</span>
19. <span id="cite_note-19">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-19) Cormack, G. V.; Clarke, C. L. A.; Büttcher, S. (2009). ‘‘Reciprocal Rank Fusion Outperforms Condorcet and Nearly Optimally Combines Rankings’’. SIGIR. DOI:10.1145/1571941.1572114.</span>
20. <span id="cite_note-20">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-20) Gao, L. et al. (2023). Табл. 1. DOI:10.18653/v1/2023.acl-long.99.</span>
21. <span id="cite_note-21">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-21) Izacard, G. et al. (2022); сводные метрики — в Gao et al., 2023, табл. 1. arXiv:2112.09118.</span>
22. <span id="cite_note-22">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-22) Gao, L. et al. (2023). Табл. 1.</span>
23. <span id="cite_note-23">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-23) Karpukhin, V. et al. (2020); сводные — в Gao et al., 2023.</span>
24. <span id="cite_note-24">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-24) Xiong, L. et al. (2021). ICLR. <a href="https://arxiv.org/abs/2007.00808" class="external text" rel="nofollow">arXiv:2007.00808</a>.</span>
25. <span id="cite_note-25">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-25) Thakur, N. et al. (2021); сводные — в Gao et al., 2023, табл. 2. arXiv:2104.08663.</span>
26. <span id="cite_note-26">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-26) Izacard, G. et al. (2022); сводные — в Gao et al., 2023, табл. 2.</span>
27. <span id="cite_note-27">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-27) Gao, L. et al. (2023). Табл. 2.</span>
28. <span id="cite_note-28">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-28) Gao, L. et al. (2023). Табл. 3. DOI:10.18653/v1/2023.acl-long.99.</span>
29. <span id="cite_note-29">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-29) Gao, L. et al. (2023). §4–5.</span>
30. <span id="cite_note-30">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-30) Gao, L. et al. (2023). §4.2–4.3.</span>
31. <span id="cite_note-31">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-31) Gao, L. et al. (2023). §4.1.</span>
32. <span id="cite_note-32">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-32) Haystack Docs. ‘‘Hypothetical Document Embeddings (HyDE)’’ (инженерная справка). <a href="https://docs.haystack.deepset.ai/docs/hypothetical-document-embeddings-hyde" class="external text" rel="nofollow">docs.haystack.deepset.ai</a></span>
33. <span id="cite_note-33">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-33) Gao, L. et al. (2023). Табл. 6.</span>
34. <span id="cite_note-34">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-34) Milvus Docs. ‘‘Similarity Metrics’’.</span>
35. <span id="cite_note-35">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-35) Haystack × Milvus Integration (официальная док.). <a href="https://haystack.deepset.ai/integrations/milvus-document-store" class="external text" rel="nofollow">haystack.deepset.ai</a></span>
36. <span id="cite_note-36">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-36) Nogueira, R.; Cho, K. (2019). arXiv:1901.04085.</span>
37. <span id="cite_note-37">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-37) Khattab, O.; Zaharia, M. (2020). DOI:10.1145/3397271.3401075.</span>
38. <span id="cite_note-38">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-38) Cormack, G. V. et al. (2009). DOI:10.1145/1571941.1572114.</span>
39. <span id="cite_note-39">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-39) Manning, C. D.; Raghavan, P.; Schütze, H. (2008). *Introduction to Information Retrieval*. Cambridge Univ. Press. ISBN 978‑0521865715.</span>
40. <span id="cite_note-40">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-40) Es, S. et al. (2023). ‘‘RAGAS: Automated Evaluation of Retrieval‑Augmented Generation’’. <a href="https://arxiv.org/abs/2309.15217" class="external text" rel="nofollow">arXiv:2309.15217</a>.</span>
41. <span id="cite_note-41">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-41) Gao, L. et al. (2023). §5.</span>
42. <span id="cite_note-42">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-42) Gao, L. et al. (2023). §3.2; §4.1. DOI:10.18653/v1/2023.acl-long.99.</span>
43. <span id="cite_note-43">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-43) Gao, L. et al. (2023). Табл. 3; §4.4.</span>
44. <span id="cite_note-44">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-44) Gao, L. et al. (2023). §4–5.</span>
45. <span id="cite_note-45">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-45) Ouyang, L. et al. (2022). ‘‘Training language models to follow instructions with human feedback (InstructGPT)’’. NeurIPS. <a href="https://arxiv.org/abs/2203.02155" class="external text" rel="nofollow">arXiv:2203.02155</a>.</span>
46. <span id="cite_note-46">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-46) Gao, L. et al. (2023). Табл. 1–2.</span>
47. <span id="cite_note-47">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-47) Robertson, S.; Zaragoza, H. (2009). ‘‘The Probabilistic Relevance Framework: BM25 and Beyond’’. Found. Trends IR. DOI:10.1561/1500000019.</span>
48. <span id="cite_note-48">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-48) Karpukhin, V. et al. (2020). DOI:10.18653/v1/2020.emnlp-main.550.</span>
49. <span id="cite_note-49">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-49) Xiong, L. et al. (2021). <a href="https://arxiv.org/abs/2007.00808" class="external text" rel="nofollow">arXiv:2007.00808</a>.</span>
50. <span id="cite_note-50">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-50) Nogueira, R. et al. (2019). arXiv:1904.08375.</span>
51. <span id="cite_note-51">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-51) Nogueira, R.; Lin, J. (2019). tech report.</span>
52. <span id="cite_note-52">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_%E2%80%94_%E5%81%87%E8%AE%BE%E6%80%A7%E6%96%87%E6%A1%A3%E6%89%A9%E5%B1%95#cite_ref-52) Rocchio, J. (1971). SMART; Lavrenko & Croft (2001) SIGIR.</span>
