---
title: "Hypothetical Document Embeddings (HyDE) (ES)"
source: "https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)"
wiki: "systems-analysis.info/int"
article: "Hypothetical_Document_Embeddings_(HyDE)_(ES)"
language: "es"
categories:
  - "Category:Large language models"
  - "Category:Prompt engineering"
  - "Category:Spanish"
revision_id: 3110
wiki_created_at: 2026-09-06T23:15:50Z
wiki_modified_at: 2026-09-06T23:15:50Z
downloaded_at: 2026-09-07T22:55:13Z
---

# Hypothetical Document Embeddings (HyDE) (ES)

**Hypothetical Document Expansion (HyDE)** es un método para mejorar la recuperación vectorial y la generación aumentada por recuperación (RAG), en el cual un modelo de lenguaje grande (LLM) genera un «documento hipotético» a partir de una consulta original; luego, este texto es vectorizado por un codificador, y la búsqueda se realiza entre documentos reales basándose en la proximidad al vector resultante. Este enfoque permite utilizar los «patrones de relevancia» codificados por el LLM y «anclarlos» al corpus mediante embeddings densos<sup>[\[1\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-1)</sup>.

## Definición e intuición

HyDE descompone la tarea de búsqueda en dos etapas:

\(1\) El LLM crea un «ejemplo de respuesta relevante» (*hypothetical document*) para la consulta, modelando así las características de relevancia;

\(2\) Un codificador contrastivo (p. ej., Contriever) convierte este texto en un vector, que se utiliza para recuperar documentos reales del índice. El texto generado puede contener errores fácticos, pero lo importante son los patrones temáticos y terminológicos que el codificador es capaz de capturar<sup>[\[2\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-2)</sup>.

## Historia y fuentes

La idea de expandir la búsqueda con textos sintéticos se remonta a trabajos sobre expansión de consultas y retroalimentación de pseudo-relevancia (PRF): el algoritmo de Rocchio y los modelos de lenguaje de relevancia<sup>[\[3\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-3)[\[4\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-4)</sup>. Para la recuperación densa (dense retrieval), se utilizaron codificadores entrenados de forma contrastiva (Contriever)<sup>[\[5\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-5)</sup> y Dense Passage Retrieval (DPR)<sup>[\[6\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-6)</sup>. El benchmark BEIR estandarizó la evaluación zero-shot<sup>[\[7\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-7)</sup>. En este contexto, se propuso HyDE como una forma de «incorporar» conocimiento de relevancia en el modo zero-shot a través de un LLM sin necesidad de reentrenar el codificador<sup>[\[8\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-8)</sup>.

## Método y formalización

Sea $\mathcal{D} = \{ d_{1},\ldots,d_{N}\}$ un corpus de documentos y $E:\text{text} \rightarrow {\mathbb{R}}^{n}$ un codificador de texto que proporciona las representaciones vectoriales de los documentos $\mathbf{v}_{d} = E(d)$. Para medir la similitud, se utiliza la similitud del coseno o el producto escalar; una observación importante: \*\*el producto escalar coincide con la similitud del coseno solo cuando ambos vectores tienen una norma L2 unitaria\*\* ($\|\mathbf{u}\| = \|\mathbf{v}\| = 1$)<sup>[\[9\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-9)</sup>.

HyDE redefine la representación de la consulta a través de un «documento hipotético» generado por un LLM. Formalmente:

$$
\begin{matrix}
 & \text{(1) Generación de texto hipotético:} & & {\overset{\sim}{d}\; = \; G\!\left( q;\,{inst} \right),} \\
 & \text{(2) Embedding del texto hipotético:} & & {\mathbf{v}_{h}\; = \; E(\overset{\sim}{d}),} \\
 & \text{(3) Búsqueda de vecinos más cercanos:} & & {\mathcal{R}_{k}(q)\; = \;{TopK}_{\, d \in \mathcal{D}}\; S\!\left( \mathbf{v}_{h},\mathbf{v}_{d} \right),}
\end{matrix}
$$

donde $G$ es un LLM con una instrucción $inst$ (por ejemplo: «Escribe un párrafo que responda a la pregunta...»), $S$ es una medida de similitud (coseno o producto interno con normalización), y $\mathcal{R}_{k}(q)$ es el conjunto de $k$ documentos con la máxima similitud<sup>[\[10\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-10)[\[11\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-11)</sup>.

En la práctica de la ingeniería, a menudo se generan \*\*varios\*\* textos hipotéticos y se agregan sus representaciones, lo que aumenta la robustez:

$$
{\overset{\sim}{d}}^{(j)} = G\!\left( q;\,{inst},\xi_{j} \right),\quad\mathbf{v}_{h}\; = \;\frac{1}{m}\sum\limits_{j = 1}^{m}E\!\left( {\overset{\sim}{d}}^{(j)} \right),
$$

donde $\xi_{j}$ son los parámetros estocásticos de decodificación (p. ej., temperature/top‑p). Este ensamblaje mejora el Recall con un aumento moderado de la latencia<sup>[\[12\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-12)</sup>.

### Pipeline básico de HyDE

    # 1) prompt(query) -> hypothetical_doc
    # 2) embed(hypothetical_doc) -> v_h
    # 3) retrieve(index, v_h, k) -> candidates
    # 4) (opcional) rerank(query, candidates) -> topN
    # 5) (para RAG) stuff / map-reduce / refine sobre topN

### Relación con otros métodos (QE, doc2query, PRF)

- **QE (expansión de consulta)** añade términos a la consulta; HyDE, en cambio, genera un «cuasi-documento» completo, lo que se alinea mejor con los codificadores densos<sup>[\[13\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-13)</sup>.
- **doc2query / docTTTTTquery** expanden los **documentos** con consultas sintéticas antes de la indexación<sup>[\[14\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-14)[\[15\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-15)</sup>; HyDE expande la **consulta** al momento, sin necesidad de reindexar.
- **PRF** (Rocchio, Relevance LM) actualiza el vector de la consulta basándose en los resultados principales; HyDE extrae el «patrón de relevancia» directamente del LLM y luego lo «ancla» mediante la recuperación en el corpus<sup>[\[16\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-16)</sup>.

## Integración en RAG y reranquinado

En RAG, HyDE se aplica como la primera etapa de recuperación: documento hipotético → embedding → k candidatos. A continuación, se utiliza el reranquinado: cross-encoders de la clase BERT<sup>[\[17\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-17)</sup> o la interacción tardía de ColBERT<sup>[\[18\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-18)</sup>. Para fusionar listas (p. ej., un híbrido de BM25+vector), se suele aplicar RRF (*reciprocal rank fusion*): $\operatorname{RRF}(d) = \sum\limits_{r \in \mathcal{R}}\frac{1}{k + \operatorname{rank}_{r}(d)},\qquad k \approx 60.$ El método RRF mejora de manera consistente la calidad agregada de las clasificaciones combinadas<sup>[\[19\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-19)</sup>.

## Evaluación en benchmarks (BEIR y otros)

El trabajo original evalúa HyDE en modo zero-shot en TREC DL’19/20 (búsqueda web) y en un subconjunto de colecciones de BEIR (Scifact, ArguAna, TREC‑COVID, FiQA, DBPedia, TREC‑NEWS, Climate‑FEVER). Fragmento de los resultados — *a fecha de 07-2023*:

| Método                    | DL19                   | DL20                   | Fuente                                                                                                           |
|---------------------------|------------------------|------------------------|------------------------------------------------------------------------------------------------------------------|
| BM25                      | 30.1 / 50.6 / 75.0     | 28.6 / 48.0 / 78.6     | <sup>[\[20\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-20)</sup> |
| Contriever (unsup.)       | 24.0 / 44.5 / 74.6     | 24.0 / 42.1 / 75.4     | <sup>[\[21\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-21)</sup> |
| **HyDE** (Contriever+LLM) | **41.8 / 61.3 / 88.0** | **38.2 / 57.9 / 84.4** | <sup>[\[22\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-22)</sup> |
| DPR (ft)                  | 36.5 / 62.2 / 76.9     | 41.8 / 65.3 / 81.4     | <sup>[\[23\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-23)</sup> |
| ANCE (ft)                 | 37.1 / 64.5 / 75.5     | 40.8 / 64.6 / 77.6     | <sup>[\[24\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-24)</sup> |

*TREC DL19/20 (búsqueda web)* — mAP / nDCG@10 / Recall@1k

| Método     | Scifact         | ArguAna         | TREC‑COVID      | FiQA        | DBPedia     | TREC‑NEWS       | Climate‑FEVER   | Fuente                                                                                                           |
|------------|-----------------|-----------------|-----------------|-------------|-------------|-----------------|-----------------|------------------------------------------------------------------------------------------------------------------|
| BM25       | 67.9 / 92.5     | 39.7 / 93.2     | **59.5 / 49.8** | 23.6 / 54.0 | 31.8 / 46.8 | 39.5 / 44.7     | 16.5 / 42.5     | <sup>[\[25\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-25)</sup> |
| Contriever | 64.9 / 92.6     | 37.9 / 90.1     | 27.3 / 17.2     | 24.5 / 56.2 | 29.2 / 45.3 | 34.8 / 42.3     | 15.5 / 44.1     | <sup>[\[26\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-26)</sup> |
| **HyDE**   | **69.1 / 96.4** | **46.6 / 97.9** | 59.3 / 41.4     | 27.3 / 62.1 | 36.8 / 47.2 | **44.0 / 50.9** | **22.3 / 53.0** | <sup>[\[27\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-27)</sup> |

*BEIR (selección de datasets)* — nDCG@10 / Recall@100

HyDE también mejora el MRR@100 en los conjuntos de datos multilingües de Mr.TyDi (sw/ko/ja/bn) en comparación con mContriever<sup>[\[28\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-28)</sup>.

## Recomendaciones prácticas

Cuándo aplicar HyDE

- Modos zero-shot o de transferencia (sin etiquetas de relevancia; «disimilitud» de dominio con los corpus de entrenamiento)<sup>[\[29\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-29)</sup>.
- Se requiere un aumento de Recall@k con una precisión aceptable: HyDE a menudo «descubre» regiones relevantes del espacio vectorial<sup>[\[30\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-30)</sup>.

Configuraciones típicas

- **LLM y prompt**: instrucción «Escribe un párrafo que responda a la pregunta...»; estocasticidad moderada (p. ej., *temperature*≈0.7)<sup>[\[31\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-31)</sup>.
- **Número de textos hipotéticos**: 1–5; promediar los embeddings aumenta la robustez<sup>[\[32\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-32)</sup>.
- **Embedder**: (m)Contriever sin fine-tuning; es posible usar codificadores con fine-tuning (el efecto de HyDE se mantiene)<sup>[\[33\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-33)</sup>.
- **Normalización de embeddings**: norma L2; el producto interno es equivalente al coseno<sup>[\[34\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-34)</sup>.
- **Recuperación híbrida**: BM25+vector con reranquinado posterior<sup>[\[35\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-35)</sup>.
- **Reranker**: Cross-Encoder (BERT re‑ranker)<sup>[\[36\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-36)</sup> o ColBERT<sup>[\[37\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-37)</sup>.
- **Fusión** de resultados de diferentes estrategias: RRF (*k*≈60)<sup>[\[38\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-38)</sup>.

Monitoreo de calidad/costo

- Recuperación: nDCG@k, Recall@k, MRR; RAG de extremo a extremo: EM/F1 o métricas de *groundedness* (RAGAS/TruLens)<sup>[\[39\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-39)[\[40\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-40)</sup>.
- Costo/latencia: dominado por la generación del LLM y (si existe) el reranquinado; se optimiza mediante el número de «hipótesis» y la longitud de la respuesta<sup>[\[41\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-41)</sup>.

## Limitaciones y preguntas abiertas

- **Alucinaciones** en el texto hipotético: el LLM puede introducir errores fácticos; el «anclaje» a través del codificador y el corpus reduce el riesgo, pero no lo elimina por completo<sup>[\[42\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-42)</sup>.
- **Limitaciones de dominio/idioma**: la ventaja de HyDE disminuye en dominios muy especializados y en idiomas con pocos recursos<sup>[\[43\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-43)</sup>.
- **Latencia y costo**: la generación del LLM añade latencia y costo por token; es crítico para escenarios en línea y «hipótesis» largas<sup>[\[44\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-44)</sup>.
- **Ética y sesgos**: es preferible usar LLMs seguros y aplicar filtrado<sup>[\[45\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-45)</sup>.

## Tabla comparativa de métodos

| Método                    | Clase             | Dónde se genera el texto                    | Codificador/índice   | Reranker (2ª etapa)          | Métricas típicas (ejemplo)                         | Costo/latencia                              | Fuentes                                                                                                                                                                                                               |
|---------------------------|-------------------|---------------------------------------------|----------------------|------------------------------|----------------------------------------------------|---------------------------------------------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| **HyDE**                  | Query→*hypo‑doc*  | Del lado de la consulta (LLM → párrafo)     | (m)Contriever; ANN   | BERT re‑rank / ColBERT / RRF | DL19 nDCG@10≈61.3; DL20≈57.9; ArguAna nDCG@10≈46.6 | \+ generación de LLM; + reranquinado (opc.) | <sup>[\[46\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-46)</sup>                                                                                                      |
| BM25                      | Léxico            | —                                           | Índice invertido     | Opcional                     | ver tabla (arriba)                                 | Baja (léxica)                               | <sup>[\[47\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-47)</sup>                                                                                                      |
| DPR / ANCE                | Denso (ft)        | —                                           | Bi‑encoder; ANN      | Opcional                     | DL19 nDCG@10≈62–65                                 | Media (sin LLM)                             | <sup>[\[48\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-48)[\[49\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-49)</sup> |
| doc2query / docTTTTTquery | Expansión de doc. | Del lado de la colección (antes de indexar) | BM25/sparse+expanded | Opcional                     | Mejoras sobre BM25 en MS MARCO                     | Alta generación offline; rápido online      | <sup>[\[50\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-50)[\[51\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-51)</sup> |
| PRF (Rocchio, RLM)        | QE por feedback   | Consulta (basado en resultados top)         | Cualquiera           | Opcional                     | Aumento de Recall/riesgos de deriva                | \+ pasada adicional de recuperación         | <sup>[\[52\]](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_note-52)</sup>                                                                                                      |

Comparación de HyDE y enfoques relacionados

## Véase también

- [RAG](https://systems-analysis.info/int/Retrieval-augmented_generation_(RAG)_(ES) "Retrieval-augmented generation (RAG) (ES)")

## Enlaces externos

- Repositorio de HyDE: <a href="https://github.com/texttron/hyde" class="external text" rel="nofollow">github.com/texttron/hyde</a>.
- Documentación: Haystack — HyDE: <a href="https://docs.haystack.deepset.ai/docs/hypothetical-document-embeddings-hyde" class="external text" rel="nofollow">docs.haystack.deepset.ai</a>.
- Documentación: LangChain — HyDE Retriever: <a href="https://docs.langchain.com/oss/javascript/integrations/retrievers/hyde" class="external text" rel="nofollow">docs.langchain.com</a>.

## Literatura

- Manning, C. D.; Raghavan, P.; Schütze, H. (2008). *Introduction to Information Retrieval*. Cambridge University Press. ISBN 978‑0521865715.
- Robertson, S.; Zaragoza, H. (2009). *The Probabilistic Relevance Framework: BM25 and Beyond*. Foundations and Trends in IR, 3(4), 333–389. DOI:10.1561/1500000019.

## Referencias

1.  <span id="cite_note-1">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-1) Gao, L.; Ma, X.; Lin, J.; Callan, J. (2023). ‘‘Precise Zero‑Shot Dense Retrieval without Relevance Labels (HyDE)’’. ACL 2023. pp. 1762–1777. DOI:10.18653/v1/2023.acl-long.99. <a href="https://arxiv.org/abs/2212.10496" class="external text" rel="nofollow">arXiv:2212.10496</a></span>
2.  <span id="cite_note-2">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-2) Gao, L. et al. (2023). ACL 2023, §3.2. DOI:10.18653/v1/2023.acl-long.99.</span>
3.  <span id="cite_note-3">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-3) Rocchio, J. (1971). ‘‘Relevance Feedback in Information Retrieval’’. In: Salton, G. (ed.) *The SMART Retrieval System*. Prentice‑Hall, pp. 313–323. ISBN 978‑0138145255.</span>
4.  <span id="cite_note-4">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-4) Lavrenko, V.; Croft, W. B. (2001). ‘‘Relevance‑Based Language Models’’. SIGIR. DOI:10.1145/383952.383972.</span>
5.  <span id="cite_note-5">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-5) Izacard, G. et al. (2021/2022). ‘‘Unsupervised Dense Information Retrieval with Contrastive Learning’’. <a href="https://arxiv.org/abs/2112.09118" class="external text" rel="nofollow">arXiv:2112.09118</a>.</span>
6.  <span id="cite_note-6">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-6) Karpukhin, V. et al. (2020). ‘‘Dense Passage Retrieval for Open‑Domain QA’’. EMNLP. DOI:10.18653/v1/2020.emnlp-main.550.</span>
7.  <span id="cite_note-7">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-7) Thakur, N. et al. (2021). ‘‘BEIR: A Heterogeneous Benchmark for Zero‑shot Evaluation of Information Retrieval Models’’. NeurIPS Datasets Track. <a href="https://arxiv.org/abs/2104.08663" class="external text" rel="nofollow">arXiv:2104.08663</a>.</span>
8.  <span id="cite_note-8">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-8) Gao, L. et al. (2023). DOI:10.18653/v1/2023.acl-long.99.</span>
9.  <span id="cite_note-9">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-9) Milvus Docs. ‘‘Similarity Metrics’’ — при L2‑нормализации векторов внутр. произведение эквивалентно косинусу. URL: <a href="https://milvus.io/docs/v2.2.x/metric.md" class="external free" rel="nofollow">https://milvus.io/docs/v2.2.x/metric.md</a></span>
10. <span id="cite_note-10">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-10) Gao, L.; Ma, X.; Lin, J.; Callan, J. (2023). ‘‘Precise Zero‑Shot Dense Retrieval without Relevance Labels (HyDE)’’. ACL 2023, §3–4. arXiv:2212.10496. DOI:10.18653/v1/2023.acl-long.99.</span>
11. <span id="cite_note-11">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-11) Izacard, G. et al. (2021/2022). ‘‘Unsupervised Dense Information Retrieval with Contrastive Learning (Contriever)’’. arXiv:2112.09118.</span>
12. <span id="cite_note-12">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-12) Gao, L. et al. (2023). Apéndice (ablation): influencia del número de textos hipotéticos y parámetros de generación. arXiv:2212.10496.</span>
13. <span id="cite_note-13">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-13) Gao, L. et al. (2023). DOI:10.18653/v1/2023.acl-long.99.</span>
14. <span id="cite_note-14">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-14) Nogueira, R. et al. (2019). ‘‘Document Expansion by Query Prediction’’ (doc2query). <a href="https://arxiv.org/abs/1904.08375" class="external text" rel="nofollow">arXiv:1904.08375</a>.</span>
15. <span id="cite_note-15">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-15) Nogueira, R.; Lin, J. (2019). ‘‘From doc2query to docTTTTTquery’’ (tech report). <a href="https://cs.uwaterloo.ca/~jimmylin/publications/Nogueira_Lin_2019_docTTTTTquery-v2.pdf" class="external text" rel="nofollow">PDF</a></span>
16. <span id="cite_note-16">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-16) Rocchio, J. (1971); Lavrenko & Croft (2001), см. выше.</span>
17. <span id="cite_note-17">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-17) Nogueira, R.; Cho, K. (2019). ‘‘Passage Re‑ranking with BERT’’. <a href="https://arxiv.org/abs/1901.04085" class="external text" rel="nofollow">arXiv:1901.04085</a>.</span>
18. <span id="cite_note-18">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-18) Khattab, O.; Zaharia, M. (2020). ‘‘ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT’’. SIGIR. DOI:10.1145/3397271.3401075; <a href="https://arxiv.org/abs/2004.12832" class="external text" rel="nofollow">arXiv:2004.12832</a>.</span>
19. <span id="cite_note-19">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-19) Cormack, G. V.; Clarke, C. L. A.; Büttcher, S. (2009). ‘‘Reciprocal Rank Fusion Outperforms Condorcet and Nearly Optimally Combines Rankings’’. SIGIR. DOI:10.1145/1571941.1572114.</span>
20. <span id="cite_note-20">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-20) Gao, L. et al. (2023). Tabla 1. DOI:10.18653/v1/2023.acl-long.99.</span>
21. <span id="cite_note-21">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-21) Izacard, G. et al. (2022); métricas resumidas en Gao et al., 2023, tabla 1. arXiv:2112.09118.</span>
22. <span id="cite_note-22">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-22) Gao, L. et al. (2023). Tabla 1.</span>
23. <span id="cite_note-23">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-23) Karpukhin, V. et al. (2020); resumido en Gao et al., 2023.</span>
24. <span id="cite_note-24">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-24) Xiong, L. et al. (2021). ICLR. <a href="https://arxiv.org/abs/2007.00808" class="external text" rel="nofollow">arXiv:2007.00808</a>.</span>
25. <span id="cite_note-25">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-25) Thakur, N. et al. (2021); métricas resumidas en Gao et al., 2023, tabla 2. arXiv:2104.08663.</span>
26. <span id="cite_note-26">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-26) Izacard, G. et al. (2022); métricas resumidas en Gao et al., 2023, tabla 2.</span>
27. <span id="cite_note-27">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-27) Gao, L. et al. (2023). Tabla 2.</span>
28. <span id="cite_note-28">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-28) Gao, L. et al. (2023). Tabla 3. DOI:10.18653/v1/2023.acl-long.99.</span>
29. <span id="cite_note-29">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-29) Gao, L. et al. (2023). §4–5.</span>
30. <span id="cite_note-30">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-30) Gao, L. et al. (2023). §4.2–4.3.</span>
31. <span id="cite_note-31">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-31) Gao, L. et al. (2023). §4.1.</span>
32. <span id="cite_note-32">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-32) Haystack Docs. ‘‘Hypothetical Document Embeddings (HyDE)’’ (referencia de ingeniería). <a href="https://docs.haystack.deepset.ai/docs/hypothetical-document-embeddings-hyde" class="external text" rel="nofollow">docs.haystack.deepset.ai</a></span>
33. <span id="cite_note-33">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-33) Gao, L. et al. (2023). Tabla 6.</span>
34. <span id="cite_note-34">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-34) Milvus Docs. ‘‘Similarity Metrics’’.</span>
35. <span id="cite_note-35">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-35) Haystack × Milvus Integration (documentación oficial). <a href="https://haystack.deepset.ai/integrations/milvus-document-store" class="external text" rel="nofollow">haystack.deepset.ai</a></span>
36. <span id="cite_note-36">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-36) Nogueira, R.; Cho, K. (2019). arXiv:1901.04085.</span>
37. <span id="cite_note-37">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-37) Khattab, O.; Zaharia, M. (2020). DOI:10.1145/3397271.3401075.</span>
38. <span id="cite_note-38">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-38) Cormack, G. V. et al. (2009). DOI:10.1145/1571941.1572114.</span>
39. <span id="cite_note-39">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-39) Manning, C. D.; Raghavan, P.; Schütze, H. (2008). *Introduction to Information Retrieval*. Cambridge Univ. Press. ISBN 978‑0521865715.</span>
40. <span id="cite_note-40">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-40) Es, S. et al. (2023). ‘‘RAGAS: Automated Evaluation of Retrieval‑Augmented Generation’’. <a href="https://arxiv.org/abs/2309.15217" class="external text" rel="nofollow">arXiv:2309.15217</a>.</span>
41. <span id="cite_note-41">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-41) Gao, L. et al. (2023). §5.</span>
42. <span id="cite_note-42">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-42) Gao, L. et al. (2023). §3.2; §4.1. DOI:10.18653/v1/2023.acl-long.99.</span>
43. <span id="cite_note-43">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-43) Gao, L. et al. (2023). Tabla 3; §4.4.</span>
44. <span id="cite_note-44">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-44) Gao, L. et al. (2023). §4–5.</span>
45. <span id="cite_note-45">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-45) Ouyang, L. et al. (2022). ‘‘Training language models to follow instructions with human feedback (InstructGPT)’’. NeurIPS. <a href="https://arxiv.org/abs/2203.02155" class="external text" rel="nofollow">arXiv:2203.02155</a>.</span>
46. <span id="cite_note-46">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-46) Gao, L. et al. (2023). Tablas 1–2.</span>
47. <span id="cite_note-47">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-47) Robertson, S.; Zaragoza, H. (2009). ‘‘The Probabilistic Relevance Framework: BM25 and Beyond’’. Found. Trends IR. DOI:10.1561/1500000019.</span>
48. <span id="cite_note-48">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-48) Karpukhin, V. et al. (2020). DOI:10.18653/v1/2020.emnlp-main.550.</span>
49. <span id="cite_note-49">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-49) Xiong, L. et al. (2021). <a href="https://arxiv.org/abs/2007.00808" class="external text" rel="nofollow">arXiv:2007.00808</a>.</span>
50. <span id="cite_note-50">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-50) Nogueira, R. et al. (2019). arXiv:1904.08375.</span>
51. <span id="cite_note-51">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-51) Nogueira, R.; Lin, J. (2019). tech report.</span>
52. <span id="cite_note-52">[↑](https://systems-analysis.info/int/Hypothetical_Document_Embeddings_(HyDE)_(ES)#cite_ref-52) Rocchio, J. (1971). SMART; Lavrenko & Croft (2001) SIGIR.</span>
