---
title: "Patrones RAG"
source: "https://systems-analysis.info/int/Patrones_RAG"
wiki: "systems-analysis.info/int"
article: "Patrones_RAG"
language: "es"
categories:
  - "Category:Large language models"
  - "Category:Prompt engineering"
  - "Category:Spanish"
revision_id: 5464
wiki_created_at: 2026-09-06T23:49:40Z
wiki_modified_at: 2026-09-06T23:49:40Z
downloaded_at: 2026-09-07T23:08:38Z
---

# Patrones RAG

**Patrones RAG** (del inglés *RAG Patterns*) son un conjunto de enfoques arquitectónicos y metodológicos para construir sistemas de **Retrieval-Augmented Generation** (RAG). Estos patrones están diseñados para resolver problemas fundamentales de los modelos de lenguaje grandes (LLM), como las alucinaciones, la obsolescencia del conocimiento y la falta de especificidad de dominio, mediante la integración de los LLM con fuentes de datos externas y accesibles dinámicamente<sup>[\[1\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-lewis2020-1)</sup>. La evolución de RAG ha pasado de simples pipelines lineales a sistemas modulares y agénticos complejos<sup>[\[2\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-survey2024-2)</sup>.

## Patrones RAG principales

Con el desarrollo de la tecnología, han surgido numerosos patrones RAG, cada uno de los cuales resuelve tareas específicas y presenta sus propias compensaciones entre calidad, velocidad y costo.

- **Classic RAG (RAG clásico)** — el enfoque básico donde la consulta del usuario se vectoriza para buscar fragmentos (chunks) relevantes en una base de datos vectorial; los chunks encontrados se envían al LLM junto con la pregunta para generar una respuesta<sup>[\[1\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-lewis2020-1)</sup>.

<!-- -->

- **Multi-Query RAG (Consultas múltiples)** — el LLM genera varias variantes parafraseadas o refinadas de la consulta original; la búsqueda se realiza sobre todas las variantes y los resultados se combinan, lo que aumenta la exhaustividad (*recall*)<sup>[\[3\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-langchain-multiquery-3)</sup>.

<!-- -->

- **HyDE (Hypothetical Document Expansion)** — para superar la "brecha semántica" entre una consulta corta y documentos largos. El LLM primero genera un documento de respuesta "hipotético", y luego su embedding se utiliza para la búsqueda, lo que a menudo mejora la calidad de la recuperación<sup>[\[4\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-hyde-4)</sup>.

<!-- -->

- **Hybrid Retrieval (Búsqueda híbrida)** — una combinación de búsqueda semántica (vectorial) y léxica (BM25). Los esquemas híbridos se han convertido en el estándar para los sistemas de producción: la búsqueda vectorial cubre las coincidencias semánticas, mientras que BM25 encuentra términos exactos, ID o acrónimos; los resultados se combinan mediante fusión (fusion)<sup>[\[5\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-weaviate-hybrid-5)[\[6\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-qdrant-hybrid-6)[\[7\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-milvus-fulltext-7)</sup>.

<!-- -->

- **Re-ranking (Reclasificación)** — un proceso de dos etapas: un recuperador rápido devuelve un conjunto de candidatos (por ejemplo, los 100 mejores), luego un cross-encoder (u otro reclasificador) recalcula la relevancia y selecciona los mejores (por ejemplo, los 5 mejores) para el LLM<sup>[\[8\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-nogueira2019-8)[\[9\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-cohere-rerank-9)</sup>.

<!-- -->

- **Query Routing (Enrutamiento de consultas)** — en sistemas con múltiples fuentes de datos heterogéneas (diferentes índices, bases de datos, API), la consulta se dirige a la mejor fuente mediante un enrutador (un selector LLM o un clasificador); incluye estrategias de fallback<sup>[\[10\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-llama-router-10)</sup>.

<!-- -->

- **Agentic/Web RAG (RAG agéntico)** — el LLM actúa como un [agente](https://systems-analysis.info/int/Agente_de_IA "Agente de IA"): descompone preguntas complejas, planifica iteraciones y utiliza herramientas (búsqueda vectorial, búsqueda web) con retroalimentación. Una implementación típica es el paradigma ReAct<sup>[\[11\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-react-11)</sup>; para la recopilación orientada a la web y la citación obligatoria, véase WebGPT<sup>[\[12\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-webgpt-12)</sup>.

### Paradigmas relacionados y en desarrollo

- **GraphRAG (RAG de grafos)** — utiliza un grafo de conocimiento como fuente y mecanismo para seleccionar el contexto; la búsqueda se realiza a través de la estructura de relaciones entre entidades y el texto, mejorando la interpretabilidad y la calidad en preguntas de múltiples saltos (multi-hop)<sup>[\[13\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-graphrag-13)[\[14\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-graphrag-project-14)</sup>.
- **MM-RAG (RAG multimodal)** — trabaja con texto y fuentes visuales (escaneos, diagramas, tablas). Ejemplo: VisRAG demuestra la recuperación y generación orientada a VLM en documentos multimodales<sup>[\[15\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-visrag-15)</sup>.
- **Empaquetado y manejo del contexto** — métodos para integrar los chunks encontrados en el prompt: *Stuff*, *Map-Reduce*, *Refine*, *Tree-of-Chunks (RAPTOR)*<sup>[\[16\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-raptor-16)</sup>.

## Tabla comparativa de patrones

| Patrón               | Cuándo aplicar                                                                | Impacto en la calidad                                                                                                                                                                                                                                                                                                       | Costo / Latencia | Riesgos y limitaciones                                                         |
|----------------------|-------------------------------------------------------------------------------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|------------------|--------------------------------------------------------------------------------|
| **Classic RAG**      | PoC y Q&A simples sobre una base de datos homogénea                           | Nivel básico; depende en gran medida de los embeddings<sup>[\[1\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-lewis2020-1)</sup>                                                                                                                                                                              | Bajo             | Sensibilidad a la redacción; riesgo de contexto irrelevante                    |
| **Hybrid Retrieval** | En la mayoría de los escenarios de producción; muchos códigos, acrónimos o ID | Aumenta la exhaustividad (recall); cubre términos exactos<sup>[\[5\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-weaviate-hybrid-5)[\[6\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-qdrant-hybrid-6)[\[7\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-milvus-fulltext-7)</sup> | Bajo/Medio       | Ajuste de los pesos de fusión; dos índices                                     |
| **Re-ranking**       | Crítico cuando la alta precisión es importante                                | Aumento significativo de la precisión (precision) en el top-k<sup>[\[8\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-nogueira2019-8)[\[9\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-cohere-rerank-9)</sup>                                                                                   | Medio/Alto       | Latencia/costo adicional                                                       |
| **Multi-Query**      | Consultas cortas o multifacéticas                                             | Aumenta el recall<sup>[\[3\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-langchain-multiquery-3)</sup>                                                                                                                                                                                                        | Medio            | Paráfrasis redundantes o ruidosas                                              |
| **HyDE**             | Consultas cortas/ambiguas con una gran "brecha semántica"                     | Mejora la calidad de la recuperación *zero-shot*<sup>[\[4\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-hyde-4)</sup>                                                                                                                                                                                         | Medio            | Depende de la calidad del texto "hipotético"                                   |
| **Query Routing**    | Múltiples fuentes (base de documentos, SQL, API, web)                         | Aumenta la relevancia al elegir la fuente correcta<sup>[\[10\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-llama-router-10)</sup>                                                                                                                                                                             | Medio            | Error de enrutamiento = fallo en la búsqueda                                   |
| **Agentic/Web RAG**  | Consultas complejas, de investigación y de varios pasos                       | Resuelve tareas más allá de un pipeline lineal<sup>[\[11\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-react-11)[\[12\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-webgpt-12)</sup>                                                                                                            | Alto             | Complejidad, riesgo de bucles infinitos; se necesitan barandillas (guardrails) |

Comparación de los patrones RAG clave

## Implementación práctica y arquitectura

### Fases de implementación

1.  **Proof of Concept (PoC):** Comience con **Classic RAG** en un conjunto de datos limitado pero representativo para verificar la calidad de los embeddings y la recuperación básica<sup>[\[1\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-lewis2020-1)</sup>.
2.  **Minimum Viable Product (MVP):** Implemente **Hybrid Retrieval** y **Re-ranking** como la mejor relación "esfuerzo/efecto"<sup>[\[5\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-weaviate-hybrid-5)[\[8\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-nogueira2019-8)</sup>.
3.  **Production:** Agregue transformaciones de consulta (**HyDE**, **Multi-Query**) y, si es necesario, **Query Routing**; configure la observabilidad (registro de recuperación/reclasificación/respuestas) y pruebas A/B<sup>[\[3\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-langchain-multiquery-3)[\[10\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-llama-router-10)</sup>.

### Componentes clave

- **Chunking (Fragmentación):** Uno de los factores más críticos para la calidad. Un tamaño fijo ingenuo a menudo rompe unidades semánticas. Se recomiendan divisores orientados a la estructura (basados en el marcado) o recursivos (párrafo → oración → palabra)<sup>[\[17\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-rcsplit-17)[\[18\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-llama-hier-18)</sup>.
- **Embeddings y metadatos:** Almacene con cada chunk el document_id, la página/sección, el título y las fechas; esto es necesario para filtrar y citar correctamente las fuentes.
- **Recuperación híbrida y reclasificación:** Utilice BM25+vector con fusión (o RRF), y luego un cross-encoder para reclasificar un pequeño grupo de candidatos<sup>[\[5\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-weaviate-hybrid-5)[\[6\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-qdrant-hybrid-6)[\[8\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-nogueira2019-8)</sup>.
- **Empaquetado del contexto:** Elija *Map-Reduce*, *Refine* o *Tree-of-Chunks* para corpus extensos<sup>[\[16\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-raptor-16)[\[18\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-llama-hier-18)</sup>.

### Errores comunes (antipatrones)

- **Solo búsqueda vectorial** sin BM25 → fallos en códigos, ID y acrónimos<sup>[\[5\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-weaviate-hybrid-5)[\[7\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-milvus-fulltext-7)</sup>.
- **Chunk demasiado grande o pequeño** → pérdida de contexto o "dilución" del embedding<sup>[\[17\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-rcsplit-17)</sup>.
- **Ausencia de reclasificación en producción** → el LLM recibe un contexto ruidoso<sup>[\[8\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-nogueira2019-8)</sup>.
- **Falta de observabilidad** y seguimiento de fuentes → imposible analizar las causas de los errores (ver evaluación de RAG).

## Evaluación de calidad y métricas

La evaluación se realiza a nivel de recuperación (offline) y de extremo a extremo (generación).

### Métricas del recuperador

- **Hit Rate, Recall@k, MRR** — cobertura y posición de los documentos relevantes.
- **Context Precision & Recall** — en qué medida el contexto recuperado está libre de "basura" y cubre todo lo necesario (implementado en RAGAS)<sup>[\[19\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-ragas-19)</sup>.

### Métricas del generador (de extremo a extremo)

- **Faithfulness / Groundedness (Fidelidad / Fundamentación)** — correspondencia de la respuesta con el contexto proporcionado.
- **Answer Relevancy (Relevancia de la respuesta)** — correspondencia con la pregunta original.

Para automatizar las métricas, se utilizan frameworks de código abierto: **RAGAS**, **TruLens** (la *tríada RAG*: context relevance, groundedness, answer relevance), **DeepEval**<sup>[\[20\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-trulens-20)[\[21\]](https://systems-analysis.info/int/Patrones_RAG#cite_note-deepeval-21)</sup>.

## Véase también

- [Retrieval-Augmented Generation (RAG)](https://systems-analysis.info/int/Retrieval-augmented_generation_(RAG)_(ES) "Retrieval-augmented generation (RAG) (ES)")
- [Bases de datos vectoriales](https://systems-analysis.info/int/Base_de_datos_vectorial "Base de datos vectorial")
- [Embedding](https://systems-analysis.info/int/Embedding_(NLP)_(ES) "Embedding (NLP) (ES)")
- [Agente de IA](https://systems-analysis.info/int/Agente_de_IA "Agente de IA")
- [MM-RAG](https://systems-analysis.info/int/MM-RAG_(Multimodal_RAG)_(ES) "MM-RAG (Multimodal RAG) (ES)")

## Referencias

1.  <span id="cite_note-lewis2020-1">↑ <sup>[1.0](https://systems-analysis.info/int/Patrones_RAG#cite_ref-lewis2020_1-0)</sup> <sup>[1.1](https://systems-analysis.info/int/Patrones_RAG#cite_ref-lewis2020_1-1)</sup> <sup>[1.2](https://systems-analysis.info/int/Patrones_RAG#cite_ref-lewis2020_1-2)</sup> <sup>[1.3](https://systems-analysis.info/int/Patrones_RAG#cite_ref-lewis2020_1-3)</sup> Lewis, P., Perez, E., et al. (2020). *Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks*. NeurIPS. arXiv:2005.11401.</span>
2.  <span id="cite_note-survey2024-2">[↑](https://systems-analysis.info/int/Patrones_RAG#cite_ref-survey2024_2-0) Fan, W., Ding, Y., et al. (2024). *A Survey on RAG Meeting LLMs: Towards Retrieval-Augmented Large Language Models*. KDD. DOI:10.1145/3637528.3671470; arXiv:2405.06211.</span>
3.  <span id="cite_note-langchain-multiquery-3">↑ <sup>[3.0](https://systems-analysis.info/int/Patrones_RAG#cite_ref-langchain-multiquery_3-0)</sup> <sup>[3.1](https://systems-analysis.info/int/Patrones_RAG#cite_ref-langchain-multiquery_3-1)</sup> <sup>[3.2](https://systems-analysis.info/int/Patrones_RAG#cite_ref-langchain-multiquery_3-2)</sup> LangChain Docs. *MultiQueryRetriever*. <a href="https://python.langchain.com/docs/how_to/MultiQueryRetriever/" class="external free" rel="nofollow">https://python.langchain.com/docs/how_to/MultiQueryRetriever/</a></span>
4.  <span id="cite_note-hyde-4">↑ <sup>[4.0](https://systems-analysis.info/int/Patrones_RAG#cite_ref-hyde_4-0)</sup> <sup>[4.1](https://systems-analysis.info/int/Patrones_RAG#cite_ref-hyde_4-1)</sup> Gao, L., Ma, X., Lin, J., Callan, J. (2023). *Precise Zero‑Shot Dense Retrieval without Relevance Labels*. ACL 2023. arXiv:2212.10496; ACL Anthology: 2023.acl‑long.99.</span>
5.  <span id="cite_note-weaviate-hybrid-5">↑ <sup>[5.0](https://systems-analysis.info/int/Patrones_RAG#cite_ref-weaviate-hybrid_5-0)</sup> <sup>[5.1](https://systems-analysis.info/int/Patrones_RAG#cite_ref-weaviate-hybrid_5-1)</sup> <sup>[5.2](https://systems-analysis.info/int/Patrones_RAG#cite_ref-weaviate-hybrid_5-2)</sup> <sup>[5.3](https://systems-analysis.info/int/Patrones_RAG#cite_ref-weaviate-hybrid_5-3)</sup> <sup>[5.4](https://systems-analysis.info/int/Patrones_RAG#cite_ref-weaviate-hybrid_5-4)</sup> Weaviate Docs. *Hybrid search (BM25+Vector)*. <a href="https://docs.weaviate.io/weaviate/concepts/search/hybrid-search" class="external free" rel="nofollow">https://docs.weaviate.io/weaviate/concepts/search/hybrid-search</a></span>
6.  <span id="cite_note-qdrant-hybrid-6">↑ <sup>[6.0](https://systems-analysis.info/int/Patrones_RAG#cite_ref-qdrant-hybrid_6-0)</sup> <sup>[6.1](https://systems-analysis.info/int/Patrones_RAG#cite_ref-qdrant-hybrid_6-1)</sup> <sup>[6.2](https://systems-analysis.info/int/Patrones_RAG#cite_ref-qdrant-hybrid_6-2)</sup> Qdrant Docs. *Hybrid Queries*. <a href="https://qdrant.tech/documentation/concepts/hybrid-queries/" class="external free" rel="nofollow">https://qdrant.tech/documentation/concepts/hybrid-queries/</a></span>
7.  <span id="cite_note-milvus-fulltext-7">↑ <sup>[7.0](https://systems-analysis.info/int/Patrones_RAG#cite_ref-milvus-fulltext_7-0)</sup> <sup>[7.1](https://systems-analysis.info/int/Patrones_RAG#cite_ref-milvus-fulltext_7-1)</sup> <sup>[7.2](https://systems-analysis.info/int/Patrones_RAG#cite_ref-milvus-fulltext_7-2)</sup> Milvus Docs. *Full‑Text Search* y *Hybrid Search*. <a href="https://milvus.io/docs/full-text-search.md" class="external free" rel="nofollow">https://milvus.io/docs/full-text-search.md</a>; <a href="https://milvus.io/docs/hybrid_search_with_milvus.md" class="external free" rel="nofollow">https://milvus.io/docs/hybrid_search_with_milvus.md</a></span>
8.  <span id="cite_note-nogueira2019-8">↑ <sup>[8.0](https://systems-analysis.info/int/Patrones_RAG#cite_ref-nogueira2019_8-0)</sup> <sup>[8.1](https://systems-analysis.info/int/Patrones_RAG#cite_ref-nogueira2019_8-1)</sup> <sup>[8.2](https://systems-analysis.info/int/Patrones_RAG#cite_ref-nogueira2019_8-2)</sup> <sup>[8.3](https://systems-analysis.info/int/Patrones_RAG#cite_ref-nogueira2019_8-3)</sup> <sup>[8.4](https://systems-analysis.info/int/Patrones_RAG#cite_ref-nogueira2019_8-4)</sup> Nogueira, R., Cho, K. (2019). *Passage Re‑ranking with BERT*. arXiv:1901.04085.</span>
9.  <span id="cite_note-cohere-rerank-9">↑ <sup>[9.0](https://systems-analysis.info/int/Patrones_RAG#cite_ref-cohere-rerank_9-0)</sup> <sup>[9.1](https://systems-analysis.info/int/Patrones_RAG#cite_ref-cohere-rerank_9-1)</sup> Cohere Docs. *Rerank — best practices*. <a href="https://docs.cohere.com/docs/reranking-best-practices" class="external free" rel="nofollow">https://docs.cohere.com/docs/reranking-best-practices</a></span>
10. <span id="cite_note-llama-router-10">↑ <sup>[10.0](https://systems-analysis.info/int/Patrones_RAG#cite_ref-llama-router_10-0)</sup> <sup>[10.1](https://systems-analysis.info/int/Patrones_RAG#cite_ref-llama-router_10-1)</sup> <sup>[10.2](https://systems-analysis.info/int/Patrones_RAG#cite_ref-llama-router_10-2)</sup> LlamaIndex Docs. *Routing (query routers/selectors)*. <a href="https://docs.llamaindex.ai/en/stable/module_guides/querying/router/" class="external free" rel="nofollow">https://docs.llamaindex.ai/en/stable/module_guides/querying/router/</a></span>
11. <span id="cite_note-react-11">↑ <sup>[11.0](https://systems-analysis.info/int/Patrones_RAG#cite_ref-react_11-0)</sup> <sup>[11.1](https://systems-analysis.info/int/Patrones_RAG#cite_ref-react_11-1)</sup> Yao, S., et al. (2023). *ReAct: Synergizing Reasoning and Acting in Language Models*. ICLR 2023. arXiv:2210.03629.</span>
12. <span id="cite_note-webgpt-12">↑ <sup>[12.0](https://systems-analysis.info/int/Patrones_RAG#cite_ref-webgpt_12-0)</sup> <sup>[12.1](https://systems-analysis.info/int/Patrones_RAG#cite_ref-webgpt_12-1)</sup> Nakano, R., et al. (2021). *WebGPT: Browser‑assisted question‑answering with human feedback*. arXiv:2112.09332.</span>
13. <span id="cite_note-graphrag-13">[↑](https://systems-analysis.info/int/Patrones_RAG#cite_ref-graphrag_13-0) Microsoft Research Blog. *GraphRAG: Unlocking LLM discovery on narrative private data*. 2024. <a href="https://www.microsoft.com/en-us/research/blog/graphrag-unlocking-llm-discovery-on-narrative-private-data/" class="external free" rel="nofollow">https://www.microsoft.com/en-us/research/blog/graphrag-unlocking-llm-discovery-on-narrative-private-data/</a></span>
14. <span id="cite_note-graphrag-project-14">[↑](https://systems-analysis.info/int/Patrones_RAG#cite_ref-graphrag-project_14-0) Microsoft Research. *Project GraphRAG*. <a href="https://www.microsoft.com/en-us/research/project/graphrag/" class="external free" rel="nofollow">https://www.microsoft.com/en-us/research/project/graphrag/</a></span>
15. <span id="cite_note-visrag-15">[↑](https://systems-analysis.info/int/Patrones_RAG#cite_ref-visrag_15-0) Yu, S., et al. (2024). *VisRAG: Vision‑based Retrieval‑augmented Generation on Multi‑modality Documents*. arXiv:2410.10594; OpenReview: zG459X3Xge.</span>
16. <span id="cite_note-raptor-16">↑ <sup>[16.0](https://systems-analysis.info/int/Patrones_RAG#cite_ref-raptor_16-0)</sup> <sup>[16.1](https://systems-analysis.info/int/Patrones_RAG#cite_ref-raptor_16-1)</sup> Sarthi, P., et al. (2024). *RAPTOR: Recursive Abstractive Processing for Tree‑Organized Retrieval*. arXiv:2401.18059.</span>
17. <span id="cite_note-rcsplit-17">↑ <sup>[17.0](https://systems-analysis.info/int/Patrones_RAG#cite_ref-rcsplit_17-0)</sup> <sup>[17.1](https://systems-analysis.info/int/Patrones_RAG#cite_ref-rcsplit_17-1)</sup> LangChain Docs. *RecursiveCharacterTextSplitter*. <a href="https://python.langchain.com/docs/how_to/recursive_text_splitter/" class="external free" rel="nofollow">https://python.langchain.com/docs/how_to/recursive_text_splitter/</a></span>
18. <span id="cite_note-llama-hier-18">↑ <sup>[18.0](https://systems-analysis.info/int/Patrones_RAG#cite_ref-llama-hier_18-0)</sup> <sup>[18.1](https://systems-analysis.info/int/Patrones_RAG#cite_ref-llama-hier_18-1)</sup> LlamaIndex Docs. *HierarchicalNodeParser* y *Tree Summarization*. <a href="https://docs.llamaindex.ai/en/stable/api/llama_index.core.node_parser.HierarchicalNodeParser.html" class="external free" rel="nofollow">https://docs.llamaindex.ai/en/stable/api/llama_index.core.node_parser.HierarchicalNodeParser.html</a>; <a href="https://docs.llamaindex.ai/en/stable/examples/low_level/response_synthesis/" class="external free" rel="nofollow">https://docs.llamaindex.ai/en/stable/examples/low_level/response_synthesis/</a></span>
19. <span id="cite_note-ragas-19">[↑](https://systems-analysis.info/int/Patrones_RAG#cite_ref-ragas_19-0) Es, S., et al. (2024). *RAGAs: Automated Evaluation of Retrieval Augmented Generation*. EACL (Demo). <a href="https://aclanthology.org/2024.eacl-demo.16/" class="external free" rel="nofollow">https://aclanthology.org/2024.eacl-demo.16/</a></span>
20. <span id="cite_note-trulens-20">[↑](https://systems-analysis.info/int/Patrones_RAG#cite_ref-trulens_20-0) TruLens Docs. *RAG Triad*. <a href="https://www.trulens.org/getting_started/core_concepts/rag_triad/" class="external free" rel="nofollow">https://www.trulens.org/getting_started/core_concepts/rag_triad/</a></span>
21. <span id="cite_note-deepeval-21">[↑](https://systems-analysis.info/int/Patrones_RAG#cite_ref-deepeval_21-0) DeepEval (GitHub). <a href="https://github.com/confident-ai/deepeval" class="external free" rel="nofollow">https://github.com/confident-ai/deepeval</a></span>

## Referencias

- Lewis, P., Perez, E., et al. (2020). *Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks*. NeurIPS. <a href="https://arxiv.org/abs/2005.11401" class="external text" rel="nofollow">arXiv:2005.11401</a>.
- Fan, W., Ding, Y., et al. (2024). *A Survey on RAG Meeting LLMs: Towards Retrieval-Augmented Large Language Models*. KDD. DOI:10.1145/3637528.3671470; <a href="https://arxiv.org/abs/2405.06211" class="external text" rel="nofollow">arXiv:2405.06211</a>.
- Gao, L., Ma, X., Lin, J., Callan, J. (2023). *Precise Zero‑Shot Dense Retrieval without Relevance Labels (HyDE)*. ACL 2023. <a href="https://aclanthology.org/2023.acl-long.99/" class="external text" rel="nofollow">ACL Anthology</a>; <a href="https://arxiv.org/abs/2212.10496" class="external text" rel="nofollow">arXiv:2212.10496</a>.
- Nogueira, R., Cho, K. (2019). *Passage Re‑ranking with BERT*. <a href="https://arxiv.org/abs/1901.04085" class="external text" rel="nofollow">arXiv:1901.04085</a>.
- Weaviate Docs. *Hybrid search (BM25+Vector)*. <a href="https://docs.weaviate.io/weaviate/concepts/search/hybrid-search" class="external autonumber" rel="nofollow">[1]</a>.
- Qdrant Docs. *Hybrid Queries*. <a href="https://qdrant.tech/documentation/concepts/hybrid-queries/" class="external autonumber" rel="nofollow">[2]</a>.
- Milvus Docs. *Full‑Text Search* / *Hybrid Search*. <a href="https://milvus.io/docs/full-text-search.md" class="external autonumber" rel="nofollow">[3]</a> / <a href="https://milvus.io/docs/hybrid_search_with_milvus.md" class="external autonumber" rel="nofollow">[4]</a>.
- LangChain Docs. *MultiQueryRetriever*. <a href="https://python.langchain.com/docs/how_to/MultiQueryRetriever/" class="external autonumber" rel="nofollow">[5]</a>.
- Cohere Docs. *Rerank — best practices*. <a href="https://docs.cohere.com/docs/reranking-best-practices" class="external autonumber" rel="nofollow">[6]</a>.
- LlamaIndex Docs. *Routing (query routers/selectors)*. <a href="https://docs.llamaindex.ai/en/stable/module_guides/querying/router/" class="external autonumber" rel="nofollow">[7]</a>.
- Yao, S., et al. (2023). *ReAct: Synergizing Reasoning and Acting in Language Models*. ICLR. <a href="https://arxiv.org/abs/2210.03629" class="external text" rel="nofollow">arXiv:2210.03629</a>.
- Nakano, R., et al. (2021). *WebGPT: Browser‑assisted question‑answering with human feedback*. <a href="https://arxiv.org/abs/2112.09332" class="external text" rel="nofollow">arXiv:2112.09332</a>.
- Microsoft Research Blog. *GraphRAG: Unlocking LLM discovery on narrative private data*. (2024). <a href="https://www.microsoft.com/en-us/research/blog/graphrag-unlocking-llm-discovery-on-narrative-private-data/" class="external autonumber" rel="nofollow">[8]</a>.
- Microsoft Research. *Project GraphRAG*. (2024). <a href="https://www.microsoft.com/en-us/research/project/graphrag/" class="external autonumber" rel="nofollow">[9]</a>.
- Yu, S., et al. (2024). *VisRAG: Vision‑based Retrieval‑augmented Generation on Multi‑modality Documents*. <a href="https://arxiv.org/abs/2410.10594" class="external text" rel="nofollow">arXiv:2410.10594</a>.
- Sarthi, P., et al. (2024). *RAPTOR: Recursive Abstractive Processing for Tree‑Organized Retrieval*. <a href="https://arxiv.org/abs/2401.18059" class="external text" rel="nofollow">arXiv:2401.18059</a>.
- Es, S., et al. (2024). *RAGAs: Automated Evaluation of Retrieval Augmented Generation*. EACL (Demo). <a href="https://aclanthology.org/2024.eacl-demo.16/" class="external autonumber" rel="nofollow">[10]</a>.
- TruLens Docs. *RAG Triad*. <a href="https://www.trulens.org/getting_started/core_concepts/rag_triad/" class="external autonumber" rel="nofollow">[11]</a>.
- DeepEval (GitHub). *The LLM Evaluation Framework*. <a href="https://github.com/confident-ai/deepeval" class="external autonumber" rel="nofollow">[12]</a>.
- LangChain Docs. *RecursiveCharacterTextSplitter*. <a href="https://python.langchain.com/docs/how_to/recursive_text_splitter/" class="external autonumber" rel="nofollow">[13]</a>.
- LlamaIndex Docs. *HierarchicalNodeParser*; *Response Synthesis (Tree/Refine)*. <a href="https://docs.llamaindex.ai/en/stable/api/llama_index.core.node_parser.HierarchicalNodeParser.html" class="external autonumber" rel="nofollow">[14]</a>; <a href="https://docs.llamaindex.ai/en/stable/examples/low_level/response_synthesis/" class="external autonumber" rel="nofollow">[15]</a>.
