---
title: "Packaging & Context Handling — 封装与上下文处理"
source: "https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86"
wiki: "systems-analysis.info/int"
article: "Packaging_&_Context_Handling_—_封装与上下文处理"
language: "zh"
categories:
  - "Category:Chinese"
  - "Category:Large language models"
  - "Category:Prompt engineering"
revision_id: 5394
wiki_created_at: 2026-09-06T23:48:40Z
wiki_modified_at: 2026-09-06T23:48:40Z
downloaded_at: 2026-09-07T23:08:14Z
---

# Packaging & Context Handling — 封装与上下文处理

**Packaging & Context Handling** — 一套在检索增强生成（Retrieval‑Augmented Generation, RAG）中，用于选择、压缩、布局和提供已检索知识片段到 LLM 上下文中的技术。其目标是最大化有限令牌预算的效用，提高答案的准确性和鲁棒性，并确保来源的可追溯引用。“封装”不仅指形成片段列表，还包括它们的压缩、排序、分组以及给模型的指令，包括 *stuff*、*map‑reduce*、*refine* 和 *tree‑of‑chunks* 等策略。<sup>[\[1\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-lewis2020-1)[\[2\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-langchain-sum-2)</sup>

## 定义与动机

在 RAG 系统中，最终答案的质量不仅取决于检索，还取决于所选片段*如何*被送入提示词（prompt）。上下文限制和令牌成本迫使我们在完整性和精确性之间进行权衡：多余的片段会增加“中间迷失”（*lost‑in‑the‑middle*）的风险并延长延迟，而激进的过滤/压缩可能会删除关键证据。<sup>[\[3\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-lostmiddle-3)</sup> 在经典的 RAG 中，信息源充当外部“非参数化内存”，确保了时效性和可引用性，前提是封装能让 LLM 可靠地处理事实和引用。<sup>[\[1\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-lewis2020-1)</sup>

## 分块与内容提取

分块（*chunking*）策略决定了块（chunk）的大小/重叠、规范化和粒度（*document*→*passage*→*sentence*）。典型方法包括：

- **固定大小规则**（按字符/令牌）并设置重叠，以保持块之间的连贯性；<sup>[\[4\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-lc-splitters-4)[\[5\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-llama-splitters-5)</sup>
- **语义分块**（根据嵌入向量的相似度确定边界），以减少“意义断裂”。<sup>[\[6\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-llama-sem-6)</sup>
- **句子窗口检索**（Sentence‑window retrieval）—— 首先对句子建立索引；检索时，提取相关句子及其前后的相邻句子作为*窗口*，以恢复局部上下文。<sup>[\[7\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-haystack-sentwin-7)</sup>
- **段落级索引**（Passage‑level indexing）—— 将维基百科分割成约100词的段落，这已成为开放域问答（DPR）的标准，表明细粒度对初始检索阶段很有帮助。<sup>[\[8\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-dpr-8)</sup>
- **规范化与清理**（删除垃圾内容、页眉/页脚、统一空格），并在元数据层面跟踪来源/页面/偏移量，以实现可追溯性。<sup>[\[9\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-haystack-pre-9)</sup>
- **去重**（Deduplication）候选块（精确和*近乎重复*）：使用 shingles + MinHash/LSH 来减少重复。<sup>[\[10\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-broder1997-10)</sup>

## 多样化与选择 (MMR 等)

在构建上下文集合时，要求高相关性和低冗余性。经典的“最大边界相关性”（Maximal Marginal Relevance）函数在选择下一个片段时，会同时考虑其与查询的相似度以及与*已选片段*的最大相似度（即对重复内容进行惩罚）：

${MMR}(d_{i}) = \arg\max\limits_{d_{i} \in D \smallsetminus S}\left\lbrack \lambda \cdot {sim}(q,d_{i}) - (1 - \lambda) \cdot \max\limits_{d_{j} \in S}{sim}(d_{i},d_{j}) \right\rbrack,\ \lambda \in \lbrack 0,1\rbrack$。<sup>[\[11\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-mmr-11)</sup>

信号组合：混合检索（BM25 + dense）→ 融合（例如，倒数排序融合，Reciprocal Rank Fusion, RRF）→ 使用交叉编码器（cross‑encoder）/ColBERT 进行重排：

- **RRF**: 一种简单有效的无监督方案，用于合并不同检索器的排名结果。<sup>[\[12\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-rrf-12)</sup>
- **交叉编码器重排器**（BERT/MonoT5/现代商业 API）能显著提高 top‑k 的准确性，但会增加延迟。<sup>[\[13\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-nogueira2019-13)[\[14\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-cohere-rerank-14)</sup>
- **多向量检索器** ColBERT（*late interaction*）在大型语料库上通常作为高效的一级重排器/检索器。<sup>[\[15\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-colbert-15)</sup>
- **混合搜索**（BM25F+向量）已在工业级搜索引擎和库中实现，并带有可配置的权重/融合方法（alpha、RRF 等）。<sup>[\[16\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-weav-hybrid-16)[\[17\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-pine-hybrid-17)</sup>

## 上下文压缩

在不丢失事实的情况下减少上下文大小，对成本和延迟至关重要：

- **抽取式压缩**（提取关键句子/短语）；**生成式摘要**（改写/压缩）。经典观点参见 Nenkova & McKeown。<sup>[\[18\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-nenkova2011-18)</sup>
- **查询导向/指令导向**压缩：根据查询/任务进行总结（突出证据并删除不相关内容）。
- **提示词/上下文压缩**（Prompt/context compression）通过 LLM 过滤/剪枝令牌（例如 LLMLingua/LLMLingua‑2）来减少令牌预算，同时质量损失较小，但需要仔细验证其*忠实性*。<sup>[\[19\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-llmlingua-19)</sup>
- 压缩是*质量↔成本↔延迟*之间的权衡：激进的压缩会增加忽略细微差别/前提的风险，并降低事实归因的准确性。<sup>[\[20\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-ji2023-20)</sup>

## 封装策略 (stuff/map-reduce/refine/tree)

以下是在提示词中组织信息源的四种基本方案及其典型应用场景（另见对比表）。

**Stuff** (直接填充)  
将选定的片段（可能经过压缩）拼接在一起，然后一次性全部送入模型。这种方法简单快捷，但受上下文长度限制，且在长输入上容易出现*中间迷失*问题。<sup>[\[2\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-langchain-sum-2)</sup>

**Map‑Reduce**  
在 *map* 阶段，针对每个片段/文档进行局部回答/总结；然后在 *reduce* 阶段进行聚合（比较、投票、合并）。这种方法在处理大量信息源时扩展性好，可降低单个提示词的负担；但如果聚合方式过于简单，可能会丢失信息源之间的交叉关联。<sup>[\[2\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-langchain-sum-2)[\[21\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-fid-21)</sup>

**Refine**  
逐步改进：根据第一个片段生成初始答案，然后结合下一个片段进行迭代式*改进*（补充/修正）。当信息源的顺序很重要时，这种方法很方便；但存在“固着”于早期错误并累积偏差的风险。<sup>[\[22\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-llama-refine-22)</sup>

**Tree‑of‑chunks**  
分层压缩/摘要：先对块进行局部摘要 → 再在章节层面进行归纳 → 最后生成最终摘要。这种方法对长文档很有用；但需要确保在不同层级间准确传递源标识符，以便正确归因。<sup>[\[23\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-llama-tree-23)</sup>

| 策略               | 思路                           | 成本/延迟            | 上下文丢失风险                           | 适用场景                  | 来源                                                                                                                                                                                                                                                                                                                                                                                 |
|--------------------|--------------------------------|----------------------|------------------------------------------|---------------------------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| **Stuff**          | 将所有片段一次性放入一个提示词 | 低（在上下文限制内） | 在长输入上风险高（*lost‑in‑the‑middle*） | 内容量小，问题简单        | <sup>[\[2\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-langchain-sum-2)[\[3\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-lostmiddle-3)</sup> |
| **Map‑Reduce**     | 局部回答 → 聚合                | 中/高（多次调用）    | 中（取决于 reduce 阶段的质量）           | 信息源多，需要可扩展性    | <sup>[\[2\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-langchain-sum-2)[\[21\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-fid-21)</sup>      |
| **Refine**         | 逐步改进答案                   | 中                   | 依赖顺序，有固化错误的风险               | 当答案的顺序/演进很重要时 | <sup>[\[22\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-llama-refine-22)</sup>                                                                                                                                                                                   |
| **Tree‑of‑chunks** | 分层摘要                       | 中/高                | 在高层级会丢失细节                       | 长文档/合集               | <sup>[\[23\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-llama-tree-23)</sup>                                                                                                                                                                                     |

封装策略对比

## 信息源的排序与定位

LLM 对长上下文中部信息的利用效率较低；有用的事实最好放在开头/结尾，按主题/来源分组，并用标题和 ID 标记。考虑到*查询感知*的重要性和多样性进行重排，有助于将关键片段移到更靠前的位置。<sup>[\[3\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-lostmiddle-3)[\[13\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-nogueira2019-13)</sup>

## RAG 流水线集成 (fusion → rerank → packaging)

典型的多阶段流水线：**混合检索**（BM25 + dense）→ **融合**（RRF/加权混合）→ **重排**（Cross‑Encoder/ColBERT）→ **封装**（采用其中一种策略）→ **生成** + **引用**。混合搜索和 RRF 对不同检索器得分不一致的情况具有鲁棒性；交叉编码器则能提高输入 LLM 的精确度，从而节省令牌。<sup>[\[16\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-weav-hybrid-16)[\[12\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-rrf-12)[\[14\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-cohere-rerank-14)[\[15\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-colbert-15)</sup>

## 质量评估与消融研究

评估在检索、封装和生成等多个层面进行：

- **检索**: Recall@k、nDCG@k、MRR — 信息检索领域的标准指标。<sup>[\[24\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-manning2008-24)</sup>
- **忠实性/有据性**（Faithfulness/groundedness）: 由引文支持的陈述比例；使用自动化框架（RAGAS, TruLens）+ 手动验证归因。<sup>[\[25\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-ragas-25)[\[26\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-trulens-26)[\[27\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-rashkin2023-27)</sup>
- **端到端问答**: EM/F1/ROUGE，具体取决于任务/数据集（NQ/HotpotQA 等）。<sup>[\[1\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-lewis2020-1)</sup>
- **效率**: 延迟 p50/p95、令牌数、美元成本；比较不同封装策略和压缩级别下的*质量↔成本*。
- **消融研究**（Ablations）: 通过关闭 MMR/去重/压缩/改变顺序来衡量每个组件的贡献（截至 2025‑09‑10，RAG 研究实践建议明确记录 k 值、λ、块大小和令牌限制）。<sup>[\[28\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-rag-survey-28)</sup>

## 实践建议与清单

- **k值与多样化**: 从混合检索中获取 k=20–40 个候选片段开始；应用 MMR，λ≈0.5–0.8；对 URL/ID/文本哈希相同的重复内容进行严格惩罚。<sup>[\[11\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-mmr-11)[\[16\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-weav-hybrid-16)</sup>
- **分块**: 固定大小分块时，使用 200–400 个令牌，重叠 10–20%；对于法律技术类文档，句子/窗口方案通常更有效。<sup>[\[4\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-lc-splitters-4)[\[7\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-haystack-sentwin-7)</sup>
- **压缩**: 使用基于查询的抽取式过滤和谨慎的生成式方法；根据数据上的*忠实度*评估结果，增减 LLMLingua 类方法的使用（必须进行 A/B 验证）。<sup>[\[19\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-llmlingua-19)[\[27\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-rashkin2023-27)</sup>
- **顺序**: 将重要/高置信度的片段放在提示词的开头；按来源/主题分组，并明确标注 ID 和标题；注意*中间迷失*效应（在开头和结尾重复关键事实可能会有帮助）。<sup>[\[3\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-lostmiddle-3)</sup>
- **重排**: 如果预算允许，在封装前对 top-k（k≈50–200）使用 Cross‑Encoder/ColBERT，这可以节省生成令牌并提高准确性。<sup>[\[13\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-nogueira2019-13)[\[15\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-colbert-15)</sup>
- **回退策略**（Fallback）: (1) 事实不足 → 请求额外来源；(2) 超过令牌限制 → 从 *stuff* 切换到 *refine* 或启用压缩；(3) 置信度低/存在矛盾 → 拒绝回答，并明确列出缺失的 ID（见下文模板）。

### 封装流水线伪代码

    # 输入: 查询 q
    cands = retrieve(q, K_sparse, K_dense)          # BM25、DPR 等检索
    cands = diversify_MMR(cands, lambda=0.7)        # 多样化 (MMR)
    snips = compress(query=q, items=cands, mode="extractive|abstractive", budget=tokens)
    pkg   = package(snips, strategy="stuff|map_reduce|refine|tree")
    resp  = generate(prompt=build_prompt(q, pkg), citations=True)  # 带引用的 LLM

### 提示词模板骨架（片段）

    [用户查询]
    {q}

    [来源]
    {# 每个片段都包含 ID、标题和链接 #}
    - [{id}] {title} — {url}
    {content_snippet}

    [要求]
    1) 仅使用来源中的事实，并引用 [ID]。
    2) 如果数据不足，请说明情况并请求澄清或提供额外来源。
    3) 保持回答结构，并列出所有使用的 [ID]。

## 局限性与开放性问题

- **map-reduce/refine 中的幻觉与聚合**: 生成式摘要可能会引入新的事实；明确的归因指令和引文验证机制至关重要。<sup>[\[20\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-ji2023-20)[\[27\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-rashkin2023-27)</sup>
- **激进压缩/分层压缩导致的细节丢失**: 保持到原始来源/页面/偏移量的反向链接很重要。
- **检索器/重排器和压缩器的领域可移植性**: 需要在特定领域的语料库上进行适配/微调。<sup>[\[28\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-rag-survey-28)</sup>
- **隐私/PII 与 LLM 的记忆效应**: 在没有严格*依据*（grounding）的情况下生成内容可能导致私密字符串泄露；应使用过滤器、私有存储和拒绝策略。<sup>[\[29\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-carlini-29)[\[30\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-shokri-30)</sup>
- **可训练的“封装器”**、自适应排序/布局、用于提高*忠实度*的 RLHF/反馈循环、多语言和超长上下文是当前活跃的研究方向。<sup>[\[28\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-rag-survey-28)[\[3\]](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_note-lostmiddle-3)</sup>

## 外部链接

- LangChain: *Summarization* (stuff/map_reduce/refine). <a href="https://python.langchain.com/docs/tutorials/summarization/" class="external autonumber" rel="nofollow">[29]</a>
- LangChain: *Text splitters*. <a href="https://python.langchain.com/docs/concepts/text_splitters/" class="external autonumber" rel="nofollow">[30]</a>
- LlamaIndex: *Response Synthesizers (refine/tree)*. <a href="https://docs.llamaindex.ai/en/stable/module_guides/deploying/response_synthesizers/" class="external autonumber" rel="nofollow">[31]</a>
- LlamaIndex: *Node Parsers / SentenceSplitter / SemanticSplitter*. <a href="https://docs.llamaindex.ai/en/stable/module_guides/loading/node_parsers/modules/" class="external autonumber" rel="nofollow">[32]</a>
- Haystack: *SentenceWindowRetriever*. <a href="https://docs.haystack.deepset.ai/docs/sentencewindowretrieval" class="external autonumber" rel="nofollow">[33]</a>
- Haystack: *PreProcessors / DocumentSplitter*. <a href="https://docs.haystack.deepset.ai/docs/preprocessors" class="external autonumber" rel="nofollow">[34]</a>
- Weaviate: *Hybrid search*. <a href="https://docs.weaviate.io/weaviate/concepts/search/hybrid-search" class="external autonumber" rel="nofollow">[35]</a>
- Pinecone: *Hybrid search*. <a href="https://docs.pinecone.io/docs/hybrid-search" class="external autonumber" rel="nofollow">[36]</a>
- Cohere: *Rerank API*. <a href="https://docs.cohere.com/docs/rerank-overview" class="external autonumber" rel="nofollow">[37]</a>
- RAGAS (repo/docs). <a href="https://arxiv.org/abs/2309.15217" class="external autonumber" rel="nofollow">[38]</a> <a href="https://github.com/explodinggradients/ragas" class="external autonumber" rel="nofollow">[39]</a>
- TruLens (docs). <a href="https://www.trulens.org/trulens_eval/getting_started/evaluation/" class="external autonumber" rel="nofollow">[40]</a>

## 参考文献

- Manning, C. D., Raghavan, P., Schütze, H. (2008). *Introduction to Information Retrieval*. Cambridge University Press. ISBN 978‑0521865715.
- Nenkova, A., McKeown, K. (2011). *Automatic Summarization*. FnT IR, 5(2–3), 103–233. DOI:10.1561/1500000015.
- Lewis, P., et al. (2020). *Retrieval‑Augmented Generation for Knowledge‑Intensive NLP Tasks*. NeurIPS. arXiv:2005.11401.
- Khattab, O., Zaharia, M. (2020). *ColBERT*. SIGIR’20. DOI:10.1145/3397271.3401075.
- Izacard, G., Grave, E. (2021). *Fusion‑in‑Decoder*. EACL. arXiv:2007.01282.
- Ji, Z., et al. (2023). *Survey of Hallucination in NLG*. ACM CS. DOI:10.1145/3571730.
- Gao, S., et al. (2024). *RAG for LLM: A Survey*. arXiv:2312.10997.

## 注释

1.  <span id="cite_note-lewis2020-1">↑ <sup>[1.0](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-lewis2020_1-0)</sup> <sup>[1.1](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-lewis2020_1-1)</sup> <sup>[1.2](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-lewis2020_1-2)</sup> Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., et al. (2020). *Retrieval‑Augmented Generation for Knowledge‑Intensive NLP Tasks*. NeurIPS. arXiv:2005.11401. <a href="https://arxiv.org/abs/2005.11401" class="external autonumber" rel="nofollow">[1]</a></span>
2.  <span id="cite_note-langchain-sum-2">↑ <sup>[2.0](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-langchain-sum_2-0)</sup> <sup>[2.1](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-langchain-sum_2-1)</sup> <sup>[2.2](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-langchain-sum_2-2)</sup> <sup>[2.3](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-langchain-sum_2-3)</sup> <sup>[2.4](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-langchain-sum_2-4)</sup> LangChain Docs. *Summarization* (stuff/map_reduce/refine/map_rerank). (доступ: 2025‑09‑10). <a href="https://python.langchain.com/docs/tutorials/summarization/" class="external autonumber" rel="nofollow">[2]</a></span>
3.  <span id="cite_note-lostmiddle-3">↑ <sup>[3.0](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-lostmiddle_3-0)</sup> <sup>[3.1](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-lostmiddle_3-1)</sup> <sup>[3.2](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-lostmiddle_3-2)</sup> <sup>[3.3](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-lostmiddle_3-3)</sup> <sup>[3.4](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-lostmiddle_3-4)</sup> Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., Liang, P. (2024). *Lost in the Middle: How Language Models Use Long Contexts*. TACL. arXiv:2307.03172. <a href="https://arxiv.org/abs/2307.03172" class="external autonumber" rel="nofollow">[3]</a></span>
4.  <span id="cite_note-lc-splitters-4">↑ <sup>[4.0](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-lc-splitters_4-0)</sup> <sup>[4.1](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-lc-splitters_4-1)</sup> LangChain Docs. *Text splitters* (RecursiveCharacter/TokenTextSplitter). (доступ: 2025‑09‑10). <a href="https://python.langchain.com/docs/concepts/text_splitters/" class="external autonumber" rel="nofollow">[4]</a></span>
5.  <span id="cite_note-llama-splitters-5">[↑](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-llama-splitters_5-0) LlamaIndex Docs. *SentenceSplitter / TokenTextSplitter / SemanticSplitter*. (доступ: 2025‑09‑10). <a href="https://docs.llamaindex.ai/en/stable/api_reference/node_parsers/sentence_splitter/" class="external autonumber" rel="nofollow">[5]</a> <a href="https://docs.llamaindex.ai/en/stable/api_reference/node_parsers/token_text_splitter/" class="external autonumber" rel="nofollow">[6]</a> <a href="https://docs.llamaindex.ai/en/stable/api_reference/node_parsers/semantic_splitter/" class="external autonumber" rel="nofollow">[7]</a></span>
6.  <span id="cite_note-llama-sem-6">[↑](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-llama-sem_6-0) LlamaIndex Docs. *SemanticSplitterNodeParser*. (доступ: 2025‑09‑10). <a href="https://docs.llamaindex.ai/en/stable/api_reference/node_parsers/semantic_splitter/" class="external autonumber" rel="nofollow">[8]</a></span>
7.  <span id="cite_note-haystack-sentwin-7">↑ <sup>[7.0](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-haystack-sentwin_7-0)</sup> <sup>[7.1](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-haystack-sentwin_7-1)</sup> Haystack Docs. *SentenceWindowRetriever*. (доступ: 2025‑09‑10). <a href="https://docs.haystack.deepset.ai/docs/sentencewindowretrieval" class="external autonumber" rel="nofollow">[9]</a></span>
8.  <span id="cite_note-dpr-8">[↑](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-dpr_8-0) Karpukhin, V., Oguz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., Yih, W.‑T. (2020). *Dense Passage Retrieval for Open‑Domain Question Answering*. EMNLP. arXiv:2004.04906. <a href="https://arxiv.org/abs/2004.04906" class="external autonumber" rel="nofollow">[10]</a></span>
9.  <span id="cite_note-haystack-pre-9">[↑](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-haystack-pre_9-0) Haystack Docs. *PreProcessors / DocumentSplitter*. (доступ: 2025‑09‑10). <a href="https://docs.haystack.deepset.ai/docs/preprocessors" class="external autonumber" rel="nofollow">[11]</a> <a href="https://docs.haystack.deepset.ai/docs/documentsplitter" class="external autonumber" rel="nofollow">[12]</a></span>
10. <span id="cite_note-broder1997-10">[↑](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-broder1997_10-0) Broder, A. Z. (1997). *On the Resemblance and Containment of Documents*. Compression and Complexity of Sequences. <a href="https://www.cs.princeton.edu/courses/archive/spring13/cos598C/broder97resemblance.pdf" class="external autonumber" rel="nofollow">[13]</a></span>
11. <span id="cite_note-mmr-11">↑ <sup>[11.0](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-mmr_11-0)</sup> <sup>[11.1](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-mmr_11-1)</sup> Carbonell, J., Goldstein, J. (1998). *The Use of MMR, Diversity‑Based Reranking for Reordering Documents and Producing Summaries*. SIGIR’98, pp. 335–336. DOI:10.1145/290941.291025.</span>
12. <span id="cite_note-rrf-12">↑ <sup>[12.0](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-rrf_12-0)</sup> <sup>[12.1](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-rrf_12-1)</sup> Cormack, G. V., Clarke, C. L. A., Büttcher, S. (2009). *Reciprocal Rank Fusion outperforms Condorcet and Individual Rank Learning Methods*. SIGIR’09, pp. 758–759. DOI:10.1145/1571941.1572114. <a href="https://cormack.uwaterloo.ca/cormacksigir09-rrf.pdf" class="external autonumber" rel="nofollow">[14]</a></span>
13. <span id="cite_note-nogueira2019-13">↑ <sup>[13.0](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-nogueira2019_13-0)</sup> <sup>[13.1](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-nogueira2019_13-1)</sup> <sup>[13.2](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-nogueira2019_13-2)</sup> Nogueira, R., Cho, K. (2019). *Passage Re‑ranking with BERT*. arXiv:1901.04085. <a href="https://arxiv.org/abs/1901.04085" class="external autonumber" rel="nofollow">[15]</a></span>
14. <span id="cite_note-cohere-rerank-14">↑ <sup>[14.0](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-cohere-rerank_14-0)</sup> <sup>[14.1](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-cohere-rerank_14-1)</sup> Cohere Docs. *Rerank API overview*. (доступ: 2025‑09‑10). <a href="https://docs.cohere.com/docs/rerank-overview" class="external autonumber" rel="nofollow">[16]</a></span>
15. <span id="cite_note-colbert-15">↑ <sup>[15.0](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-colbert_15-0)</sup> <sup>[15.1](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-colbert_15-1)</sup> <sup>[15.2](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-colbert_15-2)</sup> Khattab, O., Zaharia, M. (2020). *ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT*. SIGIR’20, pp. 39–48. DOI:10.1145/3397271.3401075. arXiv:2004.12832.</span>
16. <span id="cite_note-weav-hybrid-16">↑ <sup>[16.0](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-weav-hybrid_16-0)</sup> <sup>[16.1](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-weav-hybrid_16-1)</sup> <sup>[16.2](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-weav-hybrid_16-2)</sup> Weaviate Docs. *Hybrid search (BM25F + vector)*. (доступ: 2025‑09‑10). <a href="https://docs.weaviate.io/weaviate/concepts/search/hybrid-search" class="external autonumber" rel="nofollow">[17]</a></span>
17. <span id="cite_note-pine-hybrid-17">[↑](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-pine-hybrid_17-0) Pinecone Docs. *Hybrid search*. (доступ: 2025‑09‑10). <a href="https://docs.pinecone.io/docs/hybrid-search" class="external autonumber" rel="nofollow">[18]</a></span>
18. <span id="cite_note-nenkova2011-18">[↑](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-nenkova2011_18-0) Nenkova, A., McKeown, K. (2011). *Automatic Summarization*. Foundations and Trends in Information Retrieval, 5(2–3), 103–233. DOI:10.1561/1500000015.</span>
19. <span id="cite_note-llmlingua-19">↑ <sup>[19.0](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-llmlingua_19-0)</sup> <sup>[19.1](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-llmlingua_19-1)</sup> Zhu, Y., Shao, Z., Li, M., et al. (2023). *LLMLingua: Compressing Prompts for Accelerating LLM Inference*. arXiv:2310.05736. <a href="https://arxiv.org/abs/2310.05736" class="external autonumber" rel="nofollow">[19]</a></span>
20. <span id="cite_note-ji2023-20">↑ <sup>[20.0](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-ji2023_20-0)</sup> <sup>[20.1](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-ji2023_20-1)</sup> Ji, Z., Lee, N., Frieske, R., et al. (2023). *Survey of Hallucination in Natural Language Generation*. ACM Computing Surveys, 55(12), Art.248. DOI:10.1145/3571730.</span>
21. <span id="cite_note-fid-21">↑ <sup>[21.0](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-fid_21-0)</sup> <sup>[21.1](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-fid_21-1)</sup> Izacard, G., Grave, E. (2021). *Leveraging Passage Retrieval with Generative Models for Open‑Domain QA (Fusion‑in‑Decoder)*. EACL. arXiv:2007.01282. <a href="https://arxiv.org/abs/2007.01282" class="external autonumber" rel="nofollow">[20]</a></span>
22. <span id="cite_note-llama-refine-22">↑ <sup>[22.0](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-llama-refine_22-0)</sup> <sup>[22.1](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-llama-refine_22-1)</sup> LlamaIndex Docs. *Response Synthesizers: refine*. (доступ: 2025‑09‑10). <a href="https://docs.llamaindex.ai/en/stable/module_guides/deploying/response_synthesizers/#refine" class="external autonumber" rel="nofollow">[21]</a></span>
23. <span id="cite_note-llama-tree-23">↑ <sup>[23.0](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-llama-tree_23-0)</sup> <sup>[23.1](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-llama-tree_23-1)</sup> LlamaIndex Docs. *Tree Summarize*. (доступ: 2025‑09‑10). <a href="https://docs.llamaindex.ai/en/stable/module_guides/deploying/response_synthesizers/#tree-summarize" class="external autonumber" rel="nofollow">[22]</a></span>
24. <span id="cite_note-manning2008-24">[↑](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-manning2008_24-0) Manning, C. D., Raghavan, P., Schütze, H. (2008). *Introduction to Information Retrieval*. Cambridge Univ. Press. (см. главы о nDCG/MRR). <a href="https://nlp.stanford.edu/IR-book/" class="external autonumber" rel="nofollow">[23]</a></span>
25. <span id="cite_note-ragas-25">[↑](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-ragas_25-0) Es, S., et al. (2023). *RAGAS: Automated Evaluation of Retrieval‑Augmented Generation*. arXiv:2309.15217. <a href="https://arxiv.org/abs/2309.15217" class="external autonumber" rel="nofollow">[24]</a></span>
26. <span id="cite_note-trulens-26">[↑](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-trulens_26-0) TruLens Docs. *Evaluating RAG (groundedness, relevance)*. (доступ: 2025‑09‑10). <a href="https://www.trulens.org/trulens_eval/getting_started/evaluation/" class="external autonumber" rel="nofollow">[25]</a></span>
27. <span id="cite_note-rashkin2023-27">↑ <sup>[27.0](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-rashkin2023_27-0)</sup> <sup>[27.1](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-rashkin2023_27-1)</sup> <sup>[27.2](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-rashkin2023_27-2)</sup> Rashkin, H., Nakov, P., et al. (2023). *Measuring Attribution in Natural Language Generation*. Computational Linguistics, 49(4), 1207–1261. DOI:10.1162/coli_a_00486.</span>
28. <span id="cite_note-rag-survey-28">↑ <sup>[28.0](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-rag-survey_28-0)</sup> <sup>[28.1](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-rag-survey_28-1)</sup> <sup>[28.2](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-rag-survey_28-2)</sup> Gao, S., et al. (2024). *Retrieval‑Augmented Generation for Large Language Models: A Survey*. arXiv:2312.10997. <a href="https://arxiv.org/abs/2312.10997" class="external autonumber" rel="nofollow">[26]</a></span>
29. <span id="cite_note-carlini-29">[↑](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-carlini_29-0) Carlini, N., Tramèr, F., et al. (2021). *Extracting Training Data from Large Language Models*. USENIX Security. <a href="https://www.usenix.org/system/files/sec21-carlini-extracting.pdf" class="external autonumber" rel="nofollow">[27]</a></span>
30. <span id="cite_note-shokri-30">[↑](https://systems-analysis.info/int/Packaging_%26_Context_Handling_%E2%80%94_%E5%B0%81%E8%A3%85%E4%B8%8E%E4%B8%8A%E4%B8%8B%E6%96%87%E5%A4%84%E7%90%86#cite_ref-shokri_30-0) Shokri, R., Stronati, M., Song, C., Shmatikov, V. (2017). *Membership Inference Attacks Against ML Models*. IEEE S&P. <a href="https://www.cs.cornell.edu/~shmat/shmat_oak17.pdf" class="external autonumber" rel="nofollow">[28]</a></span>
