# S08 Context Compact 代码讲解 这一节只讲相对 S07 新增的内容: **上下文会越来越长,所以 Agent 在调用模型之前,要先把旧消息和大工具结果变小。** --- ## 1. 本节新增内容 - `CONTEXT_LIMIT`:触发自动摘要压缩的上下文大小阈值。 - `KEEP_RECENT`:保留最近几个完整 `tool_result`。 - `PERSIST_THRESHOLD`:单个工具输出超过多大时写入磁盘。 - `snip_compact()`:裁剪中间历史。 - `micro_compact()`:把旧工具结果换成占位符。 - `tool_result_budget()`:处理最后一轮过大的工具输出。 - `compact_history()`:调用 LLM 总结完整历史。 - `reactive_compact()`:API 报上下文过长时兜底压缩。 - `compact` 工具:让模型主动请求压缩。 --- ## 2. Messages 数据结构 `messages` 是: ```python messages: list[dict] ``` 一段带工具调用的历史通常长这样: ```python messages = [ {"role": "user", "content": "请读取 README.md"}, {"role": "assistant", "content": [tool_use_block]}, {"role": "user", "content": [tool_result_block]}, ] ``` `tool_result_block` 是字典: ```python { "type": "tool_result", "tool_use_id": "toolu_01xxx", "content": "工具返回的大段文本" } ``` --- ## 3. 四层压缩顺序 在 `agent_loop()` 里,压缩发生在调用模型之前: ```python messages[:] = tool_result_budget(messages) messages[:] = snip_compact(messages) messages[:] = micro_compact(messages) if estimate_size(messages) > CONTEXT_LIMIT: messages[:] = compact_history(messages) ``` 执行顺序是: ```text L3 大工具结果预算 -> L1 裁剪中间历史 -> L2 压缩旧工具结果 -> L4 LLM 摘要 ``` 核心原则: ```text 先用便宜的办法压缩,最后才调用 LLM 做摘要。 ``` --- ## 4. `tool_result_budget(messages)` 目标: **如果最后一轮工具结果总大小超过 `max_bytes`,就优先把最大的工具输出写入磁盘。** 关键代码: ```python last = messages[-1] if messages else None ``` 等价于: ```python if messages: last = messages[-1] else: last = None ``` 这里的 `last` 预期是: ```python last: dict = { "role": "user", "content": [tool_result_block, tool_result_block] } ``` 筛选工具结果: ```python blocks = [ (i, b) for i, b in enumerate(last["content"]) if isinstance(b, dict) and b.get("type") == "tool_result" ] ``` `blocks` 的类型可以理解成: ```python blocks: list[tuple[int, dict]] ``` 例如: ```python [ (0, {"type": "tool_result", "tool_use_id": "toolu_01", "content": "..."}), (1, {"type": "tool_result", "tool_use_id": "toolu_02", "content": "..."}) ] ``` 计算总大小: ```python total = sum(len(str(b.get("content", ""))) for _, b in blocks) ``` `_` 表示这个值不用。这里不关心下标,只关心 `b` 这个工具结果字典。 按输出长度排序: ```python ranked = sorted( blocks, key=lambda p: len(str(p[1].get("content", ""))), reverse=True ) ``` `p` 是 `(i, b)`,所以: ```python p[0] # 下标 p[1] # tool_result 字典 ``` 排序后,最大的工具输出排在前面,优先被写入磁盘。 --- ## 5. `persist_large_output(tool_use_id, output)` 压缩前: ```python { "type": "tool_result", "tool_use_id": "toolu_01", "content": "非常非常长的输出......" } ``` 压缩后: ```python { "type": "tool_result", "tool_use_id": "toolu_01", "content": "文件路径 + Preview" } ``` 完整输出在磁盘: ```text .task_outputs/tool-results/toolu_01.txt ``` 上下文里只保留路径和前 2000 字符预览。 --- ## 6. `snip_compact(messages)` 作用:消息数量超过阈值时,裁掉中间历史,只保留: ```text 开头几条消息 + 一个裁剪标记 + 最近几条消息 ``` 这样模型还能知道“最开始的任务是什么”,也能看到“最近正在做什么”,中间太长的历史就先省掉。 源码: ```python def snip_compact(messages, max_messages=50): if len(messages) <= max_messages: return messages keep_head, keep_tail = 3, max_messages - 3 head_end, tail_start = keep_head, len(messages) - keep_tail if head_end > 0 and _message_has_tool_use(messages[head_end - 1]): while head_end < len(messages) and _is_tool_result_message(messages[head_end]): head_end += 1 if (tail_start > 0 and tail_start < len(messages) and _is_tool_result_message(messages[tail_start]) and _message_has_tool_use(messages[tail_start - 1])): tail_start -= 1 if head_end >= tail_start: return messages snipped = tail_start - head_end return messages[:head_end] + [{"role": "user", "content": f"[snipped {snipped} messages]"}] + messages[tail_start:] ``` --- ### 6.1 最简单的裁剪例子 为了方便理解,先把 `max_messages` 想成 8。 假设现在有 12 条消息: ```python messages = [ "msg0", "msg1", "msg2", "msg3", "msg4", "msg5", "msg6", "msg7", "msg8", "msg9", "msg10", "msg11" ] ``` 代码第一步: ```python keep_head, keep_tail = 3, max_messages - 3 ``` 如果 `max_messages = 8`: ```python keep_head = 3 keep_tail = 8 - 3 = 5 ``` 意思是: ```text 保留开头 3 条 保留结尾 5 条 ``` 再看这行: ```python head_end, tail_start = keep_head, len(messages) - keep_tail ``` 代入数字: ```python head_end = 3 tail_start = 12 - 5 = 7 ``` 这两个变量可以这样理解: ```python head_end = 开头保留到哪里结束 tail_start = 结尾从哪里开始保留 ``` 对应到列表下标: ```python messages[:head_end] ``` 就是: ```python messages[:3] = ["msg0", "msg1", "msg2"] ``` 而: ```python messages[tail_start:] ``` 就是: ```python messages[7:] = ["msg7", "msg8", "msg9", "msg10", "msg11"] ``` 中间被裁掉的是: ```python messages[3:7] = ["msg3", "msg4", "msg5", "msg6"] ``` 所以: ```python snipped = tail_start - head_end ``` 就是: ```python snipped = 7 - 3 = 4 ``` 最终结果变成: ```python [ "msg0", "msg1", "msg2", {"role": "user", "content": "[snipped 4 messages]"}, "msg7", "msg8", "msg9", "msg10", "msg11" ] ``` 这就是 `snip_compact()` 的基本逻辑。 --- ### 6.2 为什么变量叫 `head_end` 和 `tail_start` 这两个名字是从“切片边界”来的。 ```python messages[:head_end] ``` 表示保留头部。 ```python messages[tail_start:] ``` 表示保留尾部。 中间要裁掉的区间是: ```python messages[head_end:tail_start] ``` 所以: ```python head_end ``` 是头部保留区的结束位置。 ```python tail_start ``` 是尾部保留区的开始位置。 --- ### 6.3 为什么不能直接裁剪 普通聊天消息可以直接裁。 但工具调用消息有配对关系: ```python assistant: tool_use user: tool_result ``` 例子: ```python messages = [ {"role": "user", "content": "读取文件"}, {"role": "assistant", "content": [tool_use_read_file]}, {"role": "user", "content": [tool_result_read_file]}, ] ``` 这里第 1 条和第 2 条是一组: ```text 模型说:我要调用工具 程序说:这是工具结果 ``` 如果裁剪后只剩: ```python {"role": "assistant", "content": [tool_use_read_file]} ``` 但对应的 `tool_result` 被裁掉了,消息结构就不完整。 所以 `snip_compact()` 里有一段逻辑,专门避免把这组消息从中间切断。 --- ### 6.4 第一段边界保护:头部结尾不能只留下 tool_use 源码: ```python if head_end > 0 and _message_has_tool_use(messages[head_end - 1]): while head_end < len(messages) and _is_tool_result_message(messages[head_end]): head_end += 1 ``` 假设: ```python head_end = 3 ``` 原本保留: ```python messages[:3] ``` 也就是保留下标: ```text 0, 1, 2 ``` 现在检查: ```python messages[head_end - 1] ``` 就是: ```python messages[2] ``` 如果 `messages[2]` 是一条 `assistant tool_use`,说明头部最后一条消息是: ```text 模型要求调用工具 ``` 那下一条很可能就是对应的工具结果: ```python messages[3] ``` 所以代码会继续看: ```python while head_end < len(messages) and _is_tool_result_message(messages[head_end]): head_end += 1 ``` 变量变化示例: ```python head_end = 3 messages[3] 是 tool_result head_end += 1 head_end = 4 ``` 这样保留头部就从: ```python messages[:3] ``` 变成: ```python messages[:4] ``` 也就是把对应的 `tool_result` 也保留下来。 --- ### 6.5 第二段边界保护:尾部开头不能只留下 tool_result 源码: ```python if (tail_start > 0 and tail_start < len(messages) and _is_tool_result_message(messages[tail_start]) and _message_has_tool_use(messages[tail_start - 1])): tail_start -= 1 ``` 假设: ```python tail_start = 7 ``` 原本保留尾部: ```python messages[7:] ``` 也就是保留下标: ```text 7, 8, 9, 10, 11 ``` 现在检查: ```python messages[tail_start] ``` 就是: ```python messages[7] ``` 如果 `messages[7]` 是 `tool_result`,说明尾部第一条就是工具结果。 再检查: ```python messages[tail_start - 1] ``` 也就是: ```python messages[6] ``` 如果 `messages[6]` 是对应的 `tool_use`,说明原本的切法会变成: ```text 裁掉 tool_use 保留 tool_result ``` 这也不完整。 所以代码做: ```python tail_start -= 1 ``` 变量变化: ```python tail_start = 7 tail_start -= 1 tail_start = 6 ``` 尾部保留范围从: ```python messages[7:] ``` 变成: ```python messages[6:] ``` 这样 `tool_use` 和 `tool_result` 就一起保留下来了。 --- ### 6.6 `head_end >= tail_start` 是什么意思 源码: ```python if head_end >= tail_start: return messages ``` 正常情况下: ```python head_end < tail_start ``` 中间才有东西可以裁。 例如: ```python head_end = 3 tail_start = 7 ``` 中间可裁: ```python messages[3:7] ``` 但如果边界保护后变成: ```python head_end = 7 tail_start = 6 ``` 说明头部和尾部已经重叠了。 这时候再裁剪就没有意义,甚至可能裁错,所以直接: ```python return messages ``` --- ### 6.7 完整执行示例 假设: ```python max_messages = 8 len(messages) = 12 ``` 先计算: ```python keep_head = 3 keep_tail = 5 head_end = 3 tail_start = 7 ``` 假设消息结构是: ```text 0 user text 1 assistant text 2 assistant tool_use 3 user tool_result 4 user text 5 assistant text 6 assistant tool_use 7 user tool_result 8 user text 9 assistant text 10 user text 11 assistant text ``` 第一段保护: ```python messages[head_end - 1] = messages[2] ``` 它是 `tool_use`。 所以检查: ```python messages[head_end] = messages[3] ``` 它是 `tool_result`。 于是: ```python head_end = 4 ``` 第二段保护: ```python messages[tail_start] = messages[7] ``` 它是 `tool_result`。 并且: ```python messages[tail_start - 1] = messages[6] ``` 它是 `tool_use`。 于是: ```python tail_start = 6 ``` 现在: ```python head_end = 4 tail_start = 6 ``` 要裁掉: ```python messages[4:6] ``` 也就是下标 4 和 5。 最终保留: ```python messages[:4] + [{"role": "user", "content": "[snipped 2 messages]"}] + messages[6:] ``` 结果是: ```text 0 user text 1 assistant text 2 assistant tool_use 3 user tool_result [snipped 2 messages] 6 assistant tool_use 7 user tool_result 8 user text 9 assistant text 10 user text 11 assistant text ``` 你会发现,两个工具调用配对都没有被切断。 --- ### 6.8 一句话总结 `snip_compact()` 做的是: ```text 先算出头部保留边界和尾部保留边界 再检查边界有没有切断 tool_use / tool_result 如果切断了,就移动边界 最后把中间消息替换成一个 [snipped ... messages] 标记 ``` 它不是为了精确总结内容,而是一个便宜、快速、不会调用 LLM 的上下文裁剪方式。 --- ## 7. `micro_compact(messages)` 作用:保留最近几个完整工具结果,把更早的大工具结果替换成短文本。 ```python block["content"] = "[早前工具结果已压缩。如有需要请重新运行。]" ``` 它不会调用 LLM,所以成本很低。 --- ## 8. `compact_history(messages)` 作用:调用 LLM,把历史总结成一条新消息。 压缩前: ```python [msg1, msg2, msg3, ..., msg100] ``` 压缩后: ```python [ { "role": "user", "content": "[Compacted]\n\n这里是历史摘要" } ] ``` --- ## 9. `messages[:] = ...` `messages[:] = ...` 表示原地替换列表内容。 如果写: ```python messages = compact_history(messages) ``` 只是函数内部变量换了一个新列表。 如果写: ```python messages[:] = compact_history(messages) ``` 外面传进来的 `history` 也会看到压缩后的内容。 这就是为什么 Agent Loop 里使用 `messages[:]`。