"Compaction"(压缩)通过在接近上下文窗口限制时自动对较早的上下文进行摘要,从而延长长时间运行的对话和任务的有效上下文长度。它还能保持活动上下文的精简:随着对话的增长,响应质量会下降,因此压缩会用简洁的摘要替换较早的内容。
这非常适合以下场景:
启用压缩后,当对话达到配置的令牌阈值时,Claude 会自动对您的对话进行摘要。API 会:
compaction 块。在后续请求中,将响应追加到您的消息中。API 会自动丢弃 compaction 块之前的所有内容块,从摘要处继续对话。
通过在 Messages API 请求的 context_management.edits 中添加 compact_20260112 策略来启用压缩。
client = anthropic.Anthropic()
messages = [{"role": "user", "content": "Help me build a website"}]
response = client.beta.messages.create(
betas=["compact-2026-01-12"],
model="claude-opus-5",
max_tokens=4096,
messages=messages,
context_management={"edits": [{"type": "compact_20260112"}]},
)
# 将响应(包括任何压缩块)追加到对话中以继续对话
messages.append({"role": "assistant", "content": response.content})| 参数 | 类型 | 默认值 | 描述 |
|---|---|---|---|
type | string | 必填 | 必须为 "compact_20260112" |
trigger | object | {"type": "input_tokens", "value": 150000} | 何时触发压缩。input_tokens 是唯一支持的触发类型。value 必须至少为 50,000 个令牌。 |
pause_after_compaction | boolean | false | 是否在生成压缩摘要后暂停 |
instructions | string | null | 自定义摘要提示。提供此参数时会完全替换默认提示。 |
使用 trigger 参数配置何时触发压缩:
client = anthropic.Anthropic()
messages = [{"role": "user", "content": "Hello, Claude"}]
response = client.beta.messages.create(
betas=["compact-2026-01-12"],
model="claude-opus-5",
max_tokens=4096,
messages=messages,
context_management={
"edits": [
{
"type": "compact_20260112",
"trigger": {"type": "input_tokens", "value": 150000},
}
]
},
)默认的摘要提示因模型而异。每个默认提示都会指示 Claude 在 <summary></summary> 标签内编写摘要,其中包含在未来的上下文窗口中继续执行任务所需的信息。例如,某些模型使用以下提示:
You have written a partial transcript for the initial task above. Please write a summary of the transcript. The purpose of this summary is to provide continuity so you can continue to make progress towards solving the task in a future context, where the raw history above may not be accessible and will be replaced with this summary. Write down anything that would be helpful, including the state, next steps, learnings etc. You must wrap your summary in a <summary></summary> block.您可以通过 instructions 参数提供自定义指令。自定义指令不是对默认提示的补充,而是会完全替换它:
client = anthropic.Anthropic()
messages = [{"role": "user", "content": "Hello, Claude"}]
response = client.beta.messages.create(
betas=["compact-2026-01-12"],
model="claude-opus-5",
max_tokens=4096,
messages=messages,
context_management={
"edits": [
{
"type": "compact_20260112",
"instructions": "Focus on preserving code snippets, variable names, and technical decisions.",
}
]
},
)使用 pause_after_compaction 可在生成压缩摘要后暂停 API。这允许您在 API 继续生成响应之前添加额外的内容块(例如保留最近的消息或特定的指令性消息)。
启用后,API 会在生成压缩块后返回一条带有 compaction 停止原因的消息:
client = anthropic.Anthropic()
messages = [{"role": "user", "content": "Hello, Claude"}]
response = client.beta.messages.create(
betas=["compact-2026-01-12"],
model="claude-opus-5",
max_tokens=4096,
messages=messages,
context_management={
"edits": [{"type": "compact_20260112", "pause_after_compaction": True}]
},
)
# 检查压缩是否触发了暂停
if response.stop_reason == "compaction":
# 响应仅包含压缩块
messages.append({"role": "assistant", "content": response.content})
# 继续请求
response = client.beta.messages.create(
betas=["compact-2026-01-12"],
model="claude-opus-5",
max_tokens=4096,
messages=messages,
context_management={"edits": [{"type": "compact_20260112"}]},
)当模型处理包含大量工具使用迭代的长任务时,总令牌消耗可能会显著增长。您可以将 pause_after_compaction 与压缩计数器结合使用,以估算累计用量,并在达到预算后优雅地结束任务。
此示例仅以 SDK 语言呈现:其价值在于围绕请求的预算跟踪逻辑。原始请求将触发器配置中的 trigger 与压缩后暂停中的 pause_after_compaction 结合使用。
client = anthropic.Anthropic()
messages = [{"role": "user", "content": "Hello, Claude"}]
TRIGGER_THRESHOLD = 100_000
TOTAL_TOKEN_BUDGET = 3_000_000
n_compactions = 0
response = client.beta.messages.create(
betas=["compact-2026-01-12"],
model="claude-opus-5",
max_tokens=4096,
messages=messages,
context_management={
"edits": [
{
"type": "compact_20260112",
"trigger": {"type": "input_tokens", "value": TRIGGER_THRESHOLD},
"pause_after_compaction": True,
}
]
},
)
if response.stop_reason == "compaction":
n_compactions += 1
messages.append({"role": "assistant", "content": response.content})
# 估算已消耗的总令牌数;若超出预算则提示收尾
if n_compactions * TRIGGER_THRESHOLD >= TOTAL_TOKEN_BUDGET:
messages.append(
{
"role": "user",
"content": "Please wrap up your current work and summarize the final state.",
}
)当触发压缩时,API 会在助手响应的开头返回一个 compaction 块。
长时间运行的对话可能会产生多次压缩。最后一个压缩块反映了提示的最终状态,用生成的摘要替换其之前的内容。
{
"content": [
{
"type": "compaction",
"content": "Summary of the conversation: The user requested help building a web scraper..."
},
{
"type": "text",
"text": "Based on our conversation so far..."
}
]
}您必须在后续请求中将 compaction 块传回 API,以便使用缩短后的提示继续对话。最简单的方法是将整个响应内容追加到您的消息中:
client = anthropic.Anthropic()
messages = [{"role": "user", "content": "Hello, Claude"}]
response = client.beta.messages.create(
betas=["compact-2026-01-12"],
model="claude-opus-5",
max_tokens=4096,
messages=messages,
context_management={"edits": [{"type": "compact_20260112"}]},
)
# 收到包含压缩块的响应后
messages.append({"role": "assistant", "content": response.content})
# 继续对话
messages.append({"role": "user", "content": "Now add error handling"})
response = client.beta.messages.create(
betas=["compact-2026-01-12"],
model="claude-opus-5",
max_tokens=4096,
messages=messages,
context_management={"edits": [{"type": "compact_20260112"}]},
)当 API 收到 compaction 块时,该块之前的所有内容块都会被忽略。您可以选择:
压缩块的流式传输方式与文本块不同。您会收到一个 content_block_start 事件,随后是一个包含完整摘要内容的单个 content_block_delta(没有中间流式传输),然后是一个 content_block_stop 事件。
client = anthropic.Anthropic()
messages = [{"role": "user", "content": "Hello, Claude"}]
with client.beta.messages.stream(
betas=["compact-2026-01-12"],
model="claude-opus-5",
max_tokens=4096,
messages=messages,
context_management={"edits": [{"type": "compact_20260112"}]},
) as stream:
for event in stream:
if event.type == "content_block_start":
if event.content_block.type == "compaction":
print("Compaction started...")
elif event.content_block.type == "text":
print("Text response started...")
elif event.type == "content_block_delta":
if event.delta.type == "compaction_delta":
print(f"Compaction complete: {len(event.delta.content or '')} chars")
elif event.delta.type == "text_delta":
print(event.delta.text, end="", flush=True)
# 获取最终累积的消息
message = stream.get_final_message()
messages.append({"role": "assistant", "content": message.content})压缩可与提示缓存良好配合。您可以在压缩块上添加 cache_control 断点来缓存摘要内容。
{
"role": "assistant",
"content": [
{
"type": "compaction",
"content": "[summary text]",
"cache_control": { "type": "ephemeral" }
},
{
"type": "text",
"text": "Based on our conversation..."
}
]
}当发生压缩时,摘要会成为需要写入缓存的新内容。如果没有额外的缓存断点,这也会使任何已缓存的系统提示失效,需要将其与压缩摘要一起重新缓存。
为了最大化缓存命中率,请在系统提示的末尾添加一个 cache_control 断点。这会使系统提示与对话分开缓存,因此当发生压缩时:
client = anthropic.Anthropic()
messages = [{"role": "user", "content": "Hello, Claude"}]
response = client.beta.messages.create(
betas=["compact-2026-01-12"],
model="claude-opus-5",
max_tokens=4096,
system=[
{
"type": "text",
"text": "You are a helpful coding assistant...",
"cache_control": {
"type": "ephemeral"
}, # Cache the system prompt separately
}
],
messages=messages,
context_management={"edits": [{"type": "compact_20260112"}]},
)这可以使较长的系统提示在整个对话的多次压缩事件中保持缓存状态。
压缩需要额外的采样步骤,这会计入速率限制和计费。API 会在响应中返回详细的用量信息:
{
"usage": {
"input_tokens": 23000,
"output_tokens": 1000,
"iterations": [
{
"type": "compaction",
"input_tokens": 180000,
"output_tokens": 3500
},
{
"type": "message",
"input_tokens": 23000,
"output_tokens": 1000
}
]
}
}iterations 数组显示每次采样迭代的用量。当发生压缩时,您会看到一个 compaction 迭代,后跟主 message 迭代。在此示例中,顶层的 input_tokens 和 output_tokens 与 message 迭代完全匹配,因为只有一个非压缩迭代。最后一次迭代的令牌计数反映了压缩后的有效上下文大小。
使用服务器工具(例如网络搜索)时,会在每次采样迭代开始时检查压缩触发器。根据您的触发阈值和生成的输出量,单个请求中可能会发生多次压缩。
令牌计数端点(/v1/messages/count_tokens)会应用提示中现有的 compaction 块,但不会触发新的压缩。使用它来检查之前压缩后的有效令牌计数:
client = anthropic.Anthropic()
messages = [{"role": "user", "content": "Hello, Claude"}]
count_response = client.beta.messages.count_tokens(
betas=["compact-2026-01-12"],
model="claude-opus-5",
messages=messages,
context_management={"edits": [{"type": "compact_20260112"}]},
)
print(f"Current tokens: {count_response.input_tokens}")
print(f"Original tokens: {count_response.context_management.original_input_tokens}")以下是一个使用压缩的长时间运行对话的完整示例:
client = anthropic.Anthropic()
messages: list[dict] = []
def chat(user_message: str) -> str:
messages.append({"role": "user", "content": user_message})
response = client.beta.messages.create(
betas=["compact-2026-01-12"],
model="claude-opus-5",
max_tokens=4096,
messages=messages,
context_management={
"edits": [
{
"type": "compact_20260112",
"trigger": {"type": "input_tokens", "value": 100000},
}
]
},
)
# 追加响应(压缩块会自动包含在内)
messages.append({"role": "assistant", "content": response.content})
# 返回文本内容
return next(block.text for block in response.content if block.type == "text")
# 运行长对话
print(chat("Help me build a Python web scraper"))
print(chat("Add support for JavaScript-rendered pages"))
print(chat("Now add rate limiting and error handling"))
# 根据对话需要持续调用 chat()以下示例使用 pause_after_compaction 来逐字保留上一轮交互和当前用户消息(共三条消息),而不是对它们进行摘要:
from typing import Any
client = anthropic.Anthropic()
messages: list[dict[str, Any]] = []
def chat(user_message: str) -> str:
messages.append({"role": "user", "content": user_message})
response = client.beta.messages.create(
betas=["compact-2026-01-12"],
model="claude-opus-5",
max_tokens=4096,
messages=messages,
context_management={
"edits": [
{
"type": "compact_20260112",
"trigger": {"type": "input_tokens", "value": 100000},
"pause_after_compaction": True,
}
]
},
)
# 检查是否发生了压缩并暂停
if response.stop_reason == "compaction":
# 从响应中获取压缩块
compaction_block = response.content[0]
# 保留之前的对话交换 + 当前用户消息(共 3 条消息)
# 方法是将它们附加在压缩块之后
preserved_messages = messages[-3:] if len(messages) >= 3 else messages
# 构建新的消息列表:压缩块 + 保留的消息
new_assistant_content = [compaction_block]
messages_after_compaction = [
{"role": "assistant", "content": new_assistant_content}
] + preserved_messages
# 使用压缩后的上下文 + 保留的消息继续请求
response = client.beta.messages.create(
betas=["compact-2026-01-12"],
model="claude-opus-5",
max_tokens=4096,
messages=messages_after_compaction,
context_management={"edits": [{"type": "compact_20260112"}]},
)
# 更新我们的消息列表以反映压缩结果
messages.clear()
messages.extend(messages_after_compaction)
# 追加最终响应
messages.append({"role": "assistant", "content": response.content})
# 返回文本内容
return next(block.text for block in response.content if block.type == "text")
# 运行一段长对话
print(chat("Help me build a Python web scraper"))
print(chat("Add support for JavaScript-rendered pages"))
print(chat("Now add rate limiting and error handling"))
# 根据对话需要持续调用 chat()摘要使用相同模型: 您在请求中指定的模型将用于生成摘要。目前无法选择使用不同的(例如更便宜的)模型来生成摘要。
定义工具时压缩可能失败: 当您的请求包含 tools 时,模型偶尔会在内部摘要步骤中调用工具,而不是编写摘要。发生这种情况时,响应会包含一个 content: null 的 compaction 块。为防止这种情况,请将 instructions 设置为明确告知模型不要调用工具的提示,例如:
Summarize the transcript inside <summary></summary> tags. Include relevant information in the summary for continuing the task in the next context window. Do not call any tools while writing this summary; respond with text only.通过上下文编辑,在对话上下文增长时自动进行管理。
了解上下文窗口大小和管理策略。
探索一个实际实现,使用后台线程和提示缓存,通过即时会话内存压缩来管理长时间运行的对话。
| Supported models |
|
|---|---|
| Supported platforms |
|
Was this page helpful?