<!-- BILINGUAL-EN-ZH -->
# Preserved Thinking - Causes, Fixes, and the Keep List / Preserved Thinking——成因、修复与保留清单

> **Reference for `shared/preserved-thinking-migration.md`** (the `preserved-thinking-migration` workflow). This file is not a workflow: it holds four lookups that the guide uses - the rules for a conversation that switches models, the "Cause -> detection -> fix" table, the keep list, and the failure modes to avoid - kept apart from the guide so that the guide fits in one Read. Read this file when the guide sends you here (Step 1.4, Step 2, or Step 3), and use the step numbers below as references into that guide.

> **`shared/preserved-thinking-migration.md` 的参考文件**（即 `preserved-thinking-migration` 工作流）。本文件不是一个工作流：它保存着指南会用到的四项查询内容——对话切换模型时的规则、"Cause -> detection -> fix" 表、保留清单，以及需要避免的失败模式——将其与指南分开存放，以便指南能一次 Read 读完。当指南把你指引到这里时（Step 1.4、Step 2 或 Step 3）阅读本文件，并以文中的步骤编号作为回指该指南的参照。

## Switching models mid-conversation / 在对话中途切换模型

A harness that routes one conversation to more than one model - a cheaper model for easy turns, a fallback when the primary is unavailable, an upgrade from the model the conversation started on - meets a second check that shares the response with the prefix check. The rules below are what the API does today, as observed on the Claude API with `claude-opus-5`, `claude-opus-5-5`, `claude-fable-5`, and `claude-fable-5-1` in both `prefix_mismatch_behavior` modes; the Preserved thinking page (under "Sources and live references" in `shared/preserved-thinking-migration.md`) states the same rules, and the Extended thinking page (section "Only for the model that produced it, or a newer one") lists which models read which blocks.

把同一段对话路由到多个模型的 harness——简单轮次交给更便宜的模型、主模型不可用时回退、或从对话最初使用的模型升级——会遇到一个与 prefix 检查并列的第二项检查。以下规则是该 API 当前的行为，观测自 Claude API 上的 `claude-opus-5`、`claude-opus-5-5`、`claude-fable-5` 与 `claude-fable-5-1`，覆盖两种 `prefix_mismatch_behavior` 模式；Preserved thinking 页面（位于 `shared/preserved-thinking-migration.md` 的 "Sources and live references" 之下）陈述了相同的规则，Extended thinking 页面（"Only for the model that produced it, or a newer one" 一节）列出了哪些模型可以读取哪些 block。

**Is the model part of what the signature records?** Yes. A `thinking` block's `signature` records the model that produced it alongside the conversation record and the previous thinking block. When the block is replayed, the API first asks whether the model now reading it reads blocks from the model that produced it (the page's rule: the model that produced it, or a newer one) - the *model check* - and only then whether the conversation before the block is unchanged - the *prefix check*. The `model` field of the request is not part of the conversation record, so changing it is not an edit: **a model switch by itself never fails the prefix check.**

**模型是否属于 signature 所记录的内容？** 是。`thinking` block 的 `signature` 在记录对话记录与上一个 thinking block 的同时，也记录了产生它的模型。当该 block 被重放时，API 首先询问当前读取它的模型能否读取由产生它的模型所产的 block（即该页面的规则：产生它的模型或更新的模型）——这是*模型检查*——然后才询问该 block 之前的对话是否未被改动——这是 *prefix 检查*。请求的 `model` 字段不属于对话记录，因此更改它不构成编辑：**仅切换模型本身永远不会导致 prefix 检查失败。**

**Does a switch break the check?** As of 2026-09-03 the API behaves as follows: a model switch does not fail a request - upgrade, downgrade, a round trip, and a downgrade combined with an edit all return 200, in `"error"` mode and in `"drop_block"` mode. `"error"` makes *edits* loud; it does not make a *model switch* loud, because the model check has no error setting: a block the current model cannot read is dropped, not rejected. Read the Preserved thinking page for the current rule. What differs by direction is whether the earlier reasoning is used:

**切换会破坏检查吗？** 截至 2026-09-03，API 的行为如下：模型切换不会使请求失败——升级、降级、往返，以及降级叠加编辑，在 `"error"` 模式与 `"drop_block"` 模式下均返回 200。`"error"` 让*编辑*变得响亮；它不会让*模型切换*变得响亮，因为模型检查没有报错设置：当前模型无法读取的 block 会被丢弃，而不是拒绝。当前规则请阅读 Preserved thinking 页面。随方向不同而不同的是：更早的推理是否被使用。

- **Upgrade (an older model's blocks replayed to Claude Fable 5.1): nothing to do.** Claude Fable 5.1 reads thinking produced by Claude Opus 5 and Claude Fable 5, among other earlier models (the Extended thinking page has the full list). The blocks are kept, fed to the model, counted in input tokens, and `input_transformations` is `[]`. An older model's block carries no conversation record for Claude Fable 5.1 to compare, so an edit made before such a block does not invalidate it; a conversation migrated from Claude Opus 5 can only break on the turns Claude Fable 5.1 produces from then on.
  - **升级（把较旧模型的 block 重放给 Claude Fable 5.1）：无需任何处理。** Claude Fable 5.1 可以读取 Claude Opus 5、Claude Fable 5 等较早模型产生的 thinking（完整清单见 Extended thinking 页面）。这些 block 会被保留、喂给模型、计入输入 token，且 `input_transformations` 为 `[]`。较旧模型的 block 没有携带可供 Claude Fable 5.1 比对的对话记录，因此在此类 block 之前做出的编辑不会使其失效；从 Claude Opus 5 迁移而来的对话，只会在 Claude Fable 5.1 自此产生的轮次上出问题。
- **Downgrade (Claude Fable 5.1's blocks replayed to Claude Opus 5 or Claude Fable 5): the reasoning is lost for that request, not the request.** The older model cannot read them, so the API leaves them out of that call, reports each one as `{"type": "thinking_dropped", "reason": "model_binding_mismatch", "path": ...}` when the beta header is on, and does not bill the dropped tokens. As of 2026-09-03 this is not a 400 in either mode, and there is no diagnosis header (that header belongs to the prefix check). Your `messages` array is never edited: the blocks stay in your history.
  - **降级（把 Claude Fable 5.1 的 block 重放给 Claude Opus 5 或 Claude Fable 5）：该请求丢失的是推理，而不是请求本身。** 较旧的模型无法读取它们，因此 API 在该次调用中将其排除，在 beta header 开启时逐个报告为 `{"type": "thinking_dropped", "reason": "model_binding_mismatch", "path": ...}`，且不对被丢弃的 token 计费。截至 2026-09-03，两种模式下这都不是 400，也没有诊断 header（那个 header 属于 prefix 检查）。你的 `messages` 数组永远不会被改动：block 仍留在你的历史中。
- **Round trip (Claude Fable 5.1, then an older model, then Claude Fable 5.1 again): nothing is lost, provided the client keeps the history intact.** Back on Claude Fable 5.1 every block is read again - the ones from before the switch, the older model's own thinking block, and the ones minted after the return. Observed through four turns: a Claude Fable 5.1 block minted after the older model's turn is valid, and the chain of Claude Fable 5.1 blocks does not care that a turn between them came from another model. The same four-turn history sent back to the older model drops exactly the Claude Fable 5.1 blocks and keeps its own.
  - **往返（先 Claude Fable 5.1，再切到较旧模型，再回到 Claude Fable 5.1）：只要客户端保持历史完整，什么都不会丢。** 回到 Claude Fable 5.1 后，每个 block 都会被再次读取——切换之前的那些、较旧模型自己的 thinking block，以及返回之后新铸的那些。经四轮观测：在较旧模型轮次之后由 Claude Fable 5.1 新铸的 block 是有效的，且 Claude Fable 5.1 的 block 链并不在意其间隔的一轮来自另一个模型。同一份四轮历史发回较旧模型时，恰好只丢弃 Claude Fable 5.1 的 block，保留它自己的。
- **Claude Opus 5.5 runs the same prefix check, and sits beside Claude Fable 5.1, not under it.** Everything in this file that names Claude Fable 5.1 as the model that runs the prefix check holds for Claude Opus 5.5 too (observed 2026-09-23: an edited history is a 400 in `"error"` mode, a `prefix_binding_mismatch` drop in `"drop_block"` mode, and a `thinking_mismatch_allowed` entry with the field unset). What differs is who reads whose blocks, per the Preserved thinking page: Claude Opus 5.5 reads thinking from Claude Opus 5 and earlier Opus, Sonnet and Haiku models, but not from Claude Fable or Claude Mythos models; on the Claude API, Claude Fable 5.1 and Claude Mythos 5.1 read Claude Opus 5.5's blocks, and no other model does. So Claude Opus 5 to Claude Opus 5.5, and Claude Opus 5.5 up to Claude Fable 5.1 on the Claude API, keep the reasoning (`[]`); Claude Fable 5.1 to Claude Opus 5.5, or Claude Opus 5.5 to any other model, is a downgrade in the sense above - `model_binding_mismatch` drops, 200 in both modes.
  - **Claude Opus 5.5 运行相同的 prefix 检查，且与 Claude Fable 5.1 并列，而非位于其下。** 本文件中所有把 Claude Fable 5.1 称为运行 prefix 检查之模型的表述，对 Claude Opus 5.5 同样成立（2026-09-23 观测：被编辑过的历史在 `"error"` 模式下返回 400，在 `"drop_block"` 模式下报告 `prefix_binding_mismatch` 丢弃，并带有一条字段未设置的 `thinking_mismatch_allowed` 条目）。差异在于谁读取谁的 block，依 Preserved thinking 页面：Claude Opus 5.5 读取来自 Claude Opus 5 及更早 Opus、Sonnet、Haiku 模型的 thinking，但不读取 Claude Fable 或 Claude Mythos 模型的；在 Claude API 上，Claude Fable 5.1 和 Claude Mythos 5.1 读取 Claude Opus 5.5 的 block，没有其他模型这么做。因此，Claude Opus 5 升到 Claude Opus 5.5，以及在 Claude API 上从 Claude Opus 5.5 升到 Claude Fable 5.1，都会保留推理（`[]`）；从 Claude Fable 5.1 降到 Claude Opus 5.5，或从 Claude Opus 5.5 到任何其他模型，则是上文意义上的降级——`model_binding_mismatch` 丢弃，两种模式下都是 200。
- **A downgrade can hide an edit for one request.** The model check runs first, so a block the older model cannot read is never prefix-judged: a request that switches down *and* edits earlier content reports only `model_binding_mismatch`, even in `"error"` mode. The edit is still there; it surfaces on the next Claude Fable 5.1 turn as an ordinary prefix break (a 400 in `"error"` mode, `prefix_binding_mismatch` in `"drop_block"` mode). Scan and measure every turn of a mixed-model conversation, not only the turn where the switch happened.
  - **降级可能把一次编辑掩盖一个请求。** 模型检查先运行，因此较旧模型无法读取的 block 永远不会被 prefix 判定：一个既切换降级*又*编辑了更早内容的请求只会报告 `model_binding_mismatch`，即便在 `"error"` 模式下也是如此。编辑依然存在；它会在下一个 Claude Fable 5.1 轮次上以普通的 prefix 断裂形式浮现（`"error"` 模式下为 400，`"drop_block"` 模式下为 `prefix_binding_mismatch`）。对混合模型对话的每一轮都要扫描和度量，而不仅仅是切换发生的那一轮。
- **Turns from a non-Claude model** do not invalidate earlier thinking, provided they are appended after the existing history as ordinary assistant messages (`text` and `tool_use` content, no `thinking` blocks) and nothing earlier changes. Claude Fable 5.1 thinking minted after such a turn stays valid.
  - **来自非 Claude 模型的轮次**不会使更早的 thinking 失效，前提是它们作为普通的 assistant 消息（`text` 与 `tool_use` 内容，不含 `thinking` block）追加在既有历史之后，且此前内容没有任何改动。在此类轮次之后新铸的 Claude Fable 5.1 thinking 仍然有效。
- **A twin model that reads the blocks but runs no conversation check** (the Extended thinking page lists which models read which): its turns report `[]` even on an edited history, so an edit made on such a turn surfaces only on the next Claude Fable 5.1 turn - the same one-request blind spot as a downgrade, with `[]` instead of a model drop to notice it by.
  - **能读取 block 但不运行对话检查的孪生模型**（Extended thinking 页面列出了谁读取谁的 block）：其轮次即使在历史被编辑过时也报告 `[]`，因此在这类轮次上做出的编辑只会在下一个 Claude Fable 5.1 轮次上浮现——与降级相同的单请求盲区，只是用以察觉它的不是模型丢弃，而是 `[]`。

**Does changing a tool description break the check?** Yes. Tools are compared as their full definitions - name, description, and `input_schema` - so changing even one tool's description invalidates every thinking block minted before the change, reported as `prefix_binding_mismatch` with `pattern=tool_schema_changed` (the API reports a description edit and a schema edit with the same word). Adding or removing a plain tool has the same effect, reported as `tool_set_changed`. Reordering tools is fine, and a `defer_loading: true` tool is outside the comparison until something references it. The fix is the `tool_schema_changed` and `tool_set_changed` rows: freeze each tool's text for the life of the conversation, and add or withdraw tools by reference - or, under `inline-tools-2026-09-15` (Claude API), append a `tool_addition` carrying the new definition ("Append-only forms under newer betas", below).

**修改工具描述会破坏检查吗？** 会。工具以其完整定义参与比较——名称、描述与 `input_schema`——因此哪怕只改一个工具的描述，也会使改动之前新铸的每一个 thinking block 失效，报告为 `prefix_binding_mismatch` 且 `pattern=tool_schema_changed`（API 用同一个词报告描述修改与 schema 修改）。新增或移除普通工具的效果相同，报告为 `tool_set_changed`。对工具重新排序没有影响，`defer_loading: true` 的工具在被引用之前不参与比较。修复方式见 `tool_schema_changed` 与 `tool_set_changed` 两行：在对话的整个生命周期内冻结每个工具的文本，通过引用来增撤工具——或者在 `inline-tools-2026-09-15`（Claude API）之下，追加一个携带新定义的 `tool_addition`（见下文"Append-only forms under newer betas"）。

**Cause -> detection -> fix.** A model switch produces findings in two shapes, and only the second is a harness bug:

**Cause -> detection -> fix（成因 -> 检测 -> 修复）。** 模型切换会产生两种形态的发现，只有第二种是 harness 缺陷：

1. **The conversation is routed to a model that cannot read its thinking.** *Detection*: `model_binding_mismatch` entries on the older model's turns; the probe prints them as `model_drops=N` per turn and `model_drop_turns` per conversation, apart from the prefix-break count, and `prefix_diff.py` prints `model_switch=A->B` on the pair where the `model` field changed (never as a MISMATCH). *Fix*: a routing decision, not a history edit. Send the same `messages` - thinking blocks included - to every model, and leave the beta header and `block_binding` field in place across the switch (the object is accepted on every model that accepts `thinking`). If the product needs the reasoning on those turns, keep the conversation on the model that produced it and switch models at conversation boundaries; if it needs the older model on those turns, accept that they run without the newer reasoning. Report it as *reasoning lost by routing* - conversations and turns affected - beside the prefix breaks, so the owner can decide.
   1. **对话被路由到无法读取其 thinking 的模型。** *检测*：较旧模型轮次上的 `model_binding_mismatch` 条目；探针按轮打印 `model_drops=N`、按对话打印 `model_drop_turns`，与 prefix 断裂计数分开；在 `model` 字段变化的那一对上，`prefix_diff.py` 打印 `model_switch=A->B`（绝不作为 MISMATCH）。*修复*：这是一个路由决策，不是历史编辑。把同一份 `messages`——包括 thinking block——发给每个模型，并在切换期间保持 beta header 与 `block_binding` 字段原位（每个接受 `thinking` 的模型都接受该对象）。如果产品需要这些轮次上的推理，就让对话留在产生它的模型上，只在对话边界切换模型；如果产品需要这些轮次由较旧模型承担，就接受它们在没有新推理的情况下运行。将其报告为*因路由而丢失的推理*——受影响的对话数与轮次数——与 prefix 断裂并列呈现，交由负责人决断。
2. **The switch triggers an edit.** Three kinds, each an ordinary prefix break with the switch as its trigger: (a) the harness strips thinking blocks on a switch, or rebuilds the history from what each model was shown - the diff shows the blocks removed (`predecessor_missing` if from the middle), and the reasoning is gone for good when the conversation returns; *fix*: stop stripping, the API already leaves out what the current model cannot read. (b) A re-rendered system prompt or tool list that *persists* past the switch - a fallback banner that stays in `system` from then on, a prompt that depends on which model answered last - reaches Claude Fable 5.1 together with blocks minted under the old text: `system_rerendered`, `tool_set_changed`, or `tool_schema_changed` on the first Claude Fable 5.1 turn after it (on a downgrade the older model's turn in between reports only the model drop). A prompt or tool set that is a pure function of the model being called is different: every Claude Fable 5.1 request then carries the same text the blocks were minted under, so the check finds nothing (verified: the original prompt restored on the return to Claude Fable 5.1 gave 200 and `[]` with the block fed, in both modes), even though the pair diff flags the switch pairs and the older model's turns still lose the reasoning through the model check. *Fix* for the persisting case: the rows for those patterns - keep each model's prompt and tool text stable across that model's own turns, and deliver anything that must change mid-conversation as an appended `role: "system"` message. (c) A request shape the target model does not accept - a `thinking` configuration it rejects (a 400 from request validation whose text names the field, not a binding failure), or mid-conversation `role: "system"` messages, and with them `tool_addition`, `tool_removal` and `clear_at`, on a model the Mid-conversation system messages page does not list (it names Claude Sonnet 5 as unsupported: "Use the top-level system field there instead"); *fix*: one request body that every model in the route accepts.
   2. **切换触发了编辑。** 三种情形，每一种都是一次普通的 prefix 断裂，切换只是其触发器：(a) harness 在切换时剥离 thinking block，或按各模型"所见"重建历史——diff 显示 block 被移除（若从中间移除则为 `predecessor_missing`），且当对话切回时推理已永久丢失；*修复*：停止剥离，API 本来就会略去当前模型读不了的内容。(b) 一次在切换之后*持续存在*的系统提示词或工具列表重渲染——一条从此驻留在 `system` 中的回退横幅、一条取决于上一个由哪个模型作答的提示词——与在旧文本下新铸的 block 一同抵达 Claude Fable 5.1：在其后的第一个 Claude Fable 5.1 轮次上出现 `system_rerendered`、`tool_set_changed` 或 `tool_schema_changed`（在降级场景下，中间那个较旧模型的轮次只报告模型丢弃）。作为"被调用模型的纯函数"的提示词或工具集则不同：此时每个 Claude Fable 5.1 请求携带的都是 block 新铸时所用的同一文本，检查一无所获（已验证：切回 Claude Fable 5.1 时恢复原始提示词，两种模式下均返回 200 和 `[]` 且 block 被正常喂入），尽管成对 diff 仍会标记切换对，且较旧模型的轮次仍经由模型检查丢失推理。*针对持续存在情形的修复*：见这些模式对应的行——让每个模型的提示词与工具文本在该模型自己的轮次间保持稳定，把任何必须在对话中途变化的内容作为追加的 `role: "system"` 消息送达。(c) 目标模型不接受某种请求形态——它拒绝的某个 `thinking` 配置（来自请求校验的 400，其文本点名字段名，而非绑定失败）、或对话中途的 `role: "system"` 消息，以及随之的 `tool_addition`、`tool_removal` 与 `clear_at`，被用在了 Mid-conversation system messages 页面未列出的模型上（该页面点名 Claude Sonnet 5 不支持："Use the top-level system field there instead"）；*修复*：让路由中的每个模型都接受同一份请求体。

【评论】该文件把"模型路由导致的推理丢失"与"harness 缺陷"明确区分为两类发现，是典型的把对未公开 API 行为的观测结论固化为工程操作规程的内部文档写法。

## Cause -> detection -> fix / 成因 -> 检测 -> 修复

The `pattern` column is the word the API's diagnosis header and the diff script both use. "Diff shows" is the attribution line from `prefix_diff.py`; "scan lead" is the heuristic `--scan` reports; the fix is the append-only form from Step 3 of `shared/preserved-thinking-migration.md`. One caution on reading the pattern word: for a summary-plus-tail compaction the header's word depends on how much was removed - the same code path reports `tail_kept`, `compaction_summary`, or `unknown` with `kind=blocks_replaced` on different conversations - so identify that cause by the kind (`blocks_removed` or `blocks_replaced`) together with the diff's attribution lines, not by one pattern word.

`pattern` 列是 API 的诊断 header 与 diff 脚本共同使用的词。"Diff shows" 是 `prefix_diff.py` 给出的归因行；"scan lead" 是 `--scan` 报告的启发式线索；修复方式即 `shared/preserved-thinking-migration.md` Step 3 的 append-only 形式。阅读 pattern 词时须注意一点：对于"摘要+尾部"式压缩，header 给出的词取决于移除了多少——同一代码路径在不同对话上会分别报告 `tail_kept`、`compaction_summary` 或 `unknown` 且 `kind=blocks_replaced`——因此要通过 kind（`blocks_removed` 或 `blocks_replaced`）结合 diff 的归因行来判定该成因，而不是靠单个 pattern 词。

| Tier | `pattern` (and `kind`) | What the harness did | Diff shows | Scan lead | Fix |
|---|---|---|---|---|---|
| 1 | `system_rerendered` (`system_changed`) | The system prompt was rebuilt with per-request content: time, cwd, account line, memory or instruction files, flags, version strings | `system[i] changed at char N` | time and environment reads near prompt builders; templates rendered per request | Render once, store the bytes with the conversation, replay; per-session facts go into the first turn; changes go out as appended `role: "system"` messages |
| 2 | `tool_set_changed` (`tools_changed`) | A tool was added or removed after the first request: a plugin or MCP server connected late, a provider disconnected, a permission changed | `tools: X added` / `removed` | `tools` mutated after session start; a tool listing fetched per request | Declare the full set at start; append a late tool with `defer_loading: true` (safe while unreferenced) and announce it with `tool_addition` in an appended system message - never append a regular tool; never remove one from the array - withdraw it with a `tool_removal` block and leave the definition in place, returning an ordinary "not available" error if the model still calls it |
| 2 | `tool_schema_changed` (`tools_changed`) | Same tool names, different description or schema text: a date, a refreshed token, a live listing, a version inside a description | `tools: X description changed at char N` | descriptions or schemas built from templates or state | Freeze each tool's text for the conversation; store and replay the definitions as sent. Without the `inline-tools-2026-09-15` beta no append-only form expresses a same-name change - a new name is the only way to offer changed text; under it, append a `tool_addition` carrying the new definition instead (the betas section below) |
| 1-2 | `system_and_tools_changed` (`multiple`) | Both re-rendered, messages untouched: a connector landing on request 2, or a restart or resume re-deriving both | both of the above | startup, resume, reconnect paths | Replay the stored prompt and tool text across restarts, and across model switches (a switch is not a boundary: a re-render that persists into later requests on the model that minted the blocks is this break, with the switch as its trigger; a prompt that is a pure function of the model called is stable on each model's own turns and is not); the only declared boundaries are a new conversation, a user-invoked reset, and the request after a full compaction |
| 1 | `first_message_rewritten` (`blocks_modified` / `blocks_removed`, often with `system` in `sections`) | The opening user message carried context rebuilt from live state: environment, instructions, a session date | `messages[0] (user) content[j] changed at char N` | `messages[0]` assigned after creation; a context header rendered per request | Announce context once and freeze it; send later changes as an appended message describing the delta |
| 3 | `rolling_truncation` (`blocks_removed`) | The oldest turns were dropped whole - a sliding window | `messages[0..k] removed` | `messages[-N:]`, keep-last, window size | Without the `compact-2026-09-04` beta no client-side form keeps the thinking (with it, on-demand compaction does, but the window must summarize rather than only drop - the betas section below). The choices: server-side compaction or context editing; simple compaction (summary plus new turn, nothing older); or keep the window and strip the retained turns' thinking as a deterministic, recorded strip, or send `drop_block` (equivalent in effect) - measured |
| 3 | `tail_kept` (`blocks_removed`) | A run of older turns removed (or replaced by a summary the record can't see) with the first message kept and the newest turns verbatim - keep-tail compaction or keep-first truncation | `messages[i..j] removed`, `messages[0]` intact | `summarize(messages[:-k])` plus `messages[-k:]` | Same as above; the retained turns' thinking cannot verify without the `compact-2026-09-04` beta (it does behind an on-demand compaction block, the betas section below) - send `drop_block` from the compaction onward or strip that thinking as a recorded decision, never mid tool-round; measure, decide, and record the decision |
| 3 | `compaction_summary` (`blocks_replaced`) | Older turns replaced in place by a shorter summary, the tail intact | `messages[i..j] replaced by 1 message(s)` | same | Same; or move the summary to simple compaction (replay nothing older than the summary); or, under `compact-2026-09-04`, on-demand compaction (the betas section below) |
| 4 | `tool_results_rewritten` (`blocks_modified`) | Old `tool_result` content trimmed or cleared after it was sent | `messages[i] (user) content[j] (tool_result) changed at char N` | tool-result truncation applied to earlier turns | Bound outputs before the first send; later clearing through server-side context editing (`clear_tool_uses_20250919`, beta `context-management-2025-06-27`); a client-side prune only at a declared boundary, as a pure function of the growing history |
| 1 | `tool_use_rewritten` (`blocks_modified`) | Old `tool_use.input` re-encoded or normalized on replay | `content[j] (tool_use) changed` | input normalizers, `to_dict` on tool calls | Echo `tool_use.input` exactly as received; normalize a copy for execution only |
| 1 | `reserialized` (`blocks_modified`) | Many blocks differ slightly across types: a lossy round trip through the app's own message model (interior whitespace, number formatting, coerced keys, trimmed text) | many `changed at char N` lines across messages | `from_dict`/`to_dict`, JSON re-encoding of history, `.strip()` on content | Persist and replay the wire JSON; never rebuild messages from domain objects |
| 2 | `reminder_stripped` / `history_block_stripped` / `block_inserted` (`blocks_removed` / `blocks_inserted`) | A per-turn text block injected into a user turn and removed on the next request (or added after the fact) | `messages[i] (user) content[j] (text) removed` / `inserted` | regex strips of reminder tags; inject-then-strip helpers | Turn-scoped system message (`clear_at: "next_user_message"`) appended after the tool results, every earlier copy left in place; without the beta, a text block after the `tool_result` blocks, left in place |
| 2 | `system_block_rerendered` / `system_blocks_stripped` | A mid-conversation `role: "system"` message re-rendered in place, or several dropped (a sub-agent transcript replayed without them) | `messages[i] (system) changed` / `removed` | transcript stores that don't keep system messages | Persist them with the transcript and replay verbatim |
| 3 | `media_stripped` (`blocks_removed` / `blocks_modified`) | Images or documents in earlier turns dropped, downsized, or replaced by a placeholder - a client media cap | `content[j] (image) removed` | image caps, resizing of stored turns | Downscale at ingestion; return images inside the producing tool's `tool_result` so server-side context editing can clear them; if a user-turn cap is unavoidable, strip deterministically to cap-minus-headroom and accept that each crossing is an edit; `file_id` only for bytes that would drift |
| 1 | `image_url_resigned` (pattern) / `media_content_changed` (kind) | A URL-sourced image or document whose block changed, or whose bytes differ from the first fetch | `content[j] (image) changed` | URL re-signing, re-uploads | The check compares the bytes, not the URL string: a rotated URL to the same bytes is fine; for content referenced across turns use a `file_id` or base64 |
| 4 | `predecessor_missing` / `predecessor_reordered` (kinds of the chain check; the header reads `kind=predecessor_missing; pattern=not_applicable`) | A thinking block removed from the middle, or re-ordered, with the rest of the prefix intact | `! messages[i] re-sent with a different set of thinking blocks` | filters on `type == "thinking"`; serializers that drop empty fields or unknown block types (a `thinking` block with empty text is still a block); a hand-rolled stream parser that loses the `signature_delta`; strip-and-retry without a record | Keep the replayed thinking blocks a contiguous window of the original (drop from the front or the back, never the middle); make any forced strip deterministic and recorded so it replays identically |
| 3 | `unknown` with kind `blocks_removed` or `blocks_replaced` | A shortening of the history that the API does not name more specifically, or several edits at once | `messages[i..j] removed` / `replaced by 1 message(s)` | the same leads as the truncation and compaction rows | Read the attribution lines; the fix is the truncation or compaction one above |
| - | `foreign_prefix` (`unrelated`) | A block from another conversation replayed (on a long conversation; a short one reports an ordinary `multiple` / `system_and_tools_changed`) | no pair diff (it is a different conversation) | session keys, multiplexed stores | Fix the session keying |
| - | (no pattern) | A drop or 400 on a pair where nothing you sent differs - the diff shows no change and the digests match | no diff | - | Not a harness bug; report the request id to Anthropic |
| 4 | `thinking_modified` (a separate 400, "cannot be modified") | A replayed thinking block's text differs from what the API returned - truncated, summarized, re-wrapped | `! the thinking text of the block in messages[i] differs` | stores that trim or reformat thinking text | Store and replay thinking blocks byte for byte |
| 0 | `model_binding_mismatch` (model check; no `pattern`, no header) | The conversation was routed to a model that cannot read its earlier thinking - a downgrade, a cheaper-model route, a fallback | `model_switch=A->B` on the pair, verdict unchanged; the probe's `model_drops` | model ids chosen per turn; fallback or router code | Not a history edit: send the same `messages` to every model, keep the field and header in place, and decide the routing - pin the conversation to the producing model, or accept that the older model's turns run without the newer reasoning; report it as reasoning lost by routing |
| 1-2 | `system_rerendered` / `tool_set_changed` / `tool_schema_changed` / `predecessor_missing`, triggered by a switch | A re-rendered prompt or tool list that persists past the switch (a fallback banner, a prompt keyed to the last responder), or thinking stripped on the switch; on a downgrade the edit is reported only on the next Claude Fable 5.1 turn. A prompt that is a pure function of the model called is not this break | the row for that pattern, on the pair after the switch (the diff also flags a per-model prompt that the API accepts - confirm with the probe) | prompts or tool lists that change with the route; thinking filtered on a switch | The row for that pattern: each model's prompt and tool text stable across its own turns, changes as an appended `role: "system"` message, and never strip thinking on a switch |

| 层级 | `pattern`（及 `kind`） | harness 做了什么 | Diff 显示 | 扫描线索 | 修复 |
|---|---|---|---|---|---|
| 1 | `system_rerendered`（`system_changed`） | 系统提示词被按请求内容重建：时间、cwd、账户信息行、记忆或指令文件、标志位、版本字符串 | `system[i] changed at char N` | 提示词构建器附近读取时间与环境；模板按请求渲染 | 只渲染一次，把字节随对话存储并重放；会话级事实放进第一轮；变更以追加的 `role: "system"` 消息发出 |
| 2 | `tool_set_changed`（`tools_changed`） | 首个请求之后新增或移除了工具：插件或 MCP 服务器晚接入、提供方断开、权限变化 | `tools: X added` / `removed` | 会话开始后 `tools` 被改动；每次请求都拉取工具列表 | 开局声明完整集合；晚到的工具以 `defer_loading: true` 追加（未被引用时安全），并用追加系统消息中的 `tool_addition` 予以宣告——绝不追加常规工具；绝不从数组中移除工具——用 `tool_removal` block 将其撤下并保留定义，若模型仍调用则返回普通的 "not available" 错误 |
| 2 | `tool_schema_changed`（`tools_changed`） | 工具名相同，描述或 schema 文本不同：日期、刷新的 token、实时清单、描述中的版本号 | `tools: X description changed at char N` | 由模板或状态构建的描述或 schema | 在整个对话中冻结每个工具的文本；按发送原样存储并重放定义。没有 `inline-tools-2026-09-15` beta 时，没有任何 append-only 形式能表达同名变更——改用新名称是提供变更文本的唯一途径；在该 beta 下，改为追加携带新定义的 `tool_addition`（见下文 betas 一节） |
| 1-2 | `system_and_tools_changed`（`multiple`） | 两者都被重渲染，消息未动：连接器落在第 2 个请求上，或重启/恢复时两者都被重新推导 | 以上两者 | 启动、恢复、重连路径 | 跨重启、跨模型切换重放存储的提示词与工具文本（切换不是边界：若重渲染持续存在于产生 block 的模型的后续请求中，即是此种断裂，切换只是其触发器；作为被调用模型纯函数的提示词在该模型自己的轮次上是稳定的，不属于此类）；唯一声明的边界是新对话、用户主动重置，以及完整压缩之后的请求 |
| 1 | `first_message_rewritten`（`blocks_modified` / `blocks_removed`，常伴有 `sections` 中的 `system`） | 开场用户消息携带了从实时状态重建的上下文：环境、指令、会话日期 | `messages[0] (user) content[j] changed at char N` | `messages[0]` 在创建之后才赋值；上下文头按请求渲染 | 上下文只宣告一次并冻结；后续变化作为描述差异（delta）的追加消息发送 |
| 3 | `rolling_truncation`（`blocks_removed`） | 最旧的轮次被整轮丢弃——滑动窗口 | `messages[0..k] removed` | `messages[-N:]`、保留末尾、窗口大小 | 没有 `compact-2026-09-04` beta 时，没有任何客户端侧形式能保住 thinking（有了它，按需压缩可以，但窗口必须做摘要而不只是丢弃——见下文 betas 一节）。可选方案：服务端压缩或上下文编辑；简单压缩（摘要加新一轮，不留更早内容）；或保留窗口，把被保留轮次的 thinking 作为确定性的、留有记录的剥离移除，或发送 `drop_block`（效果等同）——均须度量 |
| 3 | `tail_kept`（`blocks_removed`） | 移除了一段较早的轮次（或替换为记录不可见的摘要），首条消息保留、最新轮次逐字保留——保留尾部的压缩或保留首部的截断 | `messages[i..j] removed`，`messages[0]` 完好 | `summarize(messages[:-k])` 加 `messages[-k:]` | 同上；没有 `compact-2026-09-04` beta 时，被保留轮次的 thinking 无法通过校验（在按需压缩 block 之下则可以，见下文 betas 一节）——从压缩那一点起发送 `drop_block`，或将该 thinking 作为留有记录的决策剥离，绝不在工具轮中途进行；度量、决策并记录该决策 |
| 3 | `compaction_summary`（`blocks_replaced`） | 较早的轮次被原地替换为更短的摘要，尾部完好 | `messages[i..j] replaced by 1 message(s)` | 同上 | 同上；或把摘要改为简单压缩（摘要之前的内容一律不重放）；或在 `compact-2026-09-04` 下使用按需压缩（见下文 betas 一节） |
| 4 | `tool_results_rewritten`（`blocks_modified`） | 已发送的旧 `tool_result` 内容事后被裁剪或清空 | `messages[i] (user) content[j] (tool_result) changed at char N` | 对更早轮次施加的工具结果截断 | 首次发送前就限制输出规模；事后清理通过服务端上下文编辑完成（`clear_tool_uses_20250919`，beta `context-management-2025-06-27`）；客户端侧修剪只在声明的边界上、作为随历史增长的纯函数进行 |
| 1 | `tool_use_rewritten`（`blocks_modified`） | 旧的 `tool_use.input` 在重放时被重新编码或归一化 | `content[j] (tool_use) changed` | 输入归一化器、对工具调用使用 `to_dict` | `tool_use.input` 按收到时的原样回显；仅为执行而归一化一个副本 |
| 1 | `reserialized`（`blocks_modified`） | 大量 block 在各类型上略有差异：经由应用自身消息模型的有损往返（内部空白、数字格式、键被强转、文本被裁剪） | 大量跨消息的 `changed at char N` 行 | `from_dict`/`to_dict`、对历史做 JSON 重编码、对内容调用 `.strip()` | 持久化并重放线路上的 JSON；绝不从领域对象重建消息 |
| 2 | `reminder_stripped` / `history_block_stripped` / `block_inserted`（`blocks_removed` / `blocks_inserted`） | 注入到用户轮次的按轮文本 block 在下一个请求中被移除（或事后被加入） | `messages[i] (user) content[j] (text) removed` / `inserted` | 用正则剥离提醒标签；先注入后剥离的辅助函数 | 轮次作用域的系统消息（`clear_at: "next_user_message"`）追加在工具结果之后，此前每一份都原位保留；没有该 beta 时，在 `tool_result` block 之后放一个文本 block 并原位保留 |
| 2 | `system_block_rerendered` / `system_blocks_stripped` | 对话中途的 `role: "system"` 消息被原地重渲染，或几条被丢弃（子代理转写被重放时未带它们） | `messages[i] (system) changed` / `removed` | 不保存系统消息的转写存储 | 随转写持久化并逐字重放 |
| 3 | `media_stripped`（`blocks_removed` / `blocks_modified`） | 较早轮次中的图像或文档被丢弃、缩小或替换为占位符——客户端媒体上限 | `content[j] (image) removed` | 图像上限、对已存储轮次做缩放 | 在摄取时就缩小尺寸；让图像随产生它的工具的 `tool_result` 返回，以便服务端上下文编辑能清除它们；如果用户轮上限不可避免，则确定性地裁剪到"上限减去余量"并接受每次越线都是一次编辑；`file_id` 只用于本就会漂移的字节 |
| 1 | `image_url_resigned`（pattern）/ `media_content_changed`（kind） | URL 来源的图像或文档其 block 发生变化，或其字节与首次抓取不同 | `content[j] (image) changed` | URL 重新签名、重新上传 | 检查比较的是字节而非 URL 字符串：指向相同字节的轮换 URL 没有问题；跨轮引用的内容使用 `file_id` 或 base64 |
| 4 | `predecessor_missing` / `predecessor_reordered`（链检查的若干 kind；header 显示 `kind=predecessor_missing; pattern=not_applicable`） | 一个 thinking block 从中间被移除或被重新排序，其余 prefix 完好 | `! messages[i] re-sent with a different set of thinking blocks` | 针对 `type == "thinking"` 的过滤器；丢弃空字段或未知 block 类型的序列化器（文本为空的 `thinking` block 仍是 block）；丢失 `signature_delta` 的手写流解析器；不留记录的"剥离后重试" | 让重放的 thinking block 保持为原序列的连续窗口（从前部或后部丢弃，绝不从中间丢弃）；任何被迫的剥离都要确定性且有记录，以便重放完全一致 |
| 3 | `unknown`，kind 为 `blocks_removed` 或 `blocks_replaced` | API 未更具体命名的对话缩短，或同时发生多处编辑 | `messages[i..j] removed` / `replaced by 1 message(s)` | 与截断、压缩各行相同的线索 | 阅读归因行；修复方式即上文截断或压缩对应行 |
| - | `foreign_prefix`（`unrelated`） | 重放了来自另一段对话的 block（长对话上；短对话报告普通的 `multiple` / `system_and_tools_changed`） | 无成对 diff（这是另一段对话） | 会话键、多路复用的存储 | 修好会话键 |
| - | （无 pattern） | 在你发送的内容毫无差异的一对请求上出现丢弃或 400——diff 无变化、摘要一致 | 无 diff | - | 不是 harness 缺陷；将请求 id 报告给 Anthropic |
| 4 | `thinking_modified`（单独的 400，"cannot be modified"） | 重放的 thinking block 文本与 API 返回的不同——被截断、摘要、重新折行 | `! the thinking text of the block in messages[i] differs` | 会修剪或重排 thinking 文本的存储 | 逐字节存储并重放 thinking block |
| 0 | `model_binding_mismatch`（模型检查；无 `pattern`、无 header） | 对话被路由到无法读取其更早 thinking 的模型——降级、更便宜模型的路由、回退 | 该对上 `model_switch=A->B`，判定不变；探针的 `model_drops` | 按轮选择模型 id；回退或路由代码 | 不是历史编辑：向每个模型发送相同的 `messages`，保持字段与 header 原位，并对路由做出决策——把对话钉在产生推理的模型上，或接受较旧模型的轮次在没有新推理的情况下运行；将其报告为因路由而丢失的推理 |
| 1-2 | `system_rerendered` / `tool_set_changed` / `tool_schema_changed` / `predecessor_missing`，由切换触发 | 重渲染后持续存在过切换的提示词或工具列表（回退横幅、按最后作答者键控的提示词），或切换时剥离了 thinking；在降级场景下，该编辑只在下一个 Claude Fable 5.1 轮次上被报告。作为被调用模型纯函数的提示词不属于此种断裂 | 该模式对应的行，出现在切换后的那一对上（diff 也会标记 API 所接受的按模型提示词——用探针确认） | 随路由变化的提示词或工具列表；切换时对 thinking 的过滤 | 该模式对应行：每个模型的提示词与工具文本在其自己的轮次间保持稳定，变更以追加的 `role: "system"` 消息送达，且绝不在切换时剥离 thinking |

## Append-only forms under newer betas / 较新 beta 下的 append-only 形式

Two newer betas add an append-only form for shapes the table above marks as having none without them. On-demand compaction (`compact-2026-09-04`) is on the Claude API, Claude Platform on AWS, Google Cloud and Microsoft Foundry, not on Amazon Bedrock; the Compatibility list on its page names the models and platforms, and the Models API reports each model's `capabilities.compaction` with the beta header. Defining a tool inside a message (`inline-tools-2026-09-15`) is Claude API only; by-reference tool changes under the older `mid-conversation-tool-changes-2026-07-01` header also work on Amazon Bedrock and Google Cloud. Where a beta is not available, treat those shapes as the table says: measure and decide, freeze and replay.

两个较新的 beta 为上表标注"没有它们就没有对应形式"的形态补充了 append-only 形式。按需压缩（`compact-2026-09-04`）可用于 Claude API、Claude Platform on AWS、Google Cloud 与 Microsoft Foundry，不用于 Amazon Bedrock；其页面上的 Compatibility 列表给出了模型与平台，Models API 会在 beta header 下报告每个模型的 `capabilities.compaction`。在消息内定义工具（`inline-tools-2026-09-15`）仅限 Claude API；在较旧的 `mid-conversation-tool-changes-2026-07-01` header 下按引用变更工具的方式在 Amazon Bedrock 与 Google Cloud 上同样可用。在 beta 不可用之处，按表格所述对待这些形态：度量并决策，冻结并重放。

**Background and keep-tail compaction, the append-only way: on-demand compaction (beta `compact-2026-09-04`, `"compaction": {"type": "summarize"}`; the on-demand compaction page, `https://platform.claude.com/docs/en/build-with-claude/compaction-on-demand`).** Send a compaction request that carries exactly the `messages` of a request you already sent, on the conversation's model and under its `system`, `tools` and `thinking` settings, with the `compaction` field and a `max_tokens` large enough for a summary; send the beta header on it and on every request that carries the block. Exactly those messages, because the kept turns must directly follow the summarized messages and the first kept message must not be one the API would merge into the last summarized one (the same role, or a `role: "system"` message). A compaction request whose last assistant turn is still waiting on a tool result is rejected, so send the results first. Leave out `output_config.format`, `stop_sequences` and a `tool_choice` of type `any` or `tool` (the API rejects a compaction request that carries them), and never send `output_config.task_budget.remaining` on the compaction request or on any request that carries the block (a 400). The response holds one signed `compaction` block and nothing else (`stop_reason: "compaction"`). When no summary could be written there is no block: the response is still a 200 with empty `content`, and `stop_reason` is the summarization call's own - `max_tokens` (cut off), `model_context_window_exceeded` (no room for the summarization prompt), `refusal`, `tool_use`, or `end_turn` (no text) - so give `max_tokens` a few thousand tokens at least, resend with more room or fewer messages as the reason suggests, or continue without one. Custom `instructions` replace the server's summarization prompt whole (a blank value counts as absent), so ask for text only and no tool call yourself; the summarizer reads the whole conversation either way, earlier thinking included (unlike threshold compaction with custom instructions on Claude 5.1 and later models, which leaves earlier thinking out). Keep taking turns against the full history while it runs, and do not edit anything already sent.

**后台压缩与保留尾部压缩的 append-only 做法：按需压缩（beta `compact-2026-09-04`，`"compaction": {"type": "summarize"}`；按需压缩页面，`https://platform.claude.com/docs/en/build-with-claude/compaction-on-demand`）。** 发送一个压缩请求，其 `messages` 与你已发送过的某个请求完全一致，使用对话所在的模型并沿用其 `system`、`tools` 与 `thinking` 设置，附上 `compaction` 字段和一个足以容纳摘要的 `max_tokens`；在该请求以及其后每个携带该 block 的请求上都要发送 beta header。必须恰好是那些消息，因为被保留的轮次必须紧跟在被摘要的消息之后，且第一条被保留的消息不能是 API 会并入最后一条被摘要消息的那条（角色相同，或是 `role: "system"` 消息）。最后一个 assistant 轮次仍在等待工具结果的压缩请求会被拒绝，所以要先送出结果。不要带 `output_config.format`、`stop_sequences` 以及 `any` 或 `tool` 类型的 `tool_choice`（API 拒绝携带它们的压缩请求），也不要在压缩请求或任何携带该 block 的请求上发送 `output_config.task_budget.remaining`（会返回 400）。响应只包含一个已签名的 `compaction` block，别无其他（`stop_reason: "compaction"`）。当写不出摘要时则没有 block：响应仍是 200 但 `content` 为空，`stop_reason` 为摘要调用自身的取值——`max_tokens`（被截断）、`model_context_window_exceeded`（放不下摘要提示词）、`refusal`、`tool_use` 或 `end_turn`（无文本）——因此至少给 `max_tokens` 留出几千 token，按原因提示用更大空间或更少消息重发，或干脆放弃摘要继续。自定义 `instructions` 会整体替换服务端的摘要提示词（空值视同缺省），因此请自行要求只输出文本、不调用工具；无论哪种方式，摘要器都会读取整段对话，包括更早的 thinking（不同于 Claude 5.1 及以后模型上带自定义指令的阈值压缩，后者会把更早的 thinking 排除在外）。压缩运行期间可继续在完整历史上进行轮次，且不要改动任何已发送的内容。

On the first request after the block arrives, drop exactly the messages you sent to the compaction request from the front of the history and put the block first, as an `assistant` message of its own (the request that does this adopts the block; the API also accepts it as the first content block of the first kept message, and `prefix_diff.py` checks the kept turns only in the message-of-its-own form); everything appended since stays, thinking included, and keeps verifying. Keep nothing else from the dropped messages: summarized messages re-sent after the block are not rejected - the model sees them twice, summary then verbatim - so drop them yourself. Keep the block first on every later request; to compact again, send `compaction` on a request that starts with the current block, and from then on send only the newest block (a request that carries more than one `compaction` block is a 400). Do not send `compaction` and `context_management` in the same request. Text instructions in `role: "system"` messages and `tool_addition` / `tool_removal` blocks that sat inside the summarized messages stop applying at adoption (tool changes excepted when the block carries `tool_changes` - next sentence): re-declare them in a `role: "system"` message on the first request that adopts the block, directly after that request's new `user` turn (which comes after the kept turns - a system message between the block and the kept turns breaks their thinking), and leave it there afterward. Where the compaction request also carried `inline-tools-2026-09-15`, the block records the summarized messages' net tool changes in a `tool_changes` field: send the block back unmodified and they carry over by themselves, so re-declare no tool change; a block without that field carries none, so re-declare as above. The block is accepted on any model that supports the beta, with any later `system` or `tools`, but the kept turns' thinking verifies only on a model that can read it, only if every compaction request since that thinking was produced ran on a model with preserved thinking, and only while `system` and the tools other than `defer_loading: true` ones stay what the compaction request had: changing either can invalidate the kept turns' thinking and has no other effect, so to change them without losing any, compact the whole conversation first (keeping no turns), then change them on the next request. Nothing before the block is sent, but the kept turns are still checked against the summarized messages as they stood when you sent the compaction request, so do not touch them in between; `prefix_diff.py` compares them against the earlier request by aligning on a kept thinking block both carry, and when the compaction request itself is in the capture it checks that request's `system` and `tools` against the conversation's, compares the adopting request's `system` and `tools` against that request, and checks that the dropped range is the messages it carried.

在该 block 到达后的第一个请求上，把你发给压缩请求的那些消息从历史前部精确移除，并把该 block 放在首位，作为一条独立的 `assistant` 消息（执行此动作的请求即采纳该 block；API 也接受将其作为第一条被保留消息的第一个 content block，而 `prefix_diff.py` 只在"独立消息"形式下检查被保留的轮次）；此后追加的一切原样保留，thinking 也包括在内，并继续通过校验。被移除消息的其余内容一件都不要保留：block 之后再重发被摘要的消息不会被拒绝——模型会看到它们两次，先摘要后原文——所以要由你自己丢弃它们。此后的每个请求都让该 block 保持首位；要再次压缩，就在一个以当前 block 开头的请求上发送 `compaction`，从那之后只发送最新的 block（携带多于一个 `compaction` block 的请求是 400）。不要在同一请求中同时发送 `compaction` 与 `context_management`。位于被摘要消息内部的 `role: "system"` 文本消息与 `tool_addition` / `tool_removal` block，在采纳时停止生效（block 携带 `tool_changes` 时工具变更除外——见下一句）：在采纳该 block 的第一个请求上、紧跟该请求新的 `user` 轮之后（它位于被保留轮次之后——block 与被保留轮次之间插入系统消息会破坏其 thinking），用一条 `role: "system"` 消息重新声明它们，此后保持原位。若压缩请求同时带有 `inline-tools-2026-09-15`，该 block 会把被摘要消息的净工具变化记录在 `tool_changes` 字段中：把该 block 原样发回，它们会自行延续，因此不要重新声明任何工具变更；不带该字段的 block 即无工具变化，按上文重新声明即可。该 block 在任何支持该 beta 的模型上都被接受，可配任意后续的 `system` 或 `tools`，但被保留轮次的 thinking 只有在能读取它的模型上、且在该 thinking 产生之后的每个压缩请求都运行于支持 preserved thinking 的模型上、并且 `system` 与除 `defer_loading: true` 之外的工具保持与压缩请求相同时才通过校验：改动任一项都可能使被保留轮次的 thinking 失效，且没有其他作用，因此要在不丢失任何内容的前提下变更它们，先把整段对话压缩一遍（不保留任何轮次），再在下一个请求上变更。block 之前的内容不会被发送，但被保留的轮次仍会对照你发送压缩请求那一刻被摘要消息的形态进行检查，所以在此期间不要动它们；`prefix_diff.py` 通过双方共有的某个被保留 thinking block 对齐来将它们与更早的请求比较，而当压缩请求本身也在抓取中时，它会将该请求的 `system` 和 `tools` 与对话的进行比较、将采纳请求的 `system` 和 `tools` 与该请求的进行比较，并检查被移除的范围正是它所携带的那些消息。

**Tool changes by value (beta `inline-tools-2026-09-15`, Claude API; the Mid-conversation system messages and tool changes page, "Define tools in a message").** Keep `tools` exactly as the first request sent it, on every request, and make every later change by appending one `role: "system"` message: a `tool_addition` whose `tool` is `{"type": "tool_definition", "definition": {...}}` with the full `tools` entry (name, description, `input_schema`) inside `definition`, for a new tool or for a same-name tool whose description or schema changed - a different definition replaces the tool from that message on, an identical one changes nothing (safe to resend on a retry) - and a `tool_removal` by reference to withdraw one. Nothing already sent moves, so earlier thinking keeps verifying; rewording or deleting a definition message already sent is an edit like any other. The header also covers changes by reference, so it replaces `mid-conversation-tool-changes-2026-07-01`; the placement rules are the same, including no tool change directly after a paused assistant turn. A tool added this way may itself be `defer_loading: true` (inside `definition`); a tool already known at the first request belongs in `tools` with `defer_loading: true`, shown later by reference. Keep at least one non-deferred tool in `tools` (a tool search tool counts): otherwise the first tool defined by value costs one full cache miss. `cache_control` goes on the block or in the definition, not both, and never on a deferred definition. Reusing a name for a different type of tool is a 400, and during the beta some tool types (computer use among them) cannot be defined in a message - declare those in `tools` and add them by reference. A definition stays in the history after the tool is replaced or withdrawn, so a beta header that a dated tool `type` needs goes on every later request of the conversation. For an MCP connector server, add `mcp-client-2026-09-15` (in place of `mcp-client-2025-11-20`, which it includes): the `definition` can then be an `mcp_toolset` (connection details stay in `mcp_servers`), and a response for which the API fetched a server's tool list starts with one `mcp_tool_listing` block per server fetched (code that reads `content[0]` must skip them) - send the assistant message back as it came, those blocks included, keep the header on every request that carries one, and later requests reuse that list instead of asking the server again.

**按值变更工具（beta `inline-tools-2026-09-15`，Claude API；Mid-conversation system messages and tool changes 页面，"Define tools in a message"）。** 在每个请求上都让 `tools` 与第一个请求所发送的完全一致，此后的每一次变更都通过追加一条 `role: "system"` 消息完成：`tool_addition`，其 `tool` 为 `{"type": "tool_definition", "definition": {...}}`，把完整的 `tools` 条目（名称、描述、`input_schema`）放进 `definition`，用于新工具，或用于描述或 schema 发生变化的同名工具——不同的定义从该消息起替换该工具，相同的定义不改变任何东西（重试时重发是安全的）——以及按引用撤下工具的 `tool_removal`。已发送的内容一概不动，因此更早的 thinking 继续通过校验；改写或删除已发送的定义消息与其他任何编辑无异。该 header 同时覆盖按引用的变更，因此它取代了 `mid-conversation-tool-changes-2026-07-01`；放置规则相同，包括不能紧跟暂停的 assistant 轮次做工具变更。以此方式添加的工具自身可以是 `defer_loading: true`（在 `definition` 内部）；首个请求时已知的工具应放入 `tools` 并带 `defer_loading: true`，之后再按引用展示。`tools` 中至少保留一个非延迟工具（工具搜索工具也算）：否则第一个按值定义的工具要付出整整一次缓存未命中。`cache_control` 要么放在 block 上、要么放在定义里，不能两者都放，也绝不要放在延迟定义上。把同一个名称复用于不同类型的工具是 400，并且 beta 期间某些工具类型（计算机使用在其中）不能在消息中定义——在 `tools` 中声明它们，并按引用添加。定义在工具被替换或撤下之后仍留在历史中，因此带日期的工具 `type` 所需的 beta header 要挂在该对话此后每个请求上。对于 MCP 连接器服务器，添加 `mcp-client-2026-09-15`（取代 `mcp-client-2025-11-20`，前者包含后者）：此时 `definition` 可以是 `mcp_toolset`（连接细节留在 `mcp_servers` 中），而 API 为其抓取了某服务器工具列表的响应会以每个被抓取服务器一个 `mcp_tool_listing` block 开头（读取 `content[0]` 的代码必须跳过它们）——把 assistant 消息按原样发回，包括这些 block，在每个携带它们的请求上保持该 header，后续请求会复用该列表而不再询问服务器。

## Keep list - what never to flag / 保留清单——哪些情形绝不该被标记

The scan and the diff will tempt you to report things the check does not care about. These stay out of the report (or go in a "checked, fine" line):

扫描与 diff 会诱使你报告一些检查并不在意的东西。以下内容不得进入报告（或只放入一行 "checked, fine"）：

- Reordering tools in the `tools` array without changing them - compared as a name-keyed set.
  在 `tools` 数组中重新排序工具而不改动它们——按以名称为键的集合比较。
- Adding a `defer_loading: true` tool that nothing has referenced yet.
  添加尚无任何引用的 `defer_loading: true` 工具。
- Adding, moving, or removing `cache_control` markers anywhere.
  在任何位置添加、移动或移除 `cache_control` 标记。
- Changing `model`, `max_tokens`, `temperature` and other sampling parameters, `tool_choice`, `metadata`, `stop_sequences`, `stream`, `service_tier`, `output_config` (including effort), or the `thinking` configuration itself - none is part of the compared prefix today (a model change triggers the separate model check, reported as `model_binding_mismatch`, not this one). Request headers are not part of it either, except that a beta which makes the API inject a tool server-side (code execution, the web-search fallback) changes the tool set with an identical body.
  更改 `model`、`max_tokens`、`temperature` 及其他采样参数、`tool_choice`、`metadata`、`stop_sequences`、`stream`、`service_tier`、`output_config`（包括 effort）或 `thinking` 配置本身——今天这些都不属于被比较的 prefix（模型变更触发的是单独的模型检查，报告为 `model_binding_mismatch`，而不是本检查）。请求 header 同样不属于它，但有一个例外：会让 API 在服务端注入工具的 beta（代码执行、web-search 回退）会在请求体完全相同的情况下改变工具集。
- String content versus a single text block of the same text; leading or trailing whitespace of a text block; whitespace-only text blocks; JSON key order; `1` versus `1.0`.
  字符串内容与包含同一文本的单个文本 block 之间的差异；文本 block 的前导或尾随空白；纯空白的文本 block；JSON 键顺序；`1` 与 `1.0`。
- A rotated or re-signed URL for an image or document that serves the same bytes.
  为提供相同字节的图像或文档轮换或重新签名的 URL。
- Removing thinking blocks from the start of the history (oldest first), from the end, or all of them - allowed, and the Preserved thinking page documents all three. The kept blocks must stay an unbroken run of the original sequence; it costs the reasoning, not the validity of the kept blocks. An assistant turn whose `tool_use` still awaits its `tool_result` should keep its thinking (the page asks for that). Two things are not on this list: removing a block from the middle, which invalidates every block after it, and putting a removed block back, which invalidates the blocks produced while it was gone.
  从历史开头（最旧优先）、结尾移除 thinking block，或将其全部移除——都是允许的，Preserved thinking 页面对这三种情形均有说明。被保留的 block 必须仍是原序列中一段不间断的连续区间；它付出的是推理，而不是被保留 block 的有效性。`tool_use` 仍在等待其 `tool_result` 的 assistant 轮次应保留其 thinking（该页面有此要求）。有两件事不在本清单之列：从中间移除一个 block（这会使其后每个 block 失效），以及把已移除的 block 放回去（这会使它缺席期间产生的 block 失效）。
- Anything the API itself adds or rewrites server-side (server-side compaction, context editing, thinking clearing, its own injected text) - the check runs on the request as you sent it.
  API 自身在服务端添加或重写的任何内容（服务端压缩、上下文编辑、thinking 清除、它自己注入的文本）——检查针对的是你发出的请求。
- Mid-conversation `role: "system"` messages and cleared turn-scoped messages that are left in place and replayed verbatim.
  原位保留并逐字重放的对话中途 `role: "system"` 消息与已清除的轮次作用域消息。

【评论】"Keep list" 以白名单方式列出不构成 prefix 编辑的差异（参数变更、URL 重签、键序等），其目的是把自动化扫描的误报压到最低，与前文的诊断表互为补充。

## Failure modes to avoid / 需要避免的失败模式

- **Measuring on a slice that never replays thinking.** The first request of every conversation replays nothing; short turns may return no thinking block; a harness that strips thinking has nothing to check. Read the replayed-thinking count before the drop count, every run.
  - **在一个从不重放 thinking 的切片上度量。** 每个对话的第一个请求不重放任何内容；短轮次可能不返回 thinking block；剥离 thinking 的 harness 没有任何可检查的东西。每次运行都先读重放 thinking 的计数，再看丢弃计数。
- **Counting entries instead of breaks.** A stale block re-fails on every later request. Report conversations with a first break and the turn it happened at; an entries-per-request number only ever goes up with conversation length.
  - **数条目而不是数断裂。** 一个失效的 block 会在之后每个请求上重新报失败。报告"出现首次断裂的对话及其发生轮次"；"每请求条目数"只会随对话长度上升。
- **Trusting the header over the body.** The diagnosis header is best-effort and unpublished; `input_transformations` is the contract. A drop with no header is still a drop; a header-only pipeline misses every drop where the header did not arrive.
  - **信 header 胜过信响应体。** 诊断 header 是尽力而为且未公开的；`input_transformations` 才是契约。没有 header 的丢弃仍是丢弃；只看 header 的流水线会漏掉每一次 header 未抵达的丢弃。
- **Reading a header-only run as an enforcement test.** On an organization that is not enforced yet the header alone records failures as `thinking_mismatch_allowed` and drops nothing: good for finding edits in production, no measure of what enforcement costs. Set `prefix_mismatch_behavior` for that.
  - **把仅 header 的运行当作强制执行测试。** 在尚未强制执行的组织上，header 只把失败记录为 `thinking_mismatch_allowed` 而不丢弃任何东西：适合在生产中发现编辑，却不是对强制执行成本的度量。要做度量请设置 `prefix_mismatch_behavior`。
- **Misspelling the field.** Only `block_binding.prefix_mismatch_behavior` is the documented name; the probe removes any other key it finds under `block_binding` and says which; production code should write the documented name.
  - **拼错字段名。** 只有 `block_binding.prefix_mismatch_behavior` 是文档化名称；探针会移除它在 `block_binding` 下发现的任何其他键并说明是哪个；生产代码应写文档化名称。
- **Reading the capture from the app's own objects.** A capture rebuilt from ORM or domain objects hides the re-serialization the check catches. Capture at the HTTP layer.
  - **从应用自身的对象读取抓取。** 从 ORM 或领域对象重建的抓取会掩盖检查要抓的重新序列化。请在 HTTP 层抓取。
- **Replaying someone else's conversation.** A capture from another organization still gets its blocks dropped, but it is not diagnosed, and the drop teaches nothing about the harness. Replay with a key from the organization that produced the capture.
  - **重放别人的对话。** 来自其他组织的抓取仍会被丢弃 block，但不会得到诊断，而且这种丢弃对 harness 毫无教益。请用产生该抓取的组织所拥有的密钥重放。
- **Bundling fixes.** Two causes fixed in one diff cannot be attributed or reverted separately. One cause per diff, re-measured each time.
  - **捆绑修复。** 一个 diff 里修两个成因，便无法分别归因或回滚。每个 diff 只修一个成因，每次都重新度量。
- **Prescribing a compaction rewrite as if it were required.** Keep-tail and background compaction have no append-only client-side form without the `compact-2026-09-04` beta (on-demand compaction); measure the scheme the product has, choose `error` or `drop_block`, and record the decision.
  - **把某种压缩改写方案当作强制要求来开方。** 没有 `compact-2026-09-04` beta（按需压缩）时，保留尾部压缩与后台压缩没有任何 append-only 的客户端形式；度量产品现有的方案，在 `error` 与 `drop_block` 之间做选择，并记录决策。
- **Setting `drop_block` and calling it done.** `"drop_block"` hides the error but doesn't fix the edit that caused it. Dropped blocks aren't billed, but a session's token usage might still increase because Claude can sometimes think more to re-create the dropped thinking. The increase tends to be larger when more thinking blocks are dropped, or when blocks are dropped on more turns of a long session. Use `drop_block` to measure (Step 2) and as a recorded stopgap, count the responses in each session whose `input_transformations` has a `prefix_binding_mismatch` entry, and fix the edit.
  - **设了 `drop_block` 就当大功告成。** `"drop_block"` 掩盖了错误，却没有修复造成它的编辑。被丢弃的 block 不计费，但会话的 token 用量仍可能上升，因为 Claude 有时会为了重建被丢弃的 thinking 而进行更多思考。当丢弃的 thinking block 更多、或长会话中更多轮次发生丢弃时，增幅往往更大。把 `drop_block` 用于度量（Step 2）并作为留有记录的权宜之计，统计每个会话中 `input_transformations` 含 `prefix_binding_mismatch` 条目的响应数，然后修掉那个编辑。
- **Resending the refused body, or fixing a broken session per request.** The Preserved thinking page's "Handle the error in code": retry once with the beta header and `prefix_mismatch_behavior: "drop_block"`, and store that choice with the session so every later request sends it too, including after a restart; where the beta header cannot be sent, remove every `thinking` and `redacted_thinking` block from the history once and leave them out. A saved session that now fails on every request has the edit stored in it: the same remedy applies, thinking produced from then on stays valid as long as nothing before it changes again, and the edit still has to be found so new sessions do not hit it.
  - **重发被拒的请求体，或按请求逐一"修复"坏掉的会话。** 按 Preserved thinking 页面的 "Handle the error in code"：带上 beta header 与 `prefix_mismatch_behavior: "drop_block"` 重试一次，并把该选择与会话一起存储，使此后每个请求（包括重启之后）都发送它；在无法发送 beta header 之处，一次性从历史中移除所有 `thinking` 与 `redacted_thinking` block 并保持移除。一个如今每个请求都失败的已保存会话，其内部存着那个编辑：同样的补救照样适用，只要之前的内容不再变动，从那时起产生的 thinking 保持有效，但仍必须找到那个编辑，让新会话不再踩中。
- **A library, proxy or gateway that rewrites what it forwards.** Its rewrites are edits its users cannot see or fix. Forward the caller's `anthropic-beta` values and `thinking.block_binding` unchanged and return `input_transformations` to them (an options schema that rejects unknown keys stops a caller from choosing `"drop_block"`); leave a `role: "system"` message where the caller put it - moving it into the top-level `system` field invalidates every thinking block in the conversation; turn tool use off with `tool_choice: {"type": "none"}`, never by removing `tools`; and do not hide the 400 - code that catches it, strips thinking and retries on the caller's behalf logs that it did.
  - **会改写所转发内容的库、代理或网关。** 它的改写是其用户既看不见也修不了的编辑。原样转发调用方的 `anthropic-beta` 值与 `thinking.block_binding`，并把 `input_transformations` 返回给它们（拒绝未知键的 options schema 会阻止调用方选择 `"drop_block"`）；把 `role: "system"` 消息留在调用方放置的位置——把它挪进顶层 `system` 字段会使对话中每个 thinking block 失效；用 `tool_choice: {"type": "none"}` 关闭工具使用，绝不要通过移除 `tools`；并且不要把 400 藏起来——替调用方捕获它、剥离 thinking 并重试的代码要如实记录自己这么做了。
- **An unrecorded strip.** Stripping thinking after a 400 without making the strip deterministic and recorded re-sends the refused blocks on the next turn and fails again, every turn, for the rest of the conversation.
  - **不留记录的剥离。** 在 400 之后剥离 thinking 却不把剥离做成确定性且有记录，下一个轮次就会重发被拒的 block 并再次失败，一轮接一轮，贯穿对话余生。
- **Leaving the production value unset.** Defaults differ by surface and by account age; an unset field on a not-yet-enforced account means the check is only recorded (`thinking_mismatch_allowed`), and that default changes the day the account or the model is enforced. Set it, and monitor the entries or the 400s.
  - **让生产值保持未设置。** 默认值因平台与账户年龄而异；在尚未强制执行的账户上让字段保持未设置，意味着检查只做记录（`thinking_mismatch_allowed`），而当账户或模型被强制执行的那一天，这一默认值就会改变。把它设置好，并监控条目或 400。
- **Mistaking the model check for this one.** `model_binding_mismatch` entries after a model switch are expected and unbilled; only `prefix_binding_mismatch` is a harness finding.
  - **把模型检查误当成这个检查。** 模型切换后的 `model_binding_mismatch` 条目是预期内的且不计费；只有 `prefix_binding_mismatch` 才是 harness 缺陷。
