Prompt caching works when consecutive provider requests share a long, byte-stable prefix. A coding agent naturally challenges that requirement: tools connect, memory changes, results grow, and child tasks need different instructions. Claude Code therefore controls where changing data appears and records which request fields changed when reuse drops.
Reading goal: follow cache markers from request assembly, then inspect the designs that preserve old bytes—session-fixed options, fork suffixes, cached microcompact, deterministic tool-result replacement, and cache-break detection.
stable prefix = system + tools + headers + earlier messages
dynamic tail = current user turn + new tool output
addCacheBreakpoints(messages)
-> place one cache_control marker
-> keep volatile blocks after the reusable prefix
content replacement / cache_edits
-> shorten future provider view
-> keep local transcript and recovery state separate
This article verifies the public mirror’s fixed March 31, 2026 snapshot at one immutable commit. The mirror is not Anthropic’s official source repository and does not establish later product behavior. Client implementation, official API contracts, and provider-side inference remain separate.
1. Cache markers are only part of the request
getCacheControl() produces an ephemeral marker and may add a one-hour TTL or global scope. The conditional scope: "global" field is visible in this client; it is not a universal public API contract. buildSystemPromptBlocks() places a marker on the system blocks, while addCacheBreakpoints() adds at most one message-level marker. A normal request marks its final message. With skipCacheWrite, Claude Code moves the marker to the preceding message so a short-lived fork can read the shared parent prefix without writing its private suffix.
The TTL choice is itself kept stable. should1hCacheTTL() first returns true for Bedrock with an explicit ENABLE_PROMPT_CACHING_1H_BEDROCK opt-in. Other paths latch user eligibility and allowlist state in bootstrap state on first use instead of reevaluating a changing experiment or overage condition on every turn. Otherwise the marker could alternate between the default TTL and one hour even when all prompt text remained identical. Source comments treat that as a cache risk; this does not prove that every provider uses identical TTL cache-key rules.
main turn:
S = system + tools + stable history
U = current user turn; marker after S + U
fire-and-forget fork:
S = shared parent prefix
F = fork-only task; marker after S, before F
example F:
a compact summary instruction, still sent as uncached input2. Dynamic material is kept out of the stable prefix when possible
Deferred tools and late Chrome instructions illustrate the problem. Prepending a fresh meta message whenever tools appear would change every later byte. When delta attachments are enabled, Claude Code emits deferred_tools_delta or mcp_instructions_delta instead of repeatedly changing earlier messages or the system prompt.
Global caching is also disabled for the affected system path when user-specific dynamic MCP tools are rendered. A per-user tool list should not be treated as globally stable content. Advisor tools are appended after existing schemas so toggling the advisor disturbs the tail instead of reordering the whole tool list.
Memory follows the same rule. CLAUDE.md and rules enter an early meta user message through prependUserContext(); auto-memory content may enter user context or relevant-memory attachments. If the final text is unchanged, it can remain reusable. If a memory file or index changes, the provider sees different prefix bytes even though the feature is still called “memory.”
Beta headers can change cache identity without changing prompt text. Claude Code therefore keeps AFK, fast mode, cache editing, and thinking-clear headers sticky once first sent during a session. The header can stay present while current behavior—such as whether fast speed is active after cooldown—continues to change through another request field.
| Changing input | Client handling | Result |
|---|---|---|
| Deferred tools | Send a delta attachment instead of prepending a new message. | Earlier message bytes stay in place. |
| Late Chrome instructions | Move them out of the system prompt when delta support is available. | The system blocks remain stable. |
| User-specific MCP tools | Avoid global caching for the affected content. | Private dynamic schemas are not mislabeled as globally reusable. |
| Beta headers | Keep a sent header present for the rest of the session. | Mid-session feature changes do not silently change cache identity. |
3. Fork children copy the parent prefix exactly
A fork child inherits the parent's rendered system prompt, exact tool array, messages, non-interactive setting, and thinking configuration. buildForkedMessages() copies the parent assistant message and placeholder tool results, then appends a directive and task for that child.
This is why a fork is intentionally unlike a normal subagent. Rebuilding worker tools under another permission mode could change serialized tool definitions at the first schema. A new system prompt or thinking setting would do the same. useExactTools: true preserves those request fields instead.
Two implementations should remain distinct: the preceding explanation describes an AgentTool fork child, while the cover focuses on one-off internal runForkedAgent() calls such as compact summaries. Both preserve shared parameters; transcript writing, tool restrictions, and skipCacheWrite depend on the caller, so the cover is not a rule for every subagent.
4. Full compact and cached microcompact take different approaches
Full compact summarizes older work and creates a new prefix for later turns. The summary request can still reuse the current session's cached prefix through its fork-like path. Once the summary is installed, however, future requests naturally use the new compacted history.
Cached microcompact tries to shorten old tool results without rewriting local messages. When the service cache is likely still useful and the current model and thread support the feature, cachedMicrocompactPath() leaves message content untouched, queues cache_edits; the API adapter later adds cache_reference fields to tool-result blocks. Queuing is only a deletion request: cache_deleted_input_tokens from the response, minus its earlier cumulative value, provides the reported per-operation result. If the cache is probably expired, Claude Code can clear old result content directly because the full prefix will be written again anyway.
Pending edits are consumed before paramsFromContext can be called by logging, retry, or fallback code. Consuming the same edit twice during one turn could make retry requests differ from the original attempt.
cache-edit state transition (simplified):
old tool_result: tool_use_id = T, cache_reference = T, original content
queue this turn: cache_edits → delete(cache_reference = T)
API assembly: insert and pin edits; retries reuse the same batch
response: cumulative cache_deleted_input_tokens - prior total = reported deletion
Local transcript retains the result; queuing does not mean deletion completed.5. Tool-result replacement must repeat the same decision
Large tool results can be stored locally and represented by a preview in the model request. The important cache rule is consistency: an old result cannot be full text on one turn and a preview on another, and a regenerated preview cannot differ by a few bytes.
ContentReplacementState records ids that have already been evaluated and the exact preview used for replacements. Later turns reapply the preview through a map lookup without reading the file again. Previously seen results that were not replaced are frozen in that state; only fresh results participate in the next budget decision. Forks that share a prefix clone the replacement state so parent and child serialize the same old results.
The update around asynchronous file persistence is also kept atomic. If another reader observed an id in seenIds before its replacement appeared, it could send full content while the main thread sends the preview. The same historical result would then have two different wire representations.
replacement decisions (conceptual shape):
seenIds = {A, B}
replacements = {B: "saved result location + identical preview"}
next turn:
A was seen, no replacement → retain original; do not decide again
B has a replacement → reuse the exact string without file I/O
C was never seen → evaluate against this turn's budget
resume restores decisions; a cache-sharing fork copies them6. Cache-break detection compares requests with response counters
PreviousState records client-visible inputs that can affect reuse: hashes of system blocks, tools, and cache control; tool names; model; fast mode; global strategy; beta headers; overage; microcompact state; effort; extra body; and prior cache-read tokens. Separate tracking keys keep unrelated agents and SDK calls from being compared, while compact shares the main REPL key because its summary path reuses those parameters.
Before the request, recordPromptState() captures the final sticky headers and excludes deferred-loading tools that will not appear in the prompt. After the response, checkResponseForCacheBreak() compares cache_read_input_tokens. The first call has no comparison baseline. A later drop must exceed 5% and be at least 2,000 tokens to trigger diagnosis; smaller drops clear pending changes. The explicit cached-microcompact deletion flag instead resets the baseline as an expected decrease. Suspicious decreases are paired with changes in model, system, tools, cache control, headers, effort, or extra body as candidate explanations. This threshold and correlation do not prove causation; provider timing, routing, or eviction may still require server-side evidence.

7. The practical rule
Prompt-cache performance follows the serialized provider request. Stable memory text, ordered tools, sticky headers, exact fork prefixes, deterministic replacements, and carefully placed compact edits all serve the same goal: avoid changing old request bytes unless the underlying information truly changed.
| Mechanism | Cache consequence |
|---|---|
| Memory | Reloaded instructions can alter an early request prefix. |
| Context management | Selection determines the stable history, suffix, and summaries. |
| Tools / MCP | Schemas and dynamic instructions need deliberate insertion points. |
| Permissions / hooks | Results and runtime state can change subsequent request material. |
| Subagents / forks | Independent context and shared-prefix execution have different request rules. |
| Resume / replacement | Restoration must preserve prior tool-result replacement decisions. |
| Compact / microcompact | A newly installed summary differs from provider-side cache edits. |
Sources
The source-code claims in this article are based on the public mirror and the linked official documentation. Provider-side KV implementation and eviction policy are not inferred from client code.
- Anthropic prompt caching docs
- Claude Code memory chapter
- Claude Code memory docs
- Cache control and TTL selection
- Dynamic tools, global strategy, and sticky headers
- Message and system cache markers
prependUserContext()getMemoryFiles()- Fork cache-prefix reuse
- Cached microcompact
- Deterministic tool-result replacement
- Cache-break detection