Prompt caching works when consecutive provider requests share a long, byte-stable prefix. A coding agent naturally challenges that requirement: tools connect, memory changes, results grow, and child tasks need different instructions. Claude Code therefore controls where changing data appears and records which request fields changed when reuse drops.

Reading goal: follow cache markers from request assembly, then inspect the designs that preserve old bytes—session-fixed options, fork suffixes, cached microcompact, deterministic tool-result replacement, and cache-break detection.

stable prefix = system + tools + headers + earlier messages
dynamic tail = current user turn + new tool output

addCacheBreakpoints(messages)
  -> place one cache_control marker
  -> keep volatile blocks after the reusable prefix

content replacement / cache_edits
  -> shorten future provider view
  -> keep local transcript and recovery state separate
Source shape: a marker identifies a reusable prefix, but only stable request content makes that prefix reusable.

This article verifies the public mirror’s fixed March 31, 2026 snapshot at one immutable commit. The mirror is not Anthropic’s official source repository and does not establish later product behavior. Client implementation, official API contracts, and provider-side inference remain separate.

1. Cache markers are only part of the request

getCacheControl() produces an ephemeral marker and may add a one-hour TTL or global scope. The conditional scope: "global" field is visible in this client; it is not a universal public API contract. buildSystemPromptBlocks() places a marker on the system blocks, while addCacheBreakpoints() adds at most one message-level marker. A normal request marks its final message. With skipCacheWrite, Claude Code moves the marker to the preceding message so a short-lived fork can read the shared parent prefix without writing its private suffix.

The TTL choice is itself kept stable. should1hCacheTTL() first returns true for Bedrock with an explicit ENABLE_PROMPT_CACHING_1H_BEDROCK opt-in. Other paths latch user eligibility and allowlist state in bootstrap state on first use instead of reevaluating a changing experiment or overage condition on every turn. Otherwise the marker could alternate between the default TTL and one hour even when all prompt text remained identical. Source comments treat that as a cache risk; this does not prove that every provider uses identical TTL cache-key rules.

main turn:
  S = system + tools + stable history
  U = current user turn; marker after S + U
fire-and-forget fork:
  S = shared parent prefix
  F = fork-only task; marker after S, before F
example F:
  a compact summary instruction, still sent as uncached input
The parent will not resume from S + F. Moving the marker avoids writing that side-task tail; it does not omit the task from the request.

2. Dynamic material is kept out of the stable prefix when possible

Deferred tools and late Chrome instructions illustrate the problem. Prepending a fresh meta message whenever tools appear would change every later byte. When delta attachments are enabled, Claude Code emits deferred_tools_delta or mcp_instructions_delta instead of repeatedly changing earlier messages or the system prompt.

Global caching is also disabled for the affected system path when user-specific dynamic MCP tools are rendered. A per-user tool list should not be treated as globally stable content. Advisor tools are appended after existing schemas so toggling the advisor disturbs the tail instead of reordering the whole tool list.

Memory follows the same rule. CLAUDE.md and rules enter an early meta user message through prependUserContext(); auto-memory content may enter user context or relevant-memory attachments. If the final text is unchanged, it can remain reusable. If a memory file or index changes, the provider sees different prefix bytes even though the feature is still called “memory.”

Beta headers can change cache identity without changing prompt text. Claude Code therefore keeps AFK, fast mode, cache editing, and thinking-clear headers sticky once first sent during a session. The header can stay present while current behavior—such as whether fast speed is active after cooldown—continues to change through another request field.

Changing inputClient handlingResult
Deferred toolsSend a delta attachment instead of prepending a new message.Earlier message bytes stay in place.
Late Chrome instructionsMove them out of the system prompt when delta support is available.The system blocks remain stable.
User-specific MCP toolsAvoid global caching for the affected content.Private dynamic schemas are not mislabeled as globally reusable.
Beta headersKeep a sent header present for the rest of the session.Mid-session feature changes do not silently change cache identity.

3. Fork children copy the parent prefix exactly

A fork child inherits the parent's rendered system prompt, exact tool array, messages, non-interactive setting, and thinking configuration. buildForkedMessages() copies the parent assistant message and placeholder tool results, then appends a directive and task for that child.

This is why a fork is intentionally unlike a normal subagent. Rebuilding worker tools under another permission mode could change serialized tool definitions at the first schema. A new system prompt or thinking setting would do the same. useExactTools: true preserves those request fields instead.

Two implementations should remain distinct: the preceding explanation describes an AgentTool fork child, while the cover focuses on one-off internal runForkedAgent() calls such as compact summaries. Both preserve shared parameters; transcript writing, tool restrictions, and skipCacheWrite depend on the caller, so the cover is not a rule for every subagent.

4. Full compact and cached microcompact take different approaches

Full compact summarizes older work and creates a new prefix for later turns. The summary request can still reuse the current session's cached prefix through its fork-like path. Once the summary is installed, however, future requests naturally use the new compacted history.

Cached microcompact tries to shorten old tool results without rewriting local messages. When the service cache is likely still useful and the current model and thread support the feature, cachedMicrocompactPath() leaves message content untouched, queues cache_edits; the API adapter later adds cache_reference fields to tool-result blocks. Queuing is only a deletion request: cache_deleted_input_tokens from the response, minus its earlier cumulative value, provides the reported per-operation result. If the cache is probably expired, Claude Code can clear old result content directly because the full prefix will be written again anyway.

Pending edits are consumed before paramsFromContext can be called by logging, retry, or fallback code. Consuming the same edit twice during one turn could make retry requests differ from the original attempt.

cache-edit state transition (simplified):
old tool_result: tool_use_id = T, cache_reference = T, original content
queue this turn: cache_edits → delete(cache_reference = T)
API assembly: insert and pin edits; retries reuse the same batch
response: cumulative cache_deleted_input_tokens - prior total = reported deletion

Local transcript retains the result; queuing does not mean deletion completed.
This is the visible deletion-reference and response-counter contract, not a claim about how the provider recomputes or retains suffix KV states.

5. Tool-result replacement must repeat the same decision

Large tool results can be stored locally and represented by a preview in the model request. The important cache rule is consistency: an old result cannot be full text on one turn and a preview on another, and a regenerated preview cannot differ by a few bytes.

ContentReplacementState records ids that have already been evaluated and the exact preview used for replacements. Later turns reapply the preview through a map lookup without reading the file again. Previously seen results that were not replaced are frozen in that state; only fresh results participate in the next budget decision. Forks that share a prefix clone the replacement state so parent and child serialize the same old results.

The update around asynchronous file persistence is also kept atomic. If another reader observed an id in seenIds before its replacement appeared, it could send full content while the main thread sends the preview. The same historical result would then have two different wire representations.

replacement decisions (conceptual shape):
seenIds = {A, B}
replacements = {B: "saved result location + identical preview"}
next turn:
A was seen, no replacement → retain original; do not decide again
B has a replacement → reuse the exact string without file I/O
C was never seen → evaluate against this turn's budget
resume restores decisions; a cache-sharing fork copies them
The preview is deterministic persisted-result text, not a newly generated summary or an invented automatically resolved URI.

6. Cache-break detection compares requests with response counters

PreviousState records client-visible inputs that can affect reuse: hashes of system blocks, tools, and cache control; tool names; model; fast mode; global strategy; beta headers; overage; microcompact state; effort; extra body; and prior cache-read tokens. Separate tracking keys keep unrelated agents and SDK calls from being compared, while compact shares the main REPL key because its summary path reuses those parameters.

Before the request, recordPromptState() captures the final sticky headers and excludes deferred-loading tools that will not appear in the prompt. After the response, checkResponseForCacheBreak() compares cache_read_input_tokens. The first call has no comparison baseline. A later drop must exceed 5% and be at least 2,000 tokens to trigger diagnosis; smaller drops clear pending changes. The explicit cached-microcompact deletion flag instead resets the baseline as an expected decrease. Suspicious decreases are paired with changes in model, system, tools, cache control, headers, effort, or extra body as candidate explanations. This threshold and correlation do not prove causation; provider timing, routing, or eviction may still require server-side evidence.

Previous request fields and cache-read tokens compared with the next response to identify a suspected cache break
The detector separates expected deletion from suspicious decreases and offers investigation clues, not proof of a server-side cache failure.

7. The practical rule

Prompt-cache performance follows the serialized provider request. Stable memory text, ordered tools, sticky headers, exact fork prefixes, deterministic replacements, and carefully placed compact edits all serve the same goal: avoid changing old request bytes unless the underlying information truly changed.

MechanismCache consequence
MemoryReloaded instructions can alter an early request prefix.
Context managementSelection determines the stable history, suffix, and summaries.
Tools / MCPSchemas and dynamic instructions need deliberate insertion points.
Permissions / hooksResults and runtime state can change subsequent request material.
Subagents / forksIndependent context and shared-prefix execution have different request rules.
Resume / replacementRestoration must preserve prior tool-result replacement decisions.
Compact / microcompactA newly installed summary differs from provider-side cache edits.

Sources

The source-code claims in this article are based on the public mirror and the linked official documentation. Provider-side KV implementation and eviction policy are not inferred from client code.