A coding agent often feels slow long before the model itself is the only bottleneck. Each turn has to reintroduce a working scene: base instructions, tool schemas, project rules, permission state, prior messages, tool outputs, and the new user request. In a long task, the cost of rebuilding that scene becomes visible.

The official OpenAI prompt caching guide requires a matching reusable prefix and recommends stable content before dynamic input. Eligibility, billing, TTL, and key semantics depend on the model. GPT-5.6 and later route automatically and may use prompt_cache_key for separate customer or user cache accounting; earlier models use a stable key to optimize routing. Responses reports reads under usage.input_tokens_details.cached_tokens, rather than the Chat Completions usage path.

prompt_cache_key is not a cache-hit switch. It can guide routing on earlier models or separate cache accounting on newer models; neither use repairs a changed prefix. The client controls the stability of ordered input, tool definitions, and context updates.

Source scope. OpenAI docs define the server-side prompt caching conditions: exact prefixes, eligibility threshold, routing, retention, prompt_cache_key, and cached_tokens. Public Codex source shows request construction, history persistence, compaction, and token usage accounting. This article does not infer private server serialization, machine routing, or KV-cache placement from client code.

This part follows six questions:

  1. Which view is actually cacheable?
  2. Where does Codex get prompt_cache_key, and how do sessions and internal request families determine it?
  3. How do instructions, tools, and input form a stable prefix?
  4. Why should settings, environment context, and tool output land after that prefix?
  5. Why does compaction reduce pressure while changing future cache shape?
  6. Which metrics connect cache behavior to perceived latency?

1. Cache the Model View, Not the UI Transcript

Asking whether a turn “hit the cache” is too coarse. First ask: what was sent to the provider, in what order, and through which fields? The visible UI history, the rollout persisted on disk, and the request assembled for the model are related representations, but they are not identical.

Layer Implementation Role Cache relevance
Visible history Client event conversion Shows turns, tools, hooks, events, and final answers. Useful for reading; not proof of the next model input.
Durable record Rollout / thread store Supports resume, replay, fork, and audit. Seeds reconstruction, but is not the provider cache.
Model view Prompt / Responses request Contains this turn's instructions, tools, input, and controls. The provider can only reuse this ordered request view.

Codex makes this separation visible in Prompt. It carries conversation input, model-visible tools, parallel_tool_calls, base_instructions and output schema. ModelClient::build_responses_request turns that into a Responses payload: instructions, input, tools, reasoning, text, and prompt_cache_key.

Shape-level request:
Responses request
  instructions: stable base instructions
  tools:        model-visible tool schemas
  input:        prior context + dynamic turn tail
  text:         output schema / verbosity
  prompt_cache_key: session/request-family cache key

2. Cache Keys Follow Sessions and Request Families

ModelClient is session-scoped and retains authentication, provider, and transport state across turns. That does not make every key equal to the thread id. The durable thread identity, the current runtime session, and an internal helper request family must be distinguished.

ModelClient::prompt_cache_key() returns an override first. For an internal session with a parent thread id, it returns source:parent_thread_id. Otherwise it uses responses_metadata.session_id. The request builder serializes that value, so helper requests can share a request-family key while ordinary requests follow the current session; the default is no longer simply thread_id.

The key neither chooses a message breakpoint nor guarantees a hit. Public API behavior must also be separated from the ChatGPT client adapter: responses_session_id() preserves the actual session id for non-root agents and otherwise uses the selected key for request affinity. This visible header choice does not prove one server routing algorithm across all model families.

Callers such as Guardian review sessions can supply an override through with_session_options. That groups related requests; a new user message does not need a newly invented key. Observe usage separately to establish actual reuse.

3. Stable Prefix: Instructions, Tools, Context Base

OpenAI's best practice is to put stable content first. Codex has three obvious candidates: base instructions, tool schemas, and the initial or diffed context baseline.

3.1 Instructions Stay Stable Across Ordinary Turns

For ordinary Responses, build_responses_request reads prompt.base_instructions.text into instructions. The prompt caching test suite submits two ordinary turns and checks that instructions remains unchanged. When a model path needs stable apply-patch guidance because the tool itself is absent, the test also checks that the appended instruction block remains identical.

3.2 Tools Are Prefix Material

OpenAI docs say tool definitions can be cached and count toward the eligibility threshold. Ordinary Responses converts prompt.tools into the top-level tools field; Responses Lite places AdditionalTools in input. Both use the tool router captured by the current StepContext. Parts IV and VII already showed why that list can be large: core tools, MCP tools, deferred tool search, dynamic tools, and plugin or skill supplied capabilities all feed the model-visible tool list.

A large tool list has a cost. But a stable tool list is also excellent prefix material. Unnecessary per-turn tool churn hurts cacheability before the model even looks at the new user input.

3.3 Context Base: Inject Once, Diff Later

The context-management part introduced reference_context_item. It is the baseline for later context diffs. When the baseline is still valid, Codex can append an update rather than reinjecting the whole context bundle. The test named prefixes_context_and_instructions_once_and_consistently_across_requests demonstrates the effect: the second request preserves the first request's input prefix and appends the next user message.

3.4 Request shape and WebSocket deltas are separate choices

The builder has two request shapes. Ordinary Responses sends top-level instructions and tools. Responses Lite prepends AdditionalTools and base-instruction messages to input, leaving top-level instructions empty and tools absent. Prefix item ids are stable hashes within a thread namespace. Tool eligibility and the minimum cacheable length still depend on the model.

WebSocket continuation compares the complete logical request. When non-input properties are compatible and previous input plus returned output remains a prefix, it sends only new input with previous_response_id. Compaction, configuration changes, or a prefix mismatch can require full input. A full network payload does not imply a complete cache miss.

shape-level, fields omitted:
Ordinary Responses: instructions + tools + input
Responses Lite: input = AdditionalTools + base instructions + history

Logical input: S + prior output + new tool result
Reusable WebSocket continuation: previous_response_id + new tool result
No continuation: full input
Cache result: observe usage separately from transmitted bytes

4. Dynamic Tail: Change Later, Not Earlier

Dynamic material is inevitable. Users switch settings, change cwd, change approval policy, request a different model, trigger tools, receive tool outputs, or let hooks add context. Codex does not try to remove all change. It tries to keep change behind the already-stable prefix.

Two tests make that concrete. One applies thread settings overrides and checks that prompt_cache_key stays constant while updated permissions and environment messages are appended after the existing prefix. Another applies per-turn overrides and checks the same invariant while adding a model switch and new environment context.

Change source Codex handling Protected property
New user input Append a new user message. Prior prefix is not reformatted.
Thread settings override Append settings and environment updates. Stable context remains before the change.
Tool output Record history, then normalize with for_prompt(). Call/output invariants stay intact.
Hook additional context Write a contextual fragment into model context. Policy additions become visible input changes.
Compaction Install summary or replacement history as a new baseline. Future turns get a smaller but reconstructible prefix.

5. Compaction Relieves Pressure and Changes Shape

Prompt cache does not replace the context window. Once history grows too large, Codex still has to compact. ContextManager stores token usage, estimates counts, and bumps history_version when history is rewritten. The get_context_remaining tool reports the smaller remaining allowance across the auto-compact scope and effective full window, using the shared context-window status.

The current remote compact v2 attempt clones history, trims tool outputs to fit the window, retains history metadata, and appends CompactionTrigger. It uses the shared ModelClientSession streaming path. Collection requires response.completed and exactly one compaction output before replacement history can be installed. Sending a request, receiving an output item, and installing a checkpoint are distinct states.

Compaction has two performance effects: it reduces future history pressure, and it establishes a new prefix shape. The old prefix is no longer the mainline model view; the summary, retained outputs, and reinjected context become the next cache candidate.

6. Metrics: Cached Input, Non-Cached Input, First Token

The provider reports cache behavior through usage. Codex tracks input_tokens, cached_input_tokens, output_tokens, reasoning_output_tokens, and total_tokens. At turn completion, on_task_finished compares the token snapshot from turn start with current total usage to compute this turn's input, cached input, non-cached input, output, and total tokens.

The same path records tracing fields, session telemetry histograms, analytics events, time_to_first_token_ms, and total turn duration. That matters because a high cached-token count and a fast-feeling turn are related but not equivalent. Tool loops, approvals, network retries, compaction, and output length all shape user-visible latency.

Metric Meaning Do not read it as
cached_input_tokens Input tokens served from provider cache. Tokens the model did not see.
non_cached_input() Input tokens that were not cache hits. All non-cached tokens are waste.
time_to_first_token_ms Delay until the first visible model output. A value controlled only by prompt cache.
total_tokens The turn's total token count. A guarantee of semantic completeness.

7. Common Misreadings

The API fields look small, but the runtime discipline behind them is large. These are the misreadings that tend to make performance debugging noisy.

Misreading Better reading Source check
prompt_cache_key decides the hit. The key helps routing; exact prefix still decides eligibility. Key construction and request shape are separate.
If the chat shows it, the model sees it. The model sees for_prompt() output. History normalization owns the request input.
Fewer tools are always faster. Stable tool schemas are costly but cacheable prefix material. The tools field comes from prompt.tools.
Compaction automatically improves cache. Compaction changes the future prefix and must remain reconstructible. Remote compact reuses the key while rewriting history.
A cache hit guarantees a fast turn. Tool loops, approvals, compaction, output length, and network still matter. Codex records token usage, TTFT, and duration separately.

8. Why Recovery Comes Next

Prompt cache pushes performance analysis onto model view shape. But that shape has to be reconstructed from durable history: rollout items, context updates, tool outputs, compaction replacement, rollback markers, and fork state. If that durable record cannot rebuild the same history, the next request shape drifts and cache analysis loses its anchor.

The next part returns to persistence and recovery: how Codex writes a turn into rollout, reconstructs history from RolloutItem, handles historical rollback markers, fork, compact transitions, and token usage. That connects what the user saw, what disk retained, and what the model sees on the next turn.

Source References