Start with a normal Codex session. You ask it to fix a bug in a repository. It reads AGENTS.md, searches source files, runs a test, inspects the failure, and edits the patch. Halfway through, you add a constraint: do not touch generated files. Later the thread approaches the context limit and compacts. The next morning, you resume the thread and expect the agent to continue from the right place.

A plain "append every message to a chat history" model cannot explain this workflow. Repository instructions, working directory, sandbox policy, skill metadata, tool schemas, tool outputs, images, reasoning tokens, pending user input, compaction summaries, and resume baselines are stored and used by different components. Some must be durable. Some are only model-visible for one request. Some may be truncated. Some must be re-injected from scratch if the baseline becomes unreliable.

The official Codex docs describe the product behavior: a thread is a session made from prompts, model outputs, and tool calls; as the agent works, it gathers context from files, tool output, and its running record; everything must fit in the model's context window, and long threads may be compacted. The source answers the next question: where does that information live, when does it become model-visible, and what gets installed after compaction?

OpenAI's agent-loop article makes the external constraints concrete: a model request is made from prompt, tools, input, and related request fields; long-running work also has to deal with prompt cache, the context window, and compaction. The source-level question is which of those constraints is carried by ContextManager, TurnContext, and replacement history, and which state remains durable versus turn-local.

The easy-to-miss step is what happens after a tool call. The next model request does not start from a blank page. The previous reasoning, the tool call, and the returned function_call_output are appended to the next request's input in order. The second request therefore keeps the stable prefix from the first request and adds the new tool facts at the tail.

first inference:
  instructions
  tools
  input:
    - user task
  -> reasoning
  -> function_call(call_id=A)

tool execution:
  command returns:
    function_call_output(
      call_id=A
    )

next inference:
  instructions
  tools
  input:
    - user task
    - reasoning
    - function_call(call_id=A)
    - function_call_output(
        call_id=A
      )

Complete logical input does not imply a full network payload every time. The client assembles its complete input before transport; WebSocket continuation can reuse previous_response_id and send only new input when non-input properties remain compatible and previous input plus returned items remains a prefix. Without a usable response id or matching prefix, it sends full input. Client-side reconstructability remains true, while a blanket claim that Codex never uses previous_response_id does not. Cache reuse and reduced network bytes are separate questions.

Source scope. This article uses official documentation for product behavior and the public openai/codex source for implementation details. Source links point to one fixed public snapshot. Where the source exposes only the remote-compaction selection conditions, this article describes the visible lifecycle and does not infer the server-side summarization algorithm.

The article follows six questions:

  1. In one agent loop, how do model output, tool calls, and tool results return to the next request?
  2. Why is durable history not the same thing as the prompt?
  3. When does Codex inject full initial context, and when does it emit only a diff?
  4. What does for_prompt() normalize or filter?
  5. During compaction, which history is replaced and what recovery record remains?
  6. Why do legacy rollback replay, resume, and model switches affect the next context injection?

1. Separate the four kinds of data

The first step is to stop looking for one memory object. Codex has at least four relevant kinds of data: session-scoped state, per-turn configuration, durable history, and the model request. They are not interchangeable.

TurnContext keeps turn identity, initial settings, and control state. Each sampling request captures a StepContext containing one version of model settings, environments, MCP connections, and the tool router. Changes during tool execution can affect a later step without mixing the advertised capabilities and execution settings of an already captured request.

Data Stored or generated by Model-visible? What must remain true
Session state SessionState Not as a whole Keep history, permission baselines, compaction windows, and recovery state across turns.
Turn and sampling snapshots TurnContext + StepContext Partly, through context messages and request fields Keep each sampling request and its tool execution on the same captured configuration.
Durable history ContextManager.items Only after filtering and normalization Retain recoverable API messages, tool calls, outputs, and compaction records.
Model-visible prompt for_prompt() + Prompt Yes Send a legal request view adapted to the current model and tools.

The ContextManager comment calls it a transcript of thread history, but the struct also carries history_version, token_info, and reference_context_item. That last field is the baseline used to generate model-visible settings updates. If it is missing, the next regular turn falls back to full context reinjection instead of diffing against stale state.

2. A Turn Does Not Jump Straight From User Input To Model

In run_turn, the first step is not to record the new user message. Codex first runs pre-sampling compaction. A nearby source note says this happens before context updates and the new user message are recorded, and that pending incoming items should eventually be estimated before deciding whether the thread crosses the compaction threshold.

Codex request preparation showing StepContext and history processed by for_prompt() contributing to the Prompt and model request
Request preparation has two paths: StepContext supplies the turn configuration and tool environment, while for_prompt() prepares model-visible history. Together they contribute to the Prompt and model request.

After the pre-sampling compaction check, run_turn captures the first StepContext, records environment updates, and prepares skills and user input. Before sampling it clones history and applies for_prompt() using that step’s model input modalities. Later steps can capture refreshed settings and capabilities; a turn-wide object does not freeze the tool catalog for every request forever.

turn starts
  -> maybe compact old history
  -> capture StepContext for this sampling request
  -> inject full context or WorldState changes
  -> load skills/plugins chosen for this turn
  -> record user input and additional context
  -> clone history
  -> for_prompt(input_modalities)
  -> build Prompt with tools and base instructions
  -> stream model response

The turn loop also treats pending input carefully. A source comment says the turn-scoped ModelClientSession is reused across retries, while pending input is deliberately not drained at the very start of a turn or immediately after auto-compaction. That ordering preserves continuation semantics. If a steer is inserted in the wrong place, the model may resume the wrong work.

2.1 Initial Context Is Synthesized Runtime State

Full initial context comes from build_initial_context_with_world_state. It collects model-switch guidance, permission instructions, developer instructions, collaboration mode, realtime state, personality, apps instructions, available skills, plugins, extension fragments, token budget context, environment context, and user instructions. These become developer or contextual user messages. They are not the user chat message.

The structured baseline is stored as a TurnContextItem. That lets Codex later decide whether a new turn can emit only a settings diff or must re-inject the full context bundle.

2.2 Compare WorldState against the established baseline

Suppose the bug-fix task changes from workspace-write to read-only. The runtime must distinguish the old environment, the current environment, and what it has already told the model. It builds a WorldState and keeps a world_state_baseline in history. The separate reference_context_item remains the durable turn-settings baseline.

record_context_updates_and_set_reference_context_item builds WorldState for the captured step. With no reference baseline it emits full initial context plus a full WorldState record. Otherwise update_world_state produces changed fragments, with turn contributions added when turn settings change.

shape-level:
step = capture(settings + environments + MCP + tools)
world = build_world_state_for_step(step)
if reference_context_item is None:
    history += full_initial_context(step, world)
    rollout += full_world_state
else:
    history += visible_world_state_changes
    rollout += world_state_changes_if_any
persist TurnContextItem when required
advance the reference baseline

The order matters: record model-visible messages first, persist the WorldState that produced them next, then write the required TurnContextItem. A snapshot-only change with no visible fragment or turn-setting change does not require another TurnContext record. Recovery must not restore a baseline that claims the model received an update whose message was never recorded.

Both full reinjection and incremental updates remain. The current diff mechanism is WorldState-based, rather than the former fixed settings-field list. Missing baselines require complete context; trusted baselines let new facts join the history tail.

3. for_prompt() filters and normalizes history

Before sampling, Codex does not send raw ContextManager.items. It calls for_prompt(), which normalizes history and returns the items suitable for the current model request. Durable history stays recoverable, while the generated request must satisfy the current model's capabilities.

Codex prompt preparation showing durable history filtered and normalized into the model request
Durable history supports recovery. for_prompt() normalizes calls and outputs, filters unsupported modalities, and removes local history metadata before building the model request.

normalize_history supplies missing call outputs, removes orphaned paired outputs, and strips unsupported image and audio content. Function/custom outputs need their matching calls; named external outputs may stand alone. for_prompt() then unwraps ResponseItemEnvelope into API ResponseItem values. History-only provenance and truncation-budget metadata stay local instead of becoming the API envelope.

durable history, shape-level:
  user text
  function_call(call_id=A)
  function_call_output(call_id=A)
  orphan function_call_output(call_id=old)
  user image + audio

model-visible view for a text-only model:
  user text
  function_call(call_id=A)
  function_call_output(call_id=A)
  user text with unsupported image/audio stripped

The same layer also truncates large tool outputs. The record_annotated_items path applies a truncation policy to function output payloads. Then build_prompt adds model-visible tool specs, parallel tool call support, base instructions and output schema. Tools are not stored as normal chat messages, but their schemas still enter the model-visible request.

3.1 Before compaction, combine server counts and local estimates

for_prompt() explains what the model sees. The next question is when the stored history is large enough to compact. Token accounting is not just the server's last usage value. get_total_token_usage adds the latest server usage, local items appended after the last model-generated item, and, when the server has not included them, estimated non-last reasoning tokens. The local estimate is coarse: the estimator is byte-based, not tokenizer-accurate.

context_window_token_status uses that active context count together with the configured auto-compact scope, a scope limit, and the full context window limit. In BodyAfterPrefix mode, it subtracts a prefill baseline held by AutoCompactWindow. Server-observed input tokens replace an estimated baseline when they become available.

4. Compaction Installs Replacement History

Compaction is easy to under-describe as "make a summary." In Codex, the visible mechanism is a recoverable history rewrite: generate a handoff summary, preserve recent real user messages, install replacement history, advance the auto-compact window, and decide how initial context should be represented after the rewrite.

Codex compaction replacement showing old history summarized into replacement history and written to rollout
The compacted result is replacement history, not a summary dangling beside the old transcript.

4.1 Pre-Turn And Mid-Turn Compaction Differ

The pre-turn path is run_pre_sampling_compact. It checks whether compaction is required before the new user turn is sampled. There are also pre-turn compactions for compaction compatibility hash changes and model downshifts to smaller context windows.

The mid-turn path happens after sampling when the model still needs a follow-up or pending input exists. The post-sampling state collection checks token status and, if the context limit has been reached while follow-up work is needed, runs auto-compaction with InitialContextInjection::BeforeLastUserMessage.

The enum comment in InitialContextInjection explains the split. Pre-turn and manual compactions use DoNotInject, replace history with a summary, and clear reference_context_item; the next regular turn will fully reinject initial context. Mid-turn compaction injects initial context into replacement history just above the last real user message, preserving the model's expected continuation shape.

4.2 The summary prompt defines what a future turn must retain

Local inline compaction starts in run_inline_auto_compact_task. It turns turn_context.compact_prompt() into user input. The default compact prompt asks the model to create a handoff summary for another LLM that will resume the task: current progress, key decisions, constraints, remaining steps, and critical references.

In run_compact_task_inner_impl, Codex clones history, appends the compact prompt to that compact request, runs the model to completion, extracts the last assistant message as the summary suffix, collects real user messages, builds replacement history, advances the auto-compact window id, installs the compacted history, and recomputes token usage.

old history:
  initial context
  user A
  model/tool work
  user B
  model/tool work

replacement history, shape-level:
  recent real user messages, capped
  user message: SUMMARY_PREFIX + handoff summary

build_compacted_history retains recent real user messages, capped by a 20,000-token approximate budget, then appends the summary as a user message. When mid-turn compaction needs initial context, the insertion helper places canonical initial context before the last real user message when possible, or before the summary or compaction item as a fallback so the compacted item remains last.

4.3 Remote Compact V2 Has an Explicit Completion Boundary

The handoff summary above describes local compaction. A remote v2 attempt also clones history, but uses for_prompt_annotated() to retain metadata needed by this specialized request and appends CompactionTrigger. It streams through the shared ModelClientSession and receives a compaction item, rather than assuming a readable assistant summary.

Collection requires response.completed and exactly one compaction output before installing replacement history. Dispatching a request, receiving an item, and installing a checkpoint are distinct states. Failures follow bounded retries and conditional fallback; an unfinished stream is not a successful compact. How the server generates the content is outside the public client implementation.

4.4 Replacement Is Persisted To Rollout

First place rollout: it is the replayable durable record layer, not the current in-memory history. Compaction changes durable state too. replace_compacted_history replaces in-memory history, persists RolloutItem::Compacted, persists a full WorldState when supplied and a RolloutItem::TurnContext when a new reference baseline was established. A settings persistence lock and a current-settings event prevent an old turn snapshot from overwriting accepted updates. It then queues a session-start hook for the compacted state. That makes compaction part of resume and fork semantics, not just a temporary token-saving step.

5. Recovery and New Windows Invalidate Baselines

The difficult part is not appending a new turn. It is keeping the baseline valid after legacy rollback replay, resume, fork, and compaction. Codex does not keep diffing against a baseline it can no longer reconstruct. When the surviving history no longer contains the bundle that established the baseline, Codex clears the baseline.

drop_last_n_user_turns now exists to replay rollback markers in historical rollouts, not as an active rollback operation in the current submission protocol. It still removes context updates adjacent to the cut. If this trims a mixed initial-context developer bundle, the trim logic clears reference_context_item, forcing complete reinjection on the next real turn.

Resume reconstruction treats history, previous settings, reference context, and window id as a single recovery group. apply_rollout_reconstruction installs the reconstructed history and baseline. In BodyAfterPrefix mode, it also estimates prefix tokens for the recovered compaction window.

start_new_context_window advances the window number and ids, constructs initial context from the captured step and WorldState, and installs an empty-message CompactedItem with a fresh baseline. When the client-developer-message retention feature is enabled, it also preserves those messages within a budget. This path creates a fresh prefix without generating a semantic summary of previous work.

6. Common Misreadings

The mechanism becomes easier to reuse once the common shortcuts are made explicit.

Misreading What the source shows What breaks if implemented literally
History is the prompt ContextManager stores durable items; for_prompt() derives a request view. Orphan outputs, unsupported images, or bad call/output pairs can poison the request.
Initial context is repeated every turn A valid baseline allows settings diffs; missing or unreliable baselines trigger full reinjection. Either waste tokens or, worse, diff against stale runtime state.
Compaction is just a summary Compaction creates a handoff summary and installs replacement history in rollout. Resume cannot know which history was replaced or how to restore the baseline.
Token usage is only server usage Codex combines server usage, local appended items, reasoning estimates, and window baselines. Large local additions can push a request over the window without being noticed early.
Rollback just deletes user messages Rollback also trims nearby context updates and may clear the reference baseline. The next turn can send diffs based on an initial context bundle that no longer survives.

7. Runtime Rules To Carry Forward

The Codex implementation suggests a small set of general rules for coding-agent runtimes.

State type Where it belongs Before model visibility What must remain true
Conversation and tool evidence Durable history / rollout Filter by model capability and preserve valid call/output pairs. Recovery and later inspection retain complete records.
Runtime environment and permissions Turn context plus reference baseline Inject fully first, then emit trusted diffs. The model sees current settings without stale configuration.
Tool capability Tool router / prompt request Expose as model-visible specs, not as normal chat text. Tool schemas match the tools the runtime can execute.
Long-thread compaction Replacement history plus compacted rollout item Install a handoff summary and reinsert initial context when needed. Long work can continue while acknowledging compaction is lossy.
Context-window accounting token_info plus auto_compact_window Combine server usage with local estimates. Neither the server count nor the local estimate is trusted alone.

That is why context management is a natural first mechanism article after the overview. Tool calling, permissions, multi-agent work, skills, and plugins all come back to the same discipline: decide where each piece of state is stored, then decide which parts belong in the next model request.

Sources