Imagine a long coding task. The agent has read files, run tools, compacted history, changed model settings, and accumulated enough context that a simple transcript is no longer the whole story. If the process stops, the next runtime must recover more than visible messages. It needs the model-visible history, the current context baseline, session identity, lineage, and the event facts that a client will project back to the user.
That is the job of rollout. Codex writes an append-only JSONL
stream of typed RolloutItems, keeps write ordering
behind a recorder, and reconstructs the runtime state by replaying
those items through session logic. The recovery path is not a
convenience wrapper around chat history; it is part of runtime execution.
Rollout is an append-only stream of typed JSONL records. It persists enough information to rebuild model history, context baselines, token usage, state after legacy rollback markers, and fork starting points.
Source scope. This article describes public openai/codex source: rollout schema, recorder behavior, session initialization, resume, fork, legacy rollback compatibility, reconstruction, and tests. It does not infer private backend storage behavior or treat local rollout paths as a public interface.
This part follows six questions:
- Why is continuing a thread harder than showing old messages?
- What does each
RolloutItemvariant own? - How does the writer queue, persist, flush, and recover from failed writes?
- Why do resume and fork first become
InitialHistory? - Why does reconstruction scan backward before replaying forward?
- How do legacy rollback records, compaction, and prompt caching depend on these records?
1. Recovery Needs Four Kinds of State, Not Just Messages
A client needs visible facts. The next model request needs ordered input. The runtime also needs a context baseline: working directory, sandbox, permission profile, model, collaboration mode, and the latest full-context checkpoint. Those requirements overlap, but they cannot collapse into one saved string.
| Recovered state | Visible layer | Runtime requirement | Rollout role |
|---|---|---|---|
| Client state | User messages, tool progress, warnings, and legacy rollback events. | Event order, turn start/end, resume/fork source. | Persist EventMsg for client event replay. |
| Model history | Messages and tool results the model should see. | Compacted replacement history and surviving suffix. | Persist ResponseItem and CompactedItem. |
| Context baseline | cwd, sandbox, permissions, model, collaboration mode. | The last durable full-context baseline. | Persist TurnContextItem and WorldState. |
| Thread identity | Thread id, source, parent and fork lineage. | Resume keeps an id; fork creates a new id with lineage. | Persist SessionMeta. |
This is why rollout connects earlier parts of the series. Client event conversion, multi-agent spawning, context management, and prompt caching all rely on the same durable records.
2. Rollout Types Preserve More Than Messages
RolloutItem now lives in the history crate and has more than five variants. Alongside SessionMeta, ResponseItem, Compacted, TurnContext, and EventMsg, it includes WorldState, TokenUsageRecord, RetainedContext, inter-agent communication, and model-invisible realtime presentation facts. Each consumer interprets the records it needs; persistence does not make every record model-visible.
This shape-level fragment omits required fields to focus on the envelope. RolloutLine keeps timestamp and optional ordinal at the top level and flattens type/payload. The ResponseItem wire representation can store metadata beside the payload. Reconstruction dispatches on top-level type, not a nested item.type.
{"timestamp":"...","ordinal":1,"type":"session_meta","payload":{"id":"t1"}}
{"timestamp":"...","ordinal":2,"type":"world_state","payload":{"full":true,"state":{}}}
{"timestamp":"...","ordinal":3,"type":"response_item","payload":{"type":"function_call","call_id":"A","name":"shell","arguments":"{}"},"metadata":{"client_authored":false}}
{"timestamp":"...","ordinal":4,"type":"compacted","payload":{"message":"...","window_number":1,"replacement_history":[]}}
TurnContextItem preserves recoverable settings, while WorldState stores full environment snapshots or changes. The write path records model-visible context before WorldState and the required TurnContextItem. A bare settings object does not prove its context reached history; recovery also looks for a real input boundary or a valid full WorldState after compaction.
CompactedItem can hold replacement history, retained context, window numbers and ids, and the latest token usage record. The replacement gives the model-history base; the remaining records let recovery restore settings, authorization-relevant facts, and token totals without indefinitely scanning older history. Fields are optional, so older records still need compatibility handling.
3. Writes Queue Before File Writes Complete
New and resumed sessions initialize the recorder differently. A
new session precomputes rollout path and session metadata, but it
can defer file creation until persist(). A resumed
session opens the existing rollout for append immediately.
Actual writes go through a background writer task. The recorder
sends AddItems, Persist,
Flush, and Shutdown commands. Items live
in pending_items until they are written successfully.
If I/O fails, the writer drops the file handle, keeps the unwritten
suffix, and retries after reopening at the next persist or flush command.
writer discipline:
record_canonical_items(items) -> queue AddItems
persist() -> materialize file + write pending
flush() -> wait for preceding file writes (not fsync)
shutdown() -> final drain before exit
Preparing a fork from an active source thread flushes before reading so queued records are not omitted. The completion boundary is file writes and flush(): the writer does not call fsync/sync_all here. Queue acceptance, completed writes, and power-loss-safe stable storage are different guarantees.
4. Resume and Fork Enter Through InitialHistory
Session startup normalizes all starting points into
InitialHistory: New,
Cleared, Resumed(ResumedHistory), and
Forked(Vec<RolloutItem>). New and cleared sessions
defer initial context insertion until the first real turn.
Resumed and forked sessions create a default turn context and call
apply_rollout_reconstruction.
| Start | Thread id | History source | Startup behavior |
|---|---|---|---|
New / Cleared |
Fresh id | No prior rollout. | Defer initial context to the first real turn. |
Resumed |
Existing id | Recoverable history loaded through the thread store. | Rebuild history, settings, and token usage. |
Forked |
Fresh id | Snapshot items from the source thread. | Rebuild history; persist a copy or a parent-history reference. |
Resume also warns when the last recorded model differs from the current model. Reconstruction can restore the shape of history, but changing models can still change context windows, cache behavior, and runtime performance.
5. Reconstruction Scans Backward, Then Replays Forward
reconstruct_history_from_rollout does not replay the
whole file from the first line. It first scans newest-to-oldest.
The scan looks for the newest surviving replacement-history
checkpoint, previous turn settings, reference context item, and
window id. Once those are known, older items cannot affect the
rebuilt state.
Historical rollback markers become a counter during this reverse pass.
ThreadRolledBack means “drop the newest N real user
turns.” In reverse, that means skipping the next N finalized turn
segments that actually contain a user message. Tests cover the
distinction: standalone task turns should not consume rollback
skips.
After selecting the newest surviving checkpoint, forward replay installs its replacement history and replays the suffix: annotated response items, inter-agent messages, retained context, and legacy compaction/rollback records. A newer replacement checkpoint in that suffix may belong to a rolled-back turn. It must not blindly overwrite the selected base; original records preserve the boundaries needed by compatibility replay.
reconstruction:
reverse scan
find newest surviving replacement_history
recover previous_turn_settings
recover reference_context_item
account for ThreadRolledBack markers
forward replay
seed ContextManager from replacement_history
append surviving ResponseItem suffix
apply rollback markers to the rebuilt history
The shape matters. The reverse pass avoids interpreting old log prefixes after a surviving checkpoint. The forward pass preserves the exact ordering semantics of the live tail.
6. Legacy Rollback Records Limit Bounded Reads
The current submission protocol no longer exposes the former ThreadRollback operation. ThreadRolledBack remains in persistence and recovery for historical records. The drop_last_n_user_turns comment explicitly limits it to that replay role. Compatibility code does not establish a current callable rollback API.
ModelContextScan supports bounded reverse reads of the latest model context. It needs both a checkpoint containing replacement_history and window_number and a turn establishing a valid settings baseline. A raw role=user item is insufficient: contextual fragments can use that role too. Legacy rollback markers or compacted records missing either required field force a scan back to the beginning.
shape-level read decision:
Complete checkpoint + settings baseline -> bounded scan may stop
Legacy rollback marker -> scan to the beginning
Legacy compact missing required fields -> scan to the beginning
Contextual user fragment alone -> no proven user-turn boundary
The local thread store actually feeds rollout segments through this scanner. Core reconstruction then interprets the supplied item sequence. Storage selection and semantic replay are separate responsibilities; a reverse algorithm does not prove that every recovery entry avoids reading a complete file.
7. Fork Copies Replayable History
Fork shares the recovery machinery with resume, but lineage is
different. Resume keeps the old thread id. Fork creates a new id
while recording forked_from_thread_id. The manager
reads rollout-backed history, shapes it with
fork_history_from_snapshot, and starts a new thread.
Fork persistence has two strategies. Ordinary forks use ForkPersistence::Copied; prepared forks may use Referenced with a history_base and inherited item count. The child still receives reconstructable model context without necessarily rewriting the parent’s entire history into its own rollout. Forking is therefore not universally a full JSONL copy.
8. The Prompt Cache Connection
Part IX argued that prompt cache hits depend on stable request
shape. Recovery preserves the precondition for that stability.
TurnContextItem tells the runtime what context has
already been established. CompactedItem.replacement_history
gives history a new base. Rollback markers trim the tail
consistently.
If recovery were based only on visible transcript text, Codex could re-inject full context as repeated diffs or carry rolled-back tool output into the next model view. That would be a semantic bug and a prompt-cache-shape bug. The compact/resume/fork tests assert that compacted input prefixes remain predictable across recovery.
9. What to Carry Forward
Rollout/recovery gives a useful agent-design rule: a long-running agent should persist replayable facts, not only final state snapshots.
| Rule | Codex mechanism | Result |
|---|---|---|
| Append facts | RolloutLine timestamp plus typed item. |
Recovery and audit read the same records, including legacy rollback compatibility markers. |
| Separate record types | History, context, session meta, and events are distinct. | Client state and model view can rebuild differently. |
| Checkpoint history | replacement_history after compaction. |
Long threads can resume from a newer base. |
| Use semantic markers | ThreadRolledBack as an appended event. |
Old facts remain; replay computes the current world. |
| Flush before snapshots | Fork flushes the active source before reading; shutdown drains pending writes. | Snapshots are based on persisted facts. |
The internal execution path is now almost complete: requests enter the runtime, context forms the model view, tools create side effects, events go to clients, rollout persists the records, and recovery turns those records back into a usable next turn. The next part should look at the public edge: how SDK and app-server callers enter the same runtime.