The first mistake is to treat memory as a larger chat transcript.
That would make every future request carry more old text and more
ambiguity. Codex takes a different route. The official Memories
documentation describes useful context from earlier threads, and
the source splits the implementation into memories/read
and memories/write: one path brings existing memory
into the current thread, while the other turns historical rollouts
into local material that future threads can consult.
Codex extracts high-value lessons from eligible rollouts in the
background, stores
stage outputs in state, consolidates them into
~/.codex/memories/, and later exposes them through a
read path with citations and filters for threads that used unstable
external material. It reduces repeated
explanation, but it does not take over the jobs of rules files,
rollout recovery records, or the context window.
Reading goal.
This chapter follows one lifecycle: which threads may generate
memories, how the background writer extracts and consolidates
them, how a future turn reads them, and how AGENTS.md,
rollout, Memories, and Chronicle differ.
By the end, you should know where each durable fact belongs.
Source scope.
Product behavior comes from OpenAI's Codex Memories, AGENTS.md,
and Chronicle docs. Feature switches, configuration, read/write
pipelines, citations, and pollution handling come from public
openai/codex source. This article does not inspect or
cite any local ~/.codex/memories contents, and it
does not infer private service behavior.
1. Separate the Five Things That Can Feel Like Memory
Several Codex features can make the product feel like it
remembers. TurnContext, covered in Part II, is the
current runtime snapshot. Rollout, covered in Part X, is durable
execution evidence. AGENTS.md is instruction material
loaded before work starts. Memories form a separate layer: reusable
lessons extracted from previous work and consulted later when they
are likely to help.
| Feature | Implementation | Lifetime | Best fit |
|---|---|---|---|
| Current context | TurnContext and model-visible view. |
One turn, projected and compacted as needed. | Inputs, tool results, environment, and instructions the model must see now. |
| Project rules | AGENTS.md / checked-in docs. |
Read when entering the repository. | Rules, commands, permissions, and delivery requirements that must reliably apply. |
| Execution evidence | Rollout / state DB. | Written as a thread runs; reused for resume, fork, audit, and recovery. | What happened, which tool output exists, and where recovery should restart. |
| Memories | codex-memories-read / codex-memories-write. |
Generated in the background, stored locally, selectively read by future turns. | Stable preferences, recurring workflows, project habits, known pitfalls, and high-leverage steps. |
| Chronicle | Opt-in Codex app research preview. | Uses recent screen context to help memory building. | Clues about the tools, documents, dashboards, or workflows currently on screen. |
Required rules should stay in rules files. Replay evidence should stay in rollout. Memory earns its place when it reduces repeated explanation in future work without pretending to be the only source of truth.
2. Feature Switches Decide Whether Memory Runs
The official docs start with the important default: Memories are
off by default. The source matches that documented behavior. The feature spec
for Feature::MemoryTool uses the key
memories, marks it Stable, and sets
default_enabled to false. Configuration
therefore has two layers: the feature flag decides whether the
system exists, while the [memories] table decides
generation, use, external-context handling, and model overrides.
[features]
memories = true
[memories]
version = "v1"
dual_write = false
generate_memories = true
use_memories = true
disable_on_external_context = true
The three memory-specific switches answer different questions.
generate_memories controls whether newly created
threads can become future memory inputs. use_memories
controls whether existing memory instructions are injected into
the current thread. disable_on_external_context can
keep threads that used web search, tool search, or memory-polluting
MCP servers out of memory generation. In the TUI, /memories
follows the same split: if the feature is off, it prompts for
enablement; once enabled, it exposes use/generate settings for the
thread.
| Switch | Controls | Source anchor |
|---|---|---|
features.memories |
Whether Memories is enabled at all. | Feature::MemoryTool, default off. |
memories.generate_memories |
Whether a new thread may become a memory input. | Thread creation writes a memory_mode. |
memories.use_memories |
Whether an active thread can use existing memories. | The read extension injects developer instructions only when enabled. |
memories.disable_on_external_context |
Whether external context can mark a thread polluted. | Web/tool search and MCP paths can set memory_mode = polluted. |
3. Memory Writing Runs in the Background
Immediate summarization would include unfinished work, user corrections, and transient errors. Current app-server attempts start_memories_startup_task only when turn/start has nonempty user input, core starts a new turn, and the primary environment is configured. Steering an existing turn or submitting a standalone tool output does not trigger this branch. The background function then skips ephemeral sessions, disabled features, and non-root agents, and checks memory-store availability.
The background task normally runs the selected version; dual_write = true starts separate v1 and v2 pipelines. Each prepares its root and extension instructions, prunes, checks quotas, and runs Phase 1 then Phase 2. Version roots are memories/ and memories_v2/, with v1 still the default. Dual writing does not inject both summaries: version selects one namespace for reading. The example also explicitly enables disable_on_external_context; its source default is false.
turn/start: nonempty user input + new turn + primary environment ready
-> start_memories_startup_task
-> skip checks: ephemeral / feature off / sub-agent / no state DB
-> selected version (or both when dual_write)
-> create version-specific memories root
-> prune
-> rate-limit guard
-> Phase 1 extraction
-> Phase 2 consolidation
This placement is also why the memory chapter belongs after the rollout chapter. The write path selects from state-backed rollout records. Without those rollout records, memory would look like a generic chat summarizer; with rollout in view, it becomes a later extraction layer over execution evidence.
3.1 How One Reusable Lesson Moves
Use a concrete lesson: when article flow changes in a bilingual repository, the matching language pages should be updated together and the mobile layout checked before shipping. If that lesson repeats, the writer does not paste the old conversation into future context. It moves the lesson through several shrinking shapes.
| Stage | Unit at that moment | Owner and visibility | Output |
|---|---|---|---|
| Candidate rollout | Old user request, file diff, and validation results. | The state DB selects by idle window, age, and memory mode. | A job Phase 1 may claim. |
| Phase 1 | A reusable lesson extracted from old evidence. | The background writer reads the rollout as data. | raw_memory, rollout_summary, and rollout_slug. |
| Phase 2 | A workspace patch merged from several raw memories. | An internal consolidation agent maintains its version’s workspace; isolation depends on the parent permission profile. | memory_summary.md, MEMORY.md, summaries, or skills. |
| Future read | A memory entry matched by the current task. | The foreground agent opens a small number of files under the read-path prompt. | A hidden citation records which memory was used. |
4. Phase 1 Turns Rollout into Raw Memory
Phase 1 extracts structured memory from recent, sufficiently idle threads eligible for generation. The candidate query excludes the current thread and memory_mode != enabled, applies age, idle and source filters, and bounds scan/claim counts. It no longer requires history_mode = legacy; paginated history is eligible too. The example and file table below follow default v1. V2 uses a separate tiered-input serializer and extraction template, so their record shapes should not be conflated.
The prompt template adds a second check. If a rollout has no reusable insight, the memory writer may return all-empty fields. High-signal memories are stable user preferences, high-leverage procedures, task maps, decision triggers, and reliable environment facts. Routine status notes should not occupy future attention.
{
"raw_memory": "When changing article flow in this repository, update matching zh/en pages and run a mobile visual check before shipping.",
"rollout_summary": "A prior edit improved a Codex article only after the bilingual page and narrow mobile layout were checked together.",
"rollout_slug": "bilingual-article-visual-check"
}
The important word is evidence-based. The Phase 1 prompt treats rollout text and tool output as data, not instructions. It forbids storing secrets, discourages large output copies, and prefers no-op when no future agent would plausibly act better because of the saved memory.
5. Phase 2 Consolidates a Local Memory Workspace
Phase 2 turns stage-1 outputs into the file-based memory
workspace. The README describes a very concrete sequence: claim a
global lock, select bounded stage-1 inputs by usage and recency,
sync raw_memories.md and rollout_summaries/
into ~/.codex/memories/, maintain a git baseline for
the workspace diff, and exit if nothing changed.
When the workspace changes, Codex spawns an internal consolidation agent with approvals, memory regeneration, and collaboration disabled. Permission inheritance branches: a Codex-managed parent produces a memory-root write sandbox with no network; a Disabled or External parent retains that profile. Local-memory-only writes are therefore not an unconditional source guarantee. Completion must also pass artifact validation and a lock-ownership check before the new consolidation baseline is accepted.
| File or directory | Role | Read behavior |
|---|---|---|
memory_summary.md |
Dense navigational summary. | Provided to the read path as the first quick signal. |
MEMORY.md |
Searchable handbook. | Grep by relevant keywords before opening deeper files. |
rollout_summaries/ |
Compressed evidence and lessons per rollout. | Open only a small number when MEMORY.md points there. |
skills/ |
Reusable procedures or scripts. | A deeper progressive-disclosure layer. |
raw_memories.md |
Merged Phase 1 input. | Primarily used by Phase 2 consolidation. |
6. The Read Path Brings Memory Back into Future Turns
The read extension injects instructions only when both Feature::MemoryTool and memories.use_memories are enabled, and a nonempty memory_summary.md exists in the selected namespace. Configuration updates retain the version selected when the thread started, keeping its summary and retrieval tools in the same namespace. Dedicated memory tools are exposed only when dedicated_tools is also enabled.
The read-path prompt does not tell the agent to open every memory
file. It first asks whether the user query is likely to benefit
from memory. If yes, it performs a quick memory pass: skim the
provided memory_summary.md, search MEMORY.md,
and only then open one or two pointed rollout summaries or skills.
It is the same progressive-disclosure discipline this series has
already seen in skills and context management.
future user query
-> memory_summary.md hints
-> grep MEMORY.md
-> open 1-2 rollout_summaries or skills
-> answer with hidden memory citation
-> record usage for future selection
Memory citation closes the loop. When memory is used, the read-path
prompt requires a hidden <oai-mem-citation>
block at the end of the final reply. MemoryCitation
stores citation entries and rollout ids. The streaming layer strips
the hidden markup from visible output, parses the citation, and
records usage for the underlying stage-1 outputs. The user sees a
clean answer; the runtime learns which memories actually helped.
The v2 read template uses the injected summary directly. It opens a matching rollout summary only when extra evidence, chronology, or uncertainty could change the answer, and searches when a necessary route is missing. Memory is not proof of current behavior: changing claims may require fresh source checks. Cite only a read summary that informed the answer, not the injected summary itself. Explicit remember, forget, or correction requests append notes under extensions/ad_hoc/notes/; later consolidation applies them instead of directly editing generated files.
7. Why polluted Exists
The dangerous failure mode is not simply forgetting. It is
remembering unstable material too confidently. If a thread depends
on web search, tool search, or an MCP server that may pollute
memory, that external context may be stale or only valid for the
current task. Codex provides disable_on_external_context:
when enabled, matching response items or MCP tool calls can mark
the thread as polluted.
Pollution also affects material consolidated earlier. mark_thread_memory_mode_polluted schedules another consolidation when the thread participated in the last successful Phase 2 baseline, and subsequent selection excludes that source. This both prevents fresh extraction and allows the existing workspace to forget its contribution. Marking and enqueueing do not mean files have already been removed; successful background consolidation still has to happen.
Compare it with the lesson above. "Keep matching language pages in sync" is a stable workflow candidate. "A web page showed the latest price as 19.99 dollars today" is a one-time external fact with a short shelf life. The first may enter memory; the second is better kept in the current turn with its source.
This rule transfers well to any agent runtime: temporary external facts may answer the current task, but they should not automatically become durable preference memory. If they must be saved, require clearer source, freshness, and explicit confirmation.
8. Put Durable Material Back in the Right Place
The memory system is easier to reason about once each kind of durable
material has a clear home. Memories do not replace context management
or AGENTS.md. They stand after rollout: old thread
evidence is extracted into reusable lessons, consolidated into a
local workspace, and selectively read by future turns.
| Information | Where it belongs | Reason |
|---|---|---|
| Team requirements, delivery flow, mandatory validation. | AGENTS.md or checked-in docs. |
Must load reliably; cannot depend on whether memory extraction hits. |
| What exactly happened in a thread. | Rollout / state DB. | Resume, fork, audit, and recovery need replayable evidence. |
| Recurring preferences, workflows, and known pitfalls. | Memories. | They reduce repeated explanation while remaining verifiable and replaceable. |
| What is currently visible on screen. | Chronicle or explicit current-turn tool reads. | Useful for locating context; not automatically suitable for durable memory. |
That completes the long-state side of the Codex route. Part II asked what the model sees now. Part X asked how an old thread can be replayed. Part XII asks which old lessons deserve to travel into a new thread. They are connected, but different modules handle them. Keeping those responsibilities separate prevents memory from becoming an ever-growing prompt and prevents mandatory rules from being entrusted to probabilistic recall.
Sources
- OpenAI Codex Memories docs: defaults, setup, storage, and per-thread control
- OpenAI Codex AGENTS.md docs: instruction discovery and layering
- OpenAI Codex Chronicle docs: opt-in screen context for memory building
Feature::MemoryToolfeature spec and default-off stateMemoriesToml/MemoriesConfigsettings and defaults- TUI
/memoriespopup and enable prompt memories/readandmemories/writecrate responsibilities- Background memory startup task: skip conditions, rate-limit guard, Phase 1 and Phase 2 order
- app-server starts the memory startup task after turn submission
- Phase 1 startup claim candidate selection
- Phase 1 selects enabled threads, including paginated history
- Phase 1 prompt: convert rollout into raw memory / rollout summary and allow no-op
- Phase 1 prompt: high-signal memory criteria
- Phase 2 consolidation: input selection, file sync, git baseline, and internal agent
- Phase 2 memory folder structure
- Read extension injects developer policy when feature and
use_memoriesare enabled - Read path prompt: when to use memory, layout, and quick pass
- Read path prompt: memory citation and ad-hoc update rules
MemoryCitation/MemoryCitationEntryprotocol structs- Memory citation parser for citation entries and rollout ids
- Streaming output strips hidden citation markup and parses memory citation
- Completed response items record memory usage and external-context pollution
- MCP tool calls can mark thread memory pollution from server metadata
mark_thread_memory_mode_pollutedsets a thread topolluted