A terminal scrollback, a transcript file, and a memory file can all mention the same fact. They are still different objects. Scrollback helps the user review the session. Transcript preserves evidence. Memory is material the runtime may load again so future turns can behave differently. Some of it enters user context, some enters a memory prompt section, and some belongs only to a selected agent.

The thesis is simple: Claude Code memory is long-lived files whose cached contents are reused across turns and reloaded after invalidation. It can affect prompt caching because inserted memory becomes part of the API-bound request prefix. But "dynamic" does not mean "changed every turn," and the visible source does not append this material to the system prompt.

Reading goal: follow one representative rule, "this project runs pnpm test." Track where it can be stored, who may write it, which prompt path later inserts it, why <system-reminder> is a meta user message on the claudemd path, and how changed text can break a prompt-cache prefix.

Evidence comes from the official Claude Code memory docs, Anthropic's prompt caching docs, and the public Rememorio/claude-code mirror. This article verifies the mirror’s fixed March 31, 2026 snapshot; every source link points to the same commit. It is not Anthropic’s official source repository and does not establish behavior in later product versions. Visible client implementation, official API contracts, and provider-side inference remain separate; code behind a feature switch does not prove that a deployment enables it.

1. Start with one project rule

Suppose the project rule is "run pnpm test, not npm test." If that rule only lives in old chat messages, compaction, resume, and new sessions make it fragile. The stronger representation is a long-lived instruction file or memory entry.

The visible source's claudemd.ts header names four instruction-file layers: managed memory, user memory, project memory, and local memory. It also describes discovery from the current directory upward, plus @include support.

representative rule

raw sentence:
  "this project runs pnpm test"

possible stores:
  CLAUDE.md             -> team-visible project rule
  CLAUDE.local.md       -> private local rule
  auto memory file      -> learned note for future turns
  agent memory          -> scoped to one agent's future work

	later insertion:
	  CLAUDE.md / rules -> user context -> model-visible request
	  auto MEMORY.md    -> user context or relevant-memory attachments
	  auto policy text  -> memory prompt section
	  agent memory      -> selected agent system prompt
The useful split is write first, insert later. The component that stores a rule, the prompt path that reads it, and the request that shows it to the model are different steps.

1.1 CLAUDE.md is human-written memory

getMemoryFiles() loads managed and user files first, then walks project directories from root toward the current working directory. A nearer rule can add task-specific constraints, but this loader does not remove conflicting text or enforce deterministic rule overrides. This is not transcript replay; it is filesystem-backed instruction loading.

That makes CLAUDE.md a good place for rules the model would otherwise get wrong: test commands, generated-code boundaries, branch conventions, or risky directories. It is a poor place for raw command output or one-off investigation evidence. Evidence belongs in transcript; durable decision rules belong in memory.

1.2 Auto memory is learned, but still restricted

Auto memory is a different path. isAutoMemoryEnabled() shows that it is enabled by default but disabled by --bare, explicit environment switches, missing remote persistence, or settings. getAutoMemPath() resolves the durable directory for the project.

The extraction path is not an inline mutation on every answer. extractMemories.ts says it runs after a completed query loop with a final response and no further tool calls, using a forked agent that shares the parent's prompt cache. Its createAutoMemCanUseTool() tool policy allows broad reads, read-only shell commands, and writes only inside the memory directory.

Enabling auto memory does not enable every background writer. The extraction entry point also requires a separate feature gate whose fallback is false, the main agent, and a non-remote run. If the foreground already wrote memory, extraction skips that range and advances its cursor. If extraction is busy, another call only stores the latest context for one trailing run; accepting the call does not mean that its files have been written.

Project rule: "run pnpm test"
1. Foreground answer finishes → check extraction gate and execution mode.
2. Foreground already wrote memory → skip; otherwise a restricted fork may run.
3. Another call arrives while busy → coalesce the latest context for a trailing run.
4. Fork returns successfully → advance the cursor; no new memory is a valid result.
5. Fork throws → keep the cursor so later extraction reconsiders those messages.
6. Saved content enters a later request → its read path and cache reload still decide when.
This is a simplified call and state trace. A completed answer, an extraction return, and a saved memory are distinct events.

The extraction body caps the fork at five turns and skips its transcript. A run may return without finding anything durable to save. The cursor advances only after a successful return; exceptions are logged as best-effort failures. Actual writes still have to complete through tools, so scheduling extraction does not establish that memory was saved or is visible in the next turn.

1.3 Session memory and auto dream are side paths

SessionMemory maintains notes about the current conversation through a background fork. Its shouldExtractMemory() check uses token thresholds, tool-call counts, and whether the last assistant turn is at a natural break.

autoDream is even more delayed: enough time and sessions must pass, and a lock must be available, before a forked agent consolidates multiple sessions. This is background maintenance, not a hidden model-side memory that rewrites the current answer.

1.4 The memory paths do not share one prompt entrance

The cover is a conceptual inventory: it puts long-lived materials on one desk so the responsible components are visible. The source path is stricter. CLAUDE.md, .claude/rules, and the auto-memory MEMORY.md entrypoint can go through getMemoryFiles(), getUserContext(), and prependUserContext() as a meta user message. Some feature modes replace the index injection with relevant-memory attachments. Auto memory's loadMemoryPrompt() supplies the dynamic system-prompt section for write and search policy. Agent memory's loadAgentMemoryPrompt() is appended to the selected agent's prompt, not the parent turn's user context.

Long-lived Material Source Entrance How to Read It
CLAUDE.md / .claude/rules getMemoryFiles() -> getUserContext() -> prependUserContext() Per-turn user context that reaches the front of API-bound messages.
auto memory / MEMORY.md getMemoryFiles(), loadMemoryPrompt(), extraction, and dream side paths The content entrypoint can enter user context or attachments; write/search policy enters a memory prompt section.
agent memory loadAgentMemoryPrompt() Visible only to the selected agent prompt, not to the parent conversation's user context.

2. How memory enters the next request

The common shorthand says memory is wrapped in a system reminder and inserted into the system prompt. If "memory" means the CLAUDE.md / rules user-context path, the direction is close, but the source does something more specific. The text is wrapped in <system-reminder>, but prependUserContext() creates a meta user message and prepends it to the message list.

getUserContext() loads filtered CLAUDE.md / rules files into claudeMd and adds the current date. Then query() calls the model with messages: prependUserContext(messagesForQuery, userContext).

shape-level request before API call

systemPrompt:
  default system prompt + system context

messages:
  meta user:
    <system-reminder>
    # claudeMd
    ...CLAUDE.md content...
    # currentDate
    Today's date is ...
    </system-reminder>

  ...selected conversation messages...
The tag says system-reminder, but the source constructs a user message. That matters for cache reasoning.

2.1 Model-visible does not mean transcript

The meta user message is visible to the model in the current API request. It is not the user's fresh input, and it is not the whole transcript. This paragraph is about the claudemd user-context path; auto memory and agent memory use separate prompt entrances. Transcript records evidence of what happened.

This is why project memory can return after compaction. Compaction rewrites old conversation view. The instruction files remain on disk. Once user context is reloaded, the same rules can be inserted again.

2.2 Dynamic memory can affect cache, but not automatically

Prompt caching rewards stable provider-visible prefixes. Because memory is prepended into API-bound messages, changed memory can change the prefix. The message-side marker placement in addCacheBreakpoints() makes this request-shape discipline explicit: one message-level cache_control marker is placed for the request.

Memory file reload preceding request-text comparison: unchanged text may reuse a prefix, changed text rebuilds its suffix, and provider conditions still govern hits
A file write changes request text only after the read path picks it up. Stable text makes reuse possible; provider cache conditions still determine an actual hit.

That does not mean memory destroys cache. getUserContext() and getMemoryFiles() are cached, and clearMemoryFileCaches() is reserved for correctness events such as settings changes, dialogs, or reloads. If the inserted memory text is unchanged, the prefix can remain stable. If a memory file changes and is reloaded into the API view, the old prefix may need to be rewritten.

A precise formulation is: memory participates in prompt-cache prefix discipline. It is not hidden system-prompt magic, and it is not a random per-turn mutation. It is user context or memory prompt material that participates in the API-bound prefix.

3. Why not put everything in memory

If memory can influence future turns, the tempting mistake is to store everything. The source does the opposite. Auto-memory writes are limited to a designated directory. Session memory waits for thresholds. Auto dream waits for time, enough sessions, and a lock. The system avoids promoting raw evidence into durable rules too cheaply.

Content Better Store Reason
Team rules, test commands, forbidden areas CLAUDE.md / .claude/rules They should influence many future turns.
Private local preferences or machine-specific habits CLAUDE.local.md or user memory They should persist without becoming team-shared instructions.
Command output, diffs, investigation evidence transcript / runtime history They are evidence, not necessarily future policy.
Repeated user preferences or agent practices auto memory / agent memory They benefit from extraction, deduplication, and later reload.

This is also why memory belongs before context management in the reading route. Memory explains where long-lived material comes from. The next chapter explains how that material is projected, compacted around, and separated from transcript and runtime view. The final prompt-cache chapter then returns to the same issue: stable prefixes come from consistent reading, insertion, and update rules, not from a single cache switch.

4. Four rules to keep straight

  1. Memory is not transcript. Transcript records what happened; memory stores durable rules and learned preferences.
  2. <system-reminder> is not the system prompt. In the visible source it is a meta user message created by prependUserContext().
  3. Dynamic memory does not mean every turn changes. Cache impact comes from changes to memory text inserted into the request prefix.
  4. Background learning is a side path. Extract memories, session memory, and auto dream all have triggers, write restrictions, and storage locations.

In one sentence: Claude Code memory combines long-lived files, cached read results, and reloads after invalidation. Once that model is clear, context management becomes easier to read: compaction rewrites old view, transcript preserves evidence, memory reloads rules, and prompt cache rewards stable API-visible prefixes.

Sources