The previous Mem0 chapter focused on an external memory layer: extract facts, deduplicate or preserve them, retrieve a small relevant set, and add those memories to the model input. Letta is easier to read if we shift the mental model. The durable object is the running agent, with preferences, persona, tools, active messages, and file context restored together. If those pieces are only searched at answer time, the agent keeps feeling newly assembled on every turn.

Questions this article answers. After this chapter, you should be able to explain how Letta memory_blocks differ from Mem0 memory records, why a Block update can force a system prompt rebuild, and why Letta Code's MemFS is more like versioned agent context than file syncing.

To keep Block from sounding like a retrieved snippet, this chapter follows one representative item: a human block that says the user prefers shorter answers. It enters state when the agent is created, gets compiled into the system prompt at answer time, may trigger a prompt rebuild when updated, and comes back with the message window, tools, and identity when the agent is restored. The process is concrete: Letta persists AgentState, then Memory.compile() selects and renders blocks when it builds a model request.

Stage Trigger Where it persists Change to the next model input
Create agent memory_blocks in the create request AgentState.blocks The preference initializes with the Agent and can be compiled into later prompts.
Persist snapshot The runtime persists agent state AgentState.blocks Memory restores beside model, tools, and message ids.
Enter the model The runtime builds an LLM request Memory.compile() reads blocks The block becomes part of this turn's system prompt.
Update later A user, tool, or API updates the block AgentState.blocks; Letta Code can also map context into MemFS files The next compile may change the first system message.

Source scope and versions. The current Letta entry repository retains the retired V1 server on archive; active development lives in Letta Code. Sections 1–3 and 4.2 use reproducible V1 source to explain Blocks, AgentState, and prompt rebuilding. Section 4.3 follows the current Letta Code local backend from committed Git files to memory visible on later turns. Both implementations use fixed source links. Local-backend behavior is not a guarantee about Cloud/API backends or private services.

1. How V1 stores memory when the Agent is created

1.1 The create request names the blocks to keep

The Letta README creates a stateful agent with memory_blocks right in the quickstart. The example labels two blocks as human and persona, in the same create call that chooses the model and tools. That is the first signal that this memory is not a retrieval result found at answer time. It is part of the initialized agent state. You can see the same contract in the README example and in CreateAgent.memory_blocks.

class CreateAgent(BaseModel, validate_assignment=True):
    name: str = Field(default_factory=lambda: create_random_username())

    # memory creation
    memory_blocks: Optional[List[CreateBlock]] = Field(
        None,
        description="The blocks to create in the agent's in-context memory.",
    )
    tools: Optional[List[str]] = Field(None, description="The tools used by the agent.")

The creation request already includes both memory and tools. Letta first decides which durable context the Agent starts with; later, the runtime compiles that context into the current step.

That entry point changes the question. An external memory layer usually asks which historical snippets should be retrieved for this query. Letta asks an earlier question: which long-lived identity, user profile, and work constraints should be saved as part of Agent state and restored with it? In Mem0, “the user likes concise answers” is closer to a searchable memory record. In Letta, it is closer to a human block or persona constraint that must re-enter the system prompt whenever the Agent is reconstructed.

At creation:
  1. caller passes memory_blocks, model, tools
  2. server builds AgentState
  3. memory blocks become durable state

At answer time:
  1. Memory.compile() renders blocks
  2. system prompt receives current memory
  3. successful step checkpoints the message window

1.2 AgentState keeps memory, tools, and the message window in one snapshot

AgentState makes that restoration process explicit. The source describes it as the state of an Agent at a given time, persisted in the database backend, with the information needed to recreate a persisted agent. The fields include message_ids, the system prompt, agent type, model settings, compaction settings, blocks, tools, sources, tags, secrets, identities, and more. Letta memory therefore lives next to the runtime objects that make the agent work, not in a separate facts table.

That is why Letta does not read like a memory plugin. If a long-running agent has to keep working for weeks, the system must restore who the agent was, which tools it had, where the active message window stopped, which context blocks compile into the prompt, and which sources or identities still apply. AgentState puts those objects together. Memory is not only content to be found; it is material required to rebuild the agent.

class AgentState(OrmMetadataBase, validate_assignment=True):
    """Representation of an agent's state... persisted in the DB backend."""

    message_ids: Optional[List[str]] = Field(...)
    system: str = Field(..., description="The system prompt used by the agent.")
    model: Optional[str] = Field(None, description="The model handle used by the agent.")
    blocks: List[Block] = Field(..., description="The memory blocks used by the agent.")
    tools: List[Tool] = Field(..., description="The tools used by the agent.")
    secrets: List[AgentEnvironmentVariable] = Field(default_factory=list)
    identities: List[Identity] = Field([])

The snippet shows the lifecycle directly. Blocks follow Agent state, not a single retrieval request.

Question Letta mechanism Consequence
Which facts become durable? memory_blocks / blocks Stable information becomes part of agent state.
What does the model see? Memory.compile() and system prompt Blocks are rendered into context instead of waiting behind a search API.
How does the active thread continue? message_ids / in-context messages The active message window is checkpointed after a safe step.
How is memory maintained? Block update, block history, MemFS Memory edits can change prompts and can be carried by history or git-backed files.

2. Blocks are reserved context, not search hits

2.1 A Block has a budget, label, and permissions

Letta's Block is documented as a reserved section of the LLM context window. It has a value, limit, label, read_only, description, hidden state, and tags. The same file defines default Human and Persona block types. That is why Letta memory reads like a set of state slots: each block has a name, budget, description, and explicit read/write fields.

class BaseBlock(LettaBase, validate_assignment=True):
    """Base block of the LLM context"""

    value: str = Field(..., description="Value of the block.")
    limit: int = Field(CORE_MEMORY_BLOCK_CHAR_LIMIT)
    label: Optional[str] = Field(None, description="Label of the block...")
    read_only: bool = Field(False, description="Whether the agent has read-only access.")
    description: Optional[str] = Field(None, description="Description of the block.")
    hidden: Optional[bool] = Field(None, description="If set to True, the block will be hidden.")

Those fields regulate the slot. limit prevents an always-visible block from expanding without restraint. read_only makes some blocks behave more like constraints. description tells the model or tools what belongs in the slot. Hidden state leaves room for runtime-owned material that is not necessarily shown in the same way. If blocks are reduced to “retrieval hits,” the main design is missed: a Block is a structured context slot with explicit capacity, access, and display rules.

2.2 compile() renders Blocks into the current system prompt

The blocks become model-visible through Memory. The source describes Memory as in-context memory that contains labelled Block objects and tools to edit them. The standard renderer emits <memory_blocks> and writes each block's label, description, metadata, and value. Then compile() chooses the rendering mode based on agent type, model provider, and whether git-backed memory is enabled.

def compile(self, tool_usage_rules=None, sources=None, max_files_open=None, llm_config=None, client_skills=None) -> str:
    """Efficiently render memory, tool rules, and sources into a prompt string."""

    if not is_react:
        if self.git_enabled:
            self._render_memory_blocks_git(s)
        elif is_line_numbered:
            self._render_memory_blocks_line_numbered(s)
        else:
            self._render_memory_blocks_standard(s)

Durable state may contain many objects, but the model only sees what compile() produces for this step. The problem is therefore not just long-term storage, but how blocks, sources, tool rules, and the message window become a stable prompt before a step runs. Mem0 commonly adds a retrieved set to the current answer; Letta first compiles default context from Agent state, then runs the current step with it.

Letta V1 state and context: memory_blocks belong to AgentState and Memory.compile() builds the system prompt; successful steps checkpoint message_ids and state.
Blocks persist in Agent state, Memory.compile() renders them into the system prompt, and successful steps checkpoint the active message window.

3. Block updates can rebuild the system prompt

3.1 First decide whether the change affects the prompt

If a block is reserved context, changing it can change the system prompt. Letta's BlockManager lists prompt-affecting block fields: description, label, limit, read_only, and value. Its update_block_async path checks whether anything actually changed and rebuilds system prompts for connected agents after prompt-affecting updates.

PROMPT_AFFECTING_BLOCK_FIELDS = {"description", "label", "limit", "read_only", "value"}

has_prompt_changes = any(
    key in PROMPT_AFFECTING_BLOCK_FIELDS and getattr(block, key) != value
    for key, value in update_data.items()
)

if has_prompt_changes:
    await self._rebuild_system_prompts_for_connected_agents(block_id, actor)
Shape-level example: how a Block update changes the system prompt

before:
  block(label="human", value="User prefers shorter answers")
  compiled system prompt contains:
    <memory_blocks>
      <human>User prefers shorter answers</human>
    </memory_blocks>

update:
  value = "User prefers conclusion first, then details"

after rebuild:
  AgentState.blocks stores the new value
  the first system message receives the updated memory content

On the next model request, the runtime recompiles memory from AgentState.blocks, so the first system message changes. The runtime, rather than the Block itself, renders this context into the prompt.

That check means more than writing a database row. Changing block metadata may be bookkeeping. Changing value, label, or limit can change what the model sees next turn. Letta lists those fields explicitly because a memory update can be upstream of a prompt update. Without the check, durable state and the actual model input can diverge, or every harmless field edit can trigger unnecessary prompt rebuilds.

3.2 Rebuilding memory and checkpointing messages are separate operations

That is the sharpest difference from an external retrieval layer. Mem0 retrieves memories before an answer. Letta Block edits can change the Agent's default context. The _rebuild_memory path refreshes memory, files, sources, and tool rules, recompiles memory, and updates the first system message if the compiled prompt differs.

Letta still does not keep every conversation token in the prompt forever. In v3, _checkpoint_messages persists new messages only when a step has completed safely, then updates message_ids or the conversation's in-context message set. The code also treats system-prompt overflow as its own stop reason, which is a reminder that blocks, tools, and system instructions need a separate capacity calculation.

Action Object changed What breaks if they are mixed
Update block Long-term state and later system-prompt content. Treating it as an ordinary message loses the state-slot semantics.
Rebuild memory / system prompt The base context visible to the next model call. If skipped, durable state and the model's actual input can disagree.
Checkpoint messages The post-step message window and recovery position. If all history becomes blocks, the prompt grows without a useful boundary.

4. Letta Code turns state into long-running work

4.1 MemFS turns context assets into versioned files

Letta Code describes itself as a stateful agent harness: agents have memory, identity, and experience over time. Its feature table makes that model concrete. Self-improvement says agents can rewrite their own context, including memory blocks, skills, and prompts. MemFS says all context, including memory blocks, is tracked through git.

This extends the same idea to long-lived context. Ordinary memory blocks answer how long-lived information enters agent state. MemFS asks how those context assets accumulate, can be audited, can be rolled back, and can be read by other runtime mechanisms. A coding agent that learns over time is not only carrying a user profile; it is carrying prompts, skills, notes, file-backed context, and version history.

4.2 V1 rendering chooses which Blocks enter the prompt

The Letta source shows how that becomes prompt-visible. In git-backed memory, _render_memory_blocks_git renders system/persona into a dedicated <self> section, renders other system/* blocks under nested <memory> tags, and renders external blocks as a file tree. compile_available_skills also renders agent-scoped skills from skills/ blocks. MemFS is therefore not just file syncing. It is versioned Agent context that the compile path can render for the model.

Git-backed memory does not mean the model can freely rewrite every context object, and it does not mean every file belongs in the prompt. Rendering rules still decide which blocks become <self>, which enter <memory>, and which skills are exposed as usable capabilities. Letta Code's design turns long-lived context from one prompt blob into manageable assets; the current model input is still produced by the compile path.

4.3 The current local backend compiles committed memory files

Suppose the user asks for English answers and the Agent edits a memory file but has not committed it. The current Letta Code local backend does not compile that new text merely because it exists on disk. collectCommittedMemoryFiles() checks Git HEAD, lists committed Markdown with ls-tree, and reads it with git show HEAD:…. A dedicated test verifies that committed persona text appears while dirty working-tree text does not. This separates editing drafts from memory eligible for compilation. File tools may still read uncommitted text; the rule applies to this compiler.

Letta Code local memory compilation: git commit makes working drafts visible in HEAD; compilation organizes committed system text and an external file catalog.
This figure focuses on local memory compilation. A commit makes text readable; the runtime still checks its revision and chooses how to deliver it on a turn.

The renderer keeps a familiar split: system/persona.md becomes <self>; other system/ bodies enter <memory>. Remaining non-Skill files contribute an external directory tree, with bodies available for separate reading. This projection excludes skills/. The turn request appends its available Skill list separately, which does not imply that every Skill body has entered the model.

One preference across three moments:
Working file changed, HEAD unchanged → compilation still reads old text
Git committed, next turn detects a new revision → recompile memory
Base prompt unchanged + supported mid-conversation system messages → send a memory update
Other cases requiring compilation → persist and use the rebuilt system prompt

getOrCompileSystemPrompt() implements that sequence. An unchanged raw prompt hash and memory revision reuse the stored result. When only memory changes and the backend supports system messages within the conversation, it retains the original system content, updates stored memory content and revision, and returns a memory update for the turn. Otherwise it recompiles and persists the system prompt. Recovery therefore depends on committed content, persisted compilation state, and turn delivery, not just file existence. Unreadable files are skipped, and these Git reads do not establish an atomic snapshot. Compilation success alone does not prove that every file was read completely or atomically.

5. Compare Letta with Mem0

5.1 Compare where memory lives and when it enters the model

Mem0 and Letta are both memory systems, but they answer different questions. Mem0 asks how an application can delegate cross-session facts to a memory service. Letta asks how a long-running agent owns its own state, tools, active message window, and editable memory. The former centers retrieval-time context assembly; the latter centers stateful agent reconstruction.

Question More like Mem0 More like Letta
Where does memory live? External memory layer Agent state
When does memory enter the model? Retrieved before the answer Compiled into system prompt or current context
What does an update affect? Memory records and later retrieval ranking Blocks, model-visible context, and connected system prompts
Best fit Shared user facts and preferences across apps Long-running agents with tools, identity, and self-maintained context

The next chapter moves to Graphiti / Zep. The question changes again: if memory is a temporal context graph rather than a state slot, how does the system answer what was true before, what is true now, and where a fact came from?

Sources