The previous Mem0 chapter focused on an external memory layer: extract facts, deduplicate or preserve them, retrieve a small relevant set, and add those memories to the model input. Letta is easier to read if we shift the mental model. The durable object is the running agent, with preferences, persona, tools, active messages, and file context restored together. If those pieces are only searched at answer time, the agent keeps feeling newly assembled on every turn.
Questions this article answers.
After this chapter, you should be able to explain how Letta memory_blocks differ from Mem0
memory records, why a Block update can force a system prompt rebuild, and why Letta Code's MemFS
is more like versioned agent context than file syncing.
To keep Block from sounding like a retrieved snippet, this chapter follows one representative
item: a human block that says the user prefers shorter answers. It enters state when the agent
is created, gets compiled into the system prompt at answer time, may trigger a prompt rebuild when updated,
and comes back with the message window, tools, and identity when the agent is restored.
The process is concrete: Letta persists AgentState, then Memory.compile() selects
and renders blocks when it builds a model request.
| Stage | Trigger | Where it persists | Change to the next model input |
|---|---|---|---|
| Create agent | memory_blocks in the create request |
AgentState.blocks |
The preference initializes with the Agent and can be compiled into later prompts. |
| Persist snapshot | The runtime persists agent state | AgentState.blocks |
Memory restores beside model, tools, and message ids. |
| Enter the model | The runtime builds an LLM request | Memory.compile() reads blocks |
The block becomes part of this turn's system prompt. |
| Update later | A user, tool, or API updates the block | AgentState.blocks; Letta Code can also map context into MemFS files |
The next compile may change the first system message. |
Source scope and versions. The
current Letta entry repository retains the retired V1 server on archive;
active development lives in Letta Code. Sections 1–3 and 4.2 use reproducible V1 source to explain Blocks, AgentState, and prompt rebuilding.
Section 4.3 follows the current Letta Code local backend from committed Git files to memory visible on later turns.
Both implementations use fixed source links. Local-backend behavior is not a guarantee about Cloud/API backends or private services.
1. How V1 stores memory when the Agent is created
1.1 The create request names the blocks to keep
The Letta README creates a stateful agent with memory_blocks right in the quickstart.
The example labels two blocks as human and persona, in the same create call that
chooses the model and tools. That is the first signal that this memory is not a retrieval result found at
answer time. It is part of the initialized agent state. You can see the same contract in the
README example and in
CreateAgent.memory_blocks.
class CreateAgent(BaseModel, validate_assignment=True):
name: str = Field(default_factory=lambda: create_random_username())
# memory creation
memory_blocks: Optional[List[CreateBlock]] = Field(
None,
description="The blocks to create in the agent's in-context memory.",
)
tools: Optional[List[str]] = Field(None, description="The tools used by the agent.")
The creation request already includes both memory and tools. Letta first decides which durable context the Agent starts with; later, the runtime compiles that context into the current step.
That entry point changes the question. An external memory layer usually asks which historical snippets should
be retrieved for this query. Letta asks an earlier question: which long-lived identity, user profile, and work
constraints should be saved as part of Agent state and restored with it? In Mem0, “the user likes concise
answers” is closer to a searchable memory record. In Letta, it is closer to a human block or persona
constraint that must re-enter the system prompt whenever the Agent is reconstructed.
At creation:
1. caller passes memory_blocks, model, tools
2. server builds AgentState
3. memory blocks become durable state
At answer time:
1. Memory.compile() renders blocks
2. system prompt receives current memory
3. successful step checkpoints the message window
1.2 AgentState keeps memory, tools, and the message window in one snapshot
AgentState makes that restoration process explicit. The source describes it as the state of an Agent at a
given time, persisted in the database backend, with the information needed to recreate a persisted agent.
The fields include message_ids, the system prompt, agent type, model settings, compaction settings,
blocks, tools, sources, tags, secrets, identities, and more. Letta memory therefore lives next to
the runtime objects that make the agent work, not in a separate facts table.
That is why Letta does not read like a memory plugin. If a long-running agent has to keep working for weeks,
the system must restore who the agent was, which tools it had, where the active message window stopped, which
context blocks compile into the prompt, and which sources or identities still apply. AgentState puts
those objects together. Memory is not only content to be found; it is material required to rebuild
the agent.
class AgentState(OrmMetadataBase, validate_assignment=True):
"""Representation of an agent's state... persisted in the DB backend."""
message_ids: Optional[List[str]] = Field(...)
system: str = Field(..., description="The system prompt used by the agent.")
model: Optional[str] = Field(None, description="The model handle used by the agent.")
blocks: List[Block] = Field(..., description="The memory blocks used by the agent.")
tools: List[Tool] = Field(..., description="The tools used by the agent.")
secrets: List[AgentEnvironmentVariable] = Field(default_factory=list)
identities: List[Identity] = Field([])
The snippet shows the lifecycle directly. Blocks follow Agent state, not a single retrieval request.
| Question | Letta mechanism | Consequence |
|---|---|---|
| Which facts become durable? | memory_blocks / blocks |
Stable information becomes part of agent state. |
| What does the model see? | Memory.compile() and system prompt |
Blocks are rendered into context instead of waiting behind a search API. |
| How does the active thread continue? | message_ids / in-context messages |
The active message window is checkpointed after a safe step. |
| How is memory maintained? | Block update, block history, MemFS | Memory edits can change prompts and can be carried by history or git-backed files. |
2. Blocks are reserved context, not search hits
2.1 A Block has a budget, label, and permissions
Letta's Block
is documented as a reserved section of the LLM context window. It has a value, limit,
label, read_only, description, hidden state, and tags. The same file defines
default Human and Persona block types. That is why Letta memory reads like a set of state
slots: each block has a name, budget, description, and explicit read/write fields.
class BaseBlock(LettaBase, validate_assignment=True):
"""Base block of the LLM context"""
value: str = Field(..., description="Value of the block.")
limit: int = Field(CORE_MEMORY_BLOCK_CHAR_LIMIT)
label: Optional[str] = Field(None, description="Label of the block...")
read_only: bool = Field(False, description="Whether the agent has read-only access.")
description: Optional[str] = Field(None, description="Description of the block.")
hidden: Optional[bool] = Field(None, description="If set to True, the block will be hidden.")
Those fields regulate the slot. limit prevents an always-visible block from expanding without
restraint. read_only makes some blocks behave more like constraints. description tells the
model or tools what belongs in the slot. Hidden state leaves room for runtime-owned material that is not
necessarily shown in the same way. If blocks are reduced to “retrieval hits,” the main design is missed:
a Block is a structured context slot with explicit capacity, access, and display rules.
2.2 compile() renders Blocks into the current system prompt
The blocks become model-visible through Memory. The source describes
Memory
as in-context memory that contains labelled Block objects and tools to edit them. The standard renderer
emits <memory_blocks> and writes each block's label, description, metadata, and value. Then
compile()
chooses the rendering mode based on agent type, model provider, and whether git-backed memory is enabled.
def compile(self, tool_usage_rules=None, sources=None, max_files_open=None, llm_config=None, client_skills=None) -> str:
"""Efficiently render memory, tool rules, and sources into a prompt string."""
if not is_react:
if self.git_enabled:
self._render_memory_blocks_git(s)
elif is_line_numbered:
self._render_memory_blocks_line_numbered(s)
else:
self._render_memory_blocks_standard(s)
Durable state may contain many objects, but the model only sees what compile() produces for this
step. The problem is therefore not just long-term storage, but how blocks, sources, tool rules, and the message
window become a stable prompt before a step runs. Mem0 commonly adds a retrieved set to the current answer;
Letta first compiles default context from Agent state, then runs the current step with it.
Memory.compile() renders them
into the system prompt, and successful steps checkpoint the active message window.
3. Block updates can rebuild the system prompt
3.1 First decide whether the change affects the prompt
If a block is reserved context, changing it can change the system prompt. Letta's
BlockManager
lists prompt-affecting block fields: description, label, limit,
read_only, and value. Its
update_block_async
path checks whether anything actually changed and rebuilds system prompts for connected agents after
prompt-affecting updates.
PROMPT_AFFECTING_BLOCK_FIELDS = {"description", "label", "limit", "read_only", "value"}
has_prompt_changes = any(
key in PROMPT_AFFECTING_BLOCK_FIELDS and getattr(block, key) != value
for key, value in update_data.items()
)
if has_prompt_changes:
await self._rebuild_system_prompts_for_connected_agents(block_id, actor)
Shape-level example: how a Block update changes the system prompt
before:
block(label="human", value="User prefers shorter answers")
compiled system prompt contains:
<memory_blocks>
<human>User prefers shorter answers</human>
</memory_blocks>
update:
value = "User prefers conclusion first, then details"
after rebuild:
AgentState.blocks stores the new value
the first system message receives the updated memory content
On the next model request, the runtime recompiles memory from AgentState.blocks, so the first
system message changes. The runtime, rather than the Block itself, renders this context into the prompt.
That check means more than writing a database row. Changing block metadata may be bookkeeping. Changing
value, label, or limit can change what the model sees next turn.
Letta lists those fields explicitly because a memory update can be upstream of a prompt update. Without the
check, durable state and the actual model input can diverge, or every harmless field edit can trigger unnecessary prompt
rebuilds.
3.2 Rebuilding memory and checkpointing messages are separate operations
That is the sharpest difference from an external retrieval layer. Mem0 retrieves memories before an answer.
Letta Block edits can change the Agent's default context. The
_rebuild_memory
path refreshes memory, files, sources, and tool rules, recompiles memory, and updates the first system message
if the compiled prompt differs.
Letta still does not keep every conversation token in the prompt forever. In v3,
_checkpoint_messages
persists new messages only when a step has completed safely, then updates message_ids or the
conversation's in-context message set. The code also treats system-prompt overflow as its own stop reason,
which is a reminder that blocks, tools, and system instructions need a separate capacity calculation.
| Action | Object changed | What breaks if they are mixed |
|---|---|---|
| Update block | Long-term state and later system-prompt content. | Treating it as an ordinary message loses the state-slot semantics. |
| Rebuild memory / system prompt | The base context visible to the next model call. | If skipped, durable state and the model's actual input can disagree. |
| Checkpoint messages | The post-step message window and recovery position. | If all history becomes blocks, the prompt grows without a useful boundary. |
4. Letta Code turns state into long-running work
4.1 MemFS turns context assets into versioned files
Letta Code describes itself as a stateful agent harness: agents have memory, identity, and experience over time. Its feature table makes that model concrete. Self-improvement says agents can rewrite their own context, including memory blocks, skills, and prompts. MemFS says all context, including memory blocks, is tracked through git.
This extends the same idea to long-lived context. Ordinary memory blocks answer how long-lived information enters agent state. MemFS asks how those context assets accumulate, can be audited, can be rolled back, and can be read by other runtime mechanisms. A coding agent that learns over time is not only carrying a user profile; it is carrying prompts, skills, notes, file-backed context, and version history.
4.2 V1 rendering chooses which Blocks enter the prompt
The Letta source shows how that becomes prompt-visible. In git-backed memory,
_render_memory_blocks_git
renders system/persona into a dedicated <self> section, renders other
system/* blocks under nested <memory> tags, and renders external blocks as a file tree.
compile_available_skills
also renders agent-scoped skills from skills/ blocks. MemFS is therefore not just file syncing. It is
versioned Agent context that the compile path can render for the model.
Git-backed memory does not mean the model can freely rewrite every context object,
and it does not mean every file belongs in the prompt. Rendering rules still decide which blocks
become <self>, which enter <memory>, and which skills are exposed as usable
capabilities. Letta Code's design turns long-lived context from one prompt blob into manageable assets; the
current model input is still produced by the compile path.
4.3 The current local backend compiles committed memory files
Suppose the user asks for English answers and the Agent edits a memory file but has not committed it.
The current Letta Code local backend does not compile that new text merely because it exists on disk.
collectCommittedMemoryFiles()
checks Git HEAD, lists committed Markdown with ls-tree, and reads it with git show HEAD:….
A dedicated test verifies that committed persona text appears while dirty working-tree text does not.
This separates editing drafts from memory eligible for compilation. File tools may still read uncommitted text; the rule applies to this compiler.
The renderer keeps a familiar split:
system/persona.md becomes <self>; other system/ bodies enter <memory>.
Remaining non-Skill files contribute an external directory tree, with bodies available for separate reading.
This projection excludes skills/. The turn request
appends its available Skill list separately, which does not imply that every Skill body has entered the model.
One preference across three moments:
Working file changed, HEAD unchanged → compilation still reads old text
Git committed, next turn detects a new revision → recompile memory
Base prompt unchanged + supported mid-conversation system messages → send a memory update
Other cases requiring compilation → persist and use the rebuilt system prompt
getOrCompileSystemPrompt()
implements that sequence. An unchanged raw prompt hash and memory revision reuse the stored result.
When only memory changes and the backend supports system messages within the conversation, it retains the original system content,
updates stored memory content and revision, and returns a memory update for the turn. Otherwise it recompiles and persists the system prompt.
Recovery therefore depends on committed content, persisted compilation state, and turn delivery, not just file existence.
Unreadable files are skipped, and these Git reads do not establish an atomic snapshot. Compilation success alone does not prove that every file was read completely or atomically.
5. Compare Letta with Mem0
5.1 Compare where memory lives and when it enters the model
Mem0 and Letta are both memory systems, but they answer different questions. Mem0 asks how an application can delegate cross-session facts to a memory service. Letta asks how a long-running agent owns its own state, tools, active message window, and editable memory. The former centers retrieval-time context assembly; the latter centers stateful agent reconstruction.
| Question | More like Mem0 | More like Letta |
|---|---|---|
| Where does memory live? | External memory layer | Agent state |
| When does memory enter the model? | Retrieved before the answer | Compiled into system prompt or current context |
| What does an update affect? | Memory records and later retrieval ranking | Blocks, model-visible context, and connected system prompts |
| Best fit | Shared user facts and preferences across apps | Long-running agents with tools, identity, and self-maintained context |
The next chapter moves to Graphiti / Zep. The question changes again: if memory is a temporal context graph rather than a state slot, how does the system answer what was true before, what is true now, and where a fact came from?