Model internals · Five-part route
From activation to agent action
An agent changed a file, but that does not mean the model directly touched disk. This series separates weights, activations, silent and visible reasoning, tool calls, runtimes, and environment results in time order, then asks which internal evidence supports causal claims.
Reading boundary: internal readouts are instruments, not word-for-word autobiography. Every article separates correlation, causal evidence from intervention, and interpretation of what a model may be doing.

Reading route
Build the map, then read out, identify, and intervene
All five articles reuse one retention-policy task so every concept lands on an input, state change, tool call, or file result.
01 · OverviewFrom Activation to Agent ActionFollow one file change through model computation, a proposed tool call, and runtime execution, separating activations, KV caching, and durable records.
Read Part I →
02 · WorkspaceJ-space and Silent ReasoningSeparate J-lens readouts, sparse reconstruction, and causal intervention through the reference implementation.
Read Part II →
03 · ReadoutCoT Monitoring, NLA, SAEs, and Latent ReasoningCompare CoT, NLA, SAE, and J-lens evidence, then trace continuous thought through Coconut.
Read Part III →
04 · IdentityPersona, Self-Model, and Post-TrainingTrace persona extraction, the Assistant axis, and one-sided capping to see how post-training stabilizes roles and where self-model evidence stops.
Read Part IV →
05 · InterventionActivation Steering, Reflection Training, and Alignment AuditingCompare ActAdd and CAA intervention positions, separating inference-time steering, reflection training, and independent audits with their validation and rollback conditions.
Read Part V →