Bilingual series · model internals

From Activation to Agent Action

Five connected essays separate model activations, visible reasoning, external memory, post-training, and agent runtime—and then connect them through causal interventions and alignment evidence.

中文 模型内部:从 Activation 到 Agent 行为

从 J-space 与静默推理,一路读到 persona、steering、reflection training 与 alignment auditing。

English Model Internals: From Activation to Agent Action

Trace silent reasoning, interpretability readouts, persona, steering, reflection training, and alignment auditing.

Route map from activation state through silent and visible reasoning to a tool call and environment action, with separate observation, memory, and post-training feedback paths