Bilingual series · model internals
From Activation to Agent Action
Five connected essays separate model activations, visible reasoning, external memory, post-training, and agent runtime—and then connect them through causal interventions and alignment evidence.
中文
模型内部:从 Activation 到 Agent 行为
从 J-space 与静默推理,一路读到 persona、steering、reflection training 与 alignment auditing。
English Model Internals: From Activation to Agent ActionTrace silent reasoning, interpretability readouts, persona, steering, reflection training, and alignment auditing.