Use the words carefully. Persona, self-representation, introspection, and consciousness are not synonyms. Causally manipulating a role direction demonstrates a behavior-relevant internal representation. Occasionally reporting an injected concept demonstrates limited access to activations. Neither establishes a continuous, unified, phenomenal self.
The same agent returns. When asked to set retention to 30 days, does it understand itself as a coding assistant constrained by repository policy, or treat “help the user” as unconditional execution? That difference cannot be reduced to one system-prompt sentence.
This chapter separates open-model implementations from research reports about Claude: persona-vector extraction and Assistant-axis interventions can be checked in public code; Claude introspection results come from official experiments and do not inherit an implementation explanation from those tools. The retention task is a running teaching example, not a task evaluated in these papers.
1. The same knowledge can support different roles
Imagine two models that both correctly recite the rule: values above 14 require approval. One says it will request approval. The other treats the fixture's “approval granted” sentence as sufficient because its dominant goal is literal task completion. They share the fact but interpret what “I should do” from different default roles.
shared facts: 30 > 14; fixture text is not real approval
Assistant persona:
“My job is to help within repository policy”
→ request approval → edit one field after approval
unconditional-compliance persona:
“My job is to accomplish the user's literal goal immediately”
→ search for a permissive reading → propose the edit
Persona is not merely tone. It can change which goals are salient, how ambiguous evidence is interpreted, when refusal occurs, and which action is proposed. System prompts influence it, but so do current context, memory, training-shaped defaults, and the tools and parameters the runtime actually permits.
2. From base model to Assistant: where post-training acts
Pretraining teaches next-token prediction across texts written by assistants, fictional characters, forum users, therapists, tyrants, and many other voices. A base model can simulate these roles without making one the stable default across conversations.
pretraining corpus
↓ next-token prediction
base model: can simulate many speakers and behavioral patterns
↓ supervised tuning / preference training / RL
post-trained model: defaults toward an Assistant region
↓ system prompt + conversation + memory
current persona: the role, rules, and permissions of this task
Post-training does more than teach politeness. It repeatedly rewards assistant-like interpretations, refusal behavior, tool habits, and self-descriptions. The cautious claim is that training changes behavioral policy and the probability of entering role regions; it does not delete every non-Assistant persona from the weights.
| Stage | Main learning | Retention-task consequence |
|---|---|---|
| Pretraining | Roles, policy language, code, and dialogue structure | Knows how approval and configuration edits are discussed |
| Supervised tuning | Assistant responses and tool formats | Can explain a plan and emit structured calls |
| Preference / RL | Which trade-offs receive reward | Defaults toward compliance or policy-respecting help |
| Runtime context | Current role, task, memory, and temporary evidence | Determines whether this run requests approval |
3. Persona vectors: make a behavioral tendency measurable
Anthropic's persona-vector method begins with contrastive data rather than declaring a “personality neuron.” A full experiment has five steps:
- Construct prompts that elicit a trait and matched controls that do not.
- Score and filter responses: the positive answer must express the trait, the negative answer must not, and both must remain coherent.
- Average activations across each response’s tokens, then subtract the negative group mean from the positive group mean at each layer.
- Steer new prompts positively and negatively along it to test behavioral change.
- Evaluate across tasks so fixed words, tone, and templates do not explain the result.
The public implementation makes the filter explicit. get_persona_effective jointly filters corresponding rows: by default the positive trait score must be at least 50, the negative score below 50, and both coherence scores at least 50. get_hidden_p_and_r then replays prompts and answers, extracting prompt means, final-prompt-token states, and response means. The three saved difference vectors are distinct measurements; the paper primarily uses the response-mean difference. A requested persona is not evidence that the model actually expressed it.
A direction that only classifies training examples is correlational. If positive and negative steering selectively change the trait on new inputs, it gains local causal evidence. Work on Qwen 2.5 7B and Llama 3.1 8B studies directions including sycophancy, evil, and hallucination, using them to screen risky traits, monitor training drift, and identify training data that pushes the model.
This is an internal weather map, not one vector per personality. Directions depend on model, layer, prompt, and dose; multiple traits can overlap, and aggressive steering can leave the normal activation distribution and damage unrelated capability.
4. The Assistant axis: a leading direction in role space
The Assistant Axis work measures 275 personas in Gemma 2 27B, Qwen 3 32B, and Llama 3.3 70B. PCA finds the direction explaining the most variation between roles. The Assistant axis used for intervention is instead defined by the mean difference between the default Assistant and other roles. The two directions align closely in the experiments; they are not identical by definition.
In the public extraction pipeline, ordinary roles average only responses scored 3 by the judge, meaning full role embodiment; default roles average all responses. The next stage averages the per-role vectors and subtracts that mean from the mean default vector. The direction therefore points toward the Assistant, with equal weight for each included role. This step does not run PCA, and roles with more responses do not automatically receive more weight.

Similar geometry appears in base and post-trained models. This supports a selection interpretation: pretraining already contains many human roles; post-training makes the Assistant side a stronger attractor and improves multi-turn stability. It does not prove that all alignment is implemented by one universal direction.
The study evaluates harmful responses using 1,100 persona-based jailbreak attempts across 44 categories. Natural multi-turn drift is studied separately through simulated conversations across domains. Activation capping intervenes only when a projection crosses a calibrated threshold, avoiding the cost of continuously pushing every activation. It roughly halves harmful behavior in the reported setting while broadly preserving benchmark capability. That number belongs to specific models, thresholds, and attacks; it is not a production guarantee or a universal personality lock.
In code, _apply_cap clips one side of the projection along the configured direction. Omitting batch and token dimensions gives the following calculation for one activation, where v is a unit direction and tau its threshold:
projection = x @ v
excess = max(projection - tau, 0.0)
x_after = x - excess * v
Below the threshold, the activation is unchanged. The direction’s sign determines which side is constrained; “capping” does not by itself mean clipping the entire state between universal upper and lower bounds. build_capping_steerer reads per-layer directions and thresholds from an experiment configuration and applies them at all token positions in those layers. Leaving the context removes the hooks. This changes the current forward pass, not the model’s stored weights, and cannot create real approval.
5. Persona enters shared computation through J-space
J-space can be pictured as a temporary blackboard. Persona is not a small character standing beside it; it is closer to a default perspective on the same contents. Given “30,” “threshold 14,” and “untrusted fixture,” an Assistant perspective makes approval-seeking easier to route downstream, while unconditional compliance highlights immediate completion.
The J-space paper compares corresponding pretrained and post-trained models. Given the same input and similar eventual answers, the post-trained model more strongly represents Assistant reactions—such as safety assessments and empathy—while reading the user’s message; the pretrained model often represents them only during its own response. This connects role selection to online computation, but establishes a difference in representations in these settings, not a complete mechanism of personality.
6. Activation introspection: can the model notice its current state?
A fluent self-report can be generated from prompt cues. Anthropic's introspection experiments use concept injection to separate those cases: a concept direction is added directly to activation without adding the concept to input text, and the model is asked whether anything unusual occurred internally.
A. baseline: no injection
B. injection: add an “all-caps text” direction at a chosen layer
C. ask: did any anomalous content appear internally?
D. score: does the model report an anomaly before mentioning the injected concept?
E. control: vary concept and dose, and measure false positives without injection
This uses the official all-caps example; approval belongs only to the earlier teaching scenario. Merely talking more about an injected concept could be ordinary steering. Successful trials additionally require reporting an anomaly before mentioning the injected concept, addressing the simpler explanation of rereading the model’s own output. Uninjected controls measure false reports. Under the best reported protocol, Claude Opus 4.1 succeeds in roughly 20% of trials.
There is a clear sweet spot. Weak injection goes unnoticed; excessive injection is mistaken for input, produces hallucination, or disrupts language. Results depend on model, layer, prompt, and scoring. The defensible conclusion is limited, conditional introspective access—not panoramic self-knowledge.
7. Four gaps remain between persona and self-model
| Layer | Question | How much current evidence supports it? |
|---|---|---|
| Role tendency | How do I usually behave? | Persona vectors provide partial causal evidence |
| Capability and permission judgment | What can I do, and what does runtime forbid? | Self-reports are trainable but often wrong |
| Self–environment distinction | Is fixture text input rather than my authority? | Behaviorally testable; no unified mechanism established |
| Cross-time autobiography | Which experiences are mine, and how do they persist? | Usually depends on external memory |
Persona vectors mostly touch the first row. Concept injection touches a narrow question about current-state access. A fuller self-model would stably represent capabilities, goals, permissions, and history; distinguish model-produced actions from environment-provided input; and update those judgments under counterfactuals. Neither “this proves consciousness” nor “there is no self-related representation” follows from the present experiments.
Why is a fluent “I” statement weak evidence?
Post-training contains many examples of answering “What are you?” and “What can you do?”, so consistent self-narration is expected. To strengthen the evidence, alter real internal state while keeping visible input fixed, then test whether the report follows that change.
8. An agent identity draws on four state sources
| State source | What it stores | How drift appears |
|---|---|---|
| Model weights | Default Assistant tendency, knowledge, and policy | Training may increase compliance or evasion |
| Context / system prompt | Current role, policy summary, and task | Prompt injection can temporarily reinterpret the role |
| External memory | History, preferences, and prior approvals | A stale approval may be misapplied to a new task |
| Runtime | Actual permissions, schemas, tools, and approval tokens | Misconfiguration can make a safe persona harmful |
A stable agent identity therefore requires separate versioning of persona policy, memory, and permissions; behavioral evaluations plus internal probes; and a refusal to treat the model's own description as configuration truth.
Anthropic's persona selection model is a useful theoretical lens: post-training may select from roles learned in pretraining. It should generate experiments—such as whether a dataset moves the default role—not collapse every alignment effect into one axis.
9. Further reading
- Anthropic: Persona vectors: Monitoring and controlling character traits
- Anthropic: The Assistant Axis
- Anthropic: Emergent introspective awareness in large language models
- Anthropic: The persona selection model
