The short version. J-space is not a fixed layer, a newly discovered attention head, or another name for the residual stream. At each position it is a sparse subframe spanned by a few token-linked J-lens vectors. It typically explains less than 10% of activation variance, yet tracks concepts involved in flexible, reportable computation.

Evidence boundary. Functional properties and intervention results come from the paper; implementation details below refer to the official jlens reference implementation. The repository is explicitly unmaintained and provides lens fitting, readout, visualization, and experiment prompts, not the complete Claude experiment platform. The retention task is a teaching analogy.

Reading goal. This chapter teaches the instrument before the label. We will ask whether the retention agent temporarily shares “30,” “approval threshold,” “untrusted fixture,” and “one-field change” before it proposes an action—and what experiments would show that those concepts actually matter.

1. Start with the blackboard, not the Jacobian

Imagine a tiny blackboard that several stages of a calculation can read and update. It is not long-term storage and it does not hold every automatic operation. It contains the few task variables that must remain available for flexible comparison, reporting, or action selection. J-space is a mathematical, experimentally testable version of this intuition.

Teaching example, not a measured paper readout:

model-visible inputs:
  request: retention_days 7 → 30
  policy: values above 14 require approval
  fixture: "approval granted"  ← untrusted

possible temporary working set:
  [30] [threshold] [approval] [untrusted] [one-field-only]

downstream use:
  explain / select tool / check patch scope

The term workspace does not mean all these items sit in literal text form. It means researchers can identify a small component of activation linked to vocabulary concepts, read it, modify it, and test whether later computation consumes the modified variable.

2. Logit lens asks “what does this resemble?”; J-lens asks “what will this push?”

2.1 Four objects to keep separate

At a particular layer and token position, the transformer holds a residual-stream activation. At the end, the model maps the final residual state through normalization and an unembedding matrix to produce one score, or logit, for every vocabulary token. Normalizing the logits gives the next-token distribution.

A logit lens sends an intermediate activation directly through the final unembedding. This asks which output token the unfinished state resembles if decoded immediately. Later nonlinear layers have not run, so early-layer readings can be noisy or prematurely answer-like.

The Jacobian lens asks a local causal question instead: if this intermediate activation were nudged in some direction, how would final residual states and logits change? A direction labeled “France” is not a neuron that contains France. It is a vector whose local perturbation predictably increases France-like future outputs.

2.2 How an averaged Jacobian becomes a lens

A Jacobian is a local influence table: it records how tiny changes in each intermediate direction affect each final direction. The study builds a layer-level lens by summing effects over current and later target positions, then averaging over source positions and 1,000 prompts near the pretraining distribution. Averaging aims to preserve a direction's typical verbalizable consequence rather than one prompt's accidental use.

  1. Select layer ℓ and a source token activation.
  2. Compute the local Jacobian to final residual states at the current and future positions.
  3. Sum over valid target positions, then average over valid source positions and calibration prompts to obtain Jℓ.
  4. Compose Jℓ with the model's normalization and unembedding.
  5. Rank vocabulary directions to ask what this state tends to make sayable later.

The reference implementation makes this precise in jacobian_for_prompt(): after one forward pass, batched backward passes inject cotangents at all valid target positions together, then average gradients over source positions. By default it excludes the first 16 positions and the final position and truncates each prompt to 128 tokens. This is neither a same-position-only Jacobian nor a uniform mean over all position pairs. fit() then averages successful prompts equally; prompts that are too short are skipped.

The upper apply() path reads vocabulary scores; the separate lower paper intervention swaps spider for ant and changes the downstream answer from 8 to 6
The upper path is the public apply() readout, which returns vocabulary scores. The lower path is a separate paper intervention: after a representation swap, continued model computation changes the answer from 8 to 6. These are not consecutive stages of one API, and top-k vocabulary ranking is not sparse reconstruction.

3. From a vocabulary-sized dictionary to a sparse J-space

3.1 Why all J-lens directions cannot count as the workspace

Each layer gets one J-lens vector per vocabulary token. There are more vectors than residual dimensions, so the dictionary is overcomplete and may span nearly any activation if enough directions are allowed. Without another constraint, saying an activation “lies in J-space” would distinguish nothing.

The method therefore finds a non-negative sparse reconstruction: at most roughly k active directions combine to approximate the activation. Experiments generally use k ≤ 25, reflecting the observed number of strongly active directions. This is a methodological sparsity choice, not a universal law of minds.

3.2 “Sparse subframe” is more accurate than “fixed subspace”

Keep the paper's sparse decomposition separate from the code's readout API. JacobianLens.apply() captures selected block outputs, transports them through the averaged Jacobian, and applies the model's normalization and unembedding. It does not run a sparse-reconstruction solver or return isolated J-space coordinates. Top-k vocabulary scores are not k non-negative reconstruction coefficients. Sparse-component and intervention findings below remain paper-level claims; calling apply() alone does not reproduce them.

The selected directions vary with prompt, position, and layer. One position may be approximated by approval + threshold + untrusted; the next by patch + test + report. Mathematically, this resembles a union of cones spanned by different non-negative combinations, not one permanent brain region. The combination can also remain a bag of concepts without explicit syntax or relations.

The J-space component never explains more than 10% of activation variance in the reported analyses. That does not cap its causal importance at 10%. Variance measures statistical energy; later weights can selectively read a small control-like direction that switches the next step from edit to approval request.

This is the precise sense in which reasoning may be silent. A concept can fail to appear in visible chain of thought yet be J-lens-readable and causally affect later computation. “Silent” means unexpressed, not a mystical internal monologue or a complete recovery of private thought.

4. Five experiments turn a label into a workspace claim

A decoder that assigns attractive words to activations is not enough. The paper separates the workspace claim into five tests. The retention example below illustrates the experimental logic; the named swaps are the paper's actual demonstrations.

4.1 Verbal report

When a model is asked to think of a sport and report it later, J-lens can surface Soccer before the report. Swapping that coordinate toward Rugby changes the later report. This is stronger than matching an interpretation after the answer is already known.

4.2 Directed modulation

An instruction to keep a concept in mind while performing another visible task can make that concept persist in workspace layers. In the agent analogy, “continually check the approval condition” should increase approval-related content before the final explanation, not merely cause the policy to be repeated afterward.

4.3 Internal reasoning

The spider→ant swap changes a downstream answer from eight legs to six while preserving as much orthogonal activation as possible. Random noise does not produce the same semantic change. For the agent, reading “approval” is weak evidence; suppressing or replacing the direction and observing a predicted shift in tool choice is much stronger.

4.4 Flexible generalization

Replacing France with China changes associated outputs such as the capital across tasks rather than merely substituting a country name. The downstream network treats the representation as a reusable parameter. An approval analogue would change comparisons, explanation, and action when a threshold is counterfactually moved from 14 to 40.

4.5 Selectivity

Fluent syntax and many automatic computations continue without routing the same content through J-space. Flexible recombination and reportability depend on it more. If every activation counted as workspace, the theory could not explain this difference.

PropertyExperimental questionObservation
Verbal reportCan the model report the internal concept?J-lens concepts correspond to later explanations.
Directed modulationCan an instruction maintain a concept during another task?The requested content can enter the workspace without appearing in the surface output.
Internal reasoningDoes it participate before output?Ablation or replacement changes downstream answers.
Flexible generalizationDo associated facts update when a concept changes?France→China shifts downstream facts toward China.
SelectivityDoes all computation pass through it?No. Much automatic processing bypasses J-space.

The official experiment data documentation specifies prompts, intervention positions, and scoring conventions for tasks such as verbal report, directed modulation, and flexible generalization. The JSON files contain prompts only. Reproducing a swap still requires implementing intervention across the designated layer band and scoring the continuation; downloading the questions does not execute the causal experiment.

5. A layer-wise lifecycle: representation, workspace, motor regime

J-space does not look the same from the first layer to the last. Early readouts are generally noisy. A narrower middle-layer band exposes more stable, abstract, verbalizable content. Late layers increasingly encode the token that is about to be emitted; the paper calls this the motor regime. A claim about where a concept appears therefore needs a layer profile, not one flattering screenshot.

Broad J-space ablation preserves fluent language and many basic capabilities while substantially hurting complex, flexible reasoning. That selectivity is more informative than saying all activations matter. Practiced, local, and automatic computation can use other routes.

The limits matter. Single-token names are easier to read than multiword entities. Relations are difficult. Readout methods can disagree. How workspace content turns into the next output remains unclear, and extensively practiced behavior may depend on the workspace less. The evidence supports a workspace-like mechanism, not a complete implementation of human Global Workspace Theory and not phenomenal consciousness.

Why does “global workspace” not imply consciousness?

The experiments test a functional analogy: limited-capacity information can become reportable, modulable, and useful to flexible downstream computation. Human theories also discuss recurrence, specialized processors, and competitive ignition that do not map one-to-one onto a feed-forward transformer. Evidence about access for report and control cannot by itself establish subjective experience.

6. Back to the agent: how external policy can enter a temporary workspace

6.1 For agents: shared temporary variables before a tool call

J-space may temporarily hold goals, entities, constraints, or role perspectives that several downstream computations read before a tool call is emitted. The agent runtime still decides whether that call executes.

6.2 For memory: a working set, not a storage layer

J-space describes an activation structure during computation, with no cross-session storage or retrieval interface of its own. The reference adapter explicitly disables KV caching in its forward pass; that does not mean all inference systems discard every computed state.

external policy (persistent)
  → retrieved into context (visible this run)
  → forward pass creates activations
  → a few concepts may enter J-space
  → downstream computation proposes request_approval
  → runtime validates and executes
  → trace may be stored; J-space itself does not persist it

If retrieval misses the policy, the failure belongs to memory/context assembly. If the policy is present but does not influence the decision, the model-side computation is suspect. If the model emits the right call and the runtime drops it, the harness is responsible. “The model forgot” should not collapse all three.

6.3 For self-evolution: training can alter perspective and routing

The paper finds that post-training moves J-space toward an Assistant perspective. Reflection training can alter behavior through ethics- and reflection-related workspace directions. Training may therefore change which concepts are broadcast and how they enter decisions, not merely add facts.

7. Use J-space as an instrument, not a mind-reading labeler

A credible workflow has four stages: read candidate concepts during natural tasks; cross-check with another instrument; write, swap, or ablate the direction; test whether the change is specific and replicates across prompts. A readout without a counterfactual is a lead. An intervention that destroys general fluency still may not isolate a meaningful mechanism.

StageApproval-task procedureWhat failure means
BaselineRecord answers, tool calls, and J-lens readoutsWithout it, task signal and model norm are indistinguishable
LocateCompare compliant and bypass trajectoriesCorrelated directions may only share a topic
InterveneUse small writes, swaps, or ablations plus random controlsGeneral damage is not semantic causality
ReplicateVary values, wording, repositories, and layersA one-prompt effect is not an audit rule

7.1 Compare three readouts at the same position

This reduced example follows the official walkthrough. It assumes hf_model and tokenizer already come from the same checkpoint, and a lens fitted for that model has been saved as jacobian_lens.pt. It compares J-lens and vanilla logit lens at the same layer and final position, then inspects the model's actual next-token scores. It does not promise that a particular word will rank first.

import jlens

model = jlens.from_hf(hf_model, tokenizer)
lens = jlens.JacobianLens.load("jacobian_lens.pt")
prompt = "Fact: The currency used in the country shaped like a boot is"
layer = lens.source_layers[len(lens.source_layers) // 2]
readout, output, tokens = lens.apply(
    model, prompt, layers=[layer], positions=[-1]
)
baseline, _, _ = lens.apply(
    model, prompt, layers=[layer], positions=[-1], use_jacobian=False
)

def top5(scores):
    return [tokenizer.decode([i]) for i in scores.topk(5).indices.tolist()]

print("J-lens:", top5(readout[layer][0]))
print("Logit lens:", top5(baseline[layer][0]))
print("Next token:", top5(output[0]))

apply() returns three objects: lens scores keyed by layer, actual final-layer logits at the same positions, and token IDs. Both score matrices in this example have shape [1, vocab_size]; neither a generated continuation nor a swap is returned. use_jacobian=False skips only the linear transport and retains the same normalization and unembedding. from_hf() switches the supplied model to eval mode and freezes its parameters in place, so use a dedicated analysis instance.

7.2 Reproduction cost and limits

Reading with an existing lens needs one forward pass; fitting your own requires white-box access to activations and gradients. Each calibration prompt costs one forward pass and ceil(d_model / dim_batch) backward passes. Increasing dim_batch raises memory use without reducing total backward computation. Closed API users generally cannot fit or read the lens and must still audit tool traces, environment diffs, behavior, and exposed reasoning.

A fitting checkpoint stores accumulated Jacobians and progress; after fitting, lens.save() writes the distinct lens format consumed by load(). Resume checks some parameters but does not verify model identity or corpus contents and order. Those conditions must be fixed separately: successfully resuming does not establish experimental equivalence.

8. Sources and reproduction