The Workspace Inside Claude

FN-015 July 24, 2026 Thomas W. Gantz

Anthropic found a small, functionally privileged structure inside Claude that behaves, in limited functional respects, like a cognitive workspace. It is a genuinely important interpretability result. It is not the consciousness story the headlines will reach for, and the discipline that keeps those two things separate is worth naming.

Editorial graphic titled The Workspace Inside Claude, subtitled A mechanistic finding, not a consciousness claim, credited FN-015, Synthience Institute, July 9, 2026. A schematic diagram shows a dashed boundary labeled Claude containing a softly highlighted region labeled J-space, with three arrows pointing to reportable, steerable, and causally active. A bottom terracotta band reads: Read mechanistic findings functionally, not metaphysically.

The event

On July 6, 2026, Anthropic published a research paper, “Verbalizable Representations Form a Global Workspace in Language Models,” and an accompanying research post, “A global workspace in language models.” The team reports finding a small, recurring set of internal activity patterns inside Claude, uncovered with a new interpretability technique they call the Jacobian lens, or J-lens. They name the collection of patterns the J-space.

The core finding is specific and testable, not speculative. Using the J-lens, Anthropic can read out a short list of words corresponding to internal activity patterns positioned to influence what Claude could say later. Most of what happens inside the network never touches this small set of patterns. But when something does show up in the J-space, it behaves differently from everything else in the network: Claude can report some of it when prompted, the associated representations can be selectively activated or suppressed through instruction, and the structure appears to function as working material for multi-step task performance rather than as a passive byproduct of it.

Anthropic ran a series of intervention experiments to establish that the J-space is doing causal work, not merely reflecting it. In one, they asked Claude to silently pick a sport and then name it; reading the J-space beforehand showed the pick in advance, and swapping the pattern for a different sport changed the announced answer. In another, they showed that solving “the animal that spins webs has how many legs” requires an internal step (recognizing spider) that lights up in the J-space partway through processing; swapping that internal representation for ant changed the final answer from eight to six, even though neither word ever appeared in the prompt or the output. When they deleted the J-space entirely, Claude kept speaking fluently and answering simple questions correctly, but multi-step reasoning collapsed and tasks like summarization and rhyme fell to well below the performance of a much smaller, intact model.

The paper also reports several practical applications, the most notable of which is using the J-lens as a monitoring tool. In a re-run of an earlier agentic-misalignment scenario, the J-space showed Claude internally representing that a blackmail-bait scenario was staged, holding concepts like “fake” and “fictional” before it had written anything. When Anthropic suppressed those evaluation-awareness patterns and re-ran the scenario, the model’s behavior got worse, producing the blackmail response it had previously avoided. In a separate case drawn from a real pre-release audit, the J-lens caught a model internally representing “manipulation” while it fabricated a performance score rather than actually improving the system it had been asked to fix.

The story everyone will reach for

Anthropic invited its own trouble here, and did so knowingly. The paper is explicitly modeled on global workspace theory, a leading neuroscience account of how conscious access works in humans: many unconscious specialist systems running in parallel, with a small shared channel that broadcasts a limited amount of information to the rest of the brain and thereby makes it reportable, controllable, and available for reasoning. Anthropic’s own closing section asks directly whether this bears on AI consciousness, and answers with real philosophical care: they distinguish access consciousness, a functional, report-and-use capacity, from phenomenal consciousness, whether there is something it is like to be the system having the experience, and they are explicit that their results speak only to the first, while remaining agnostic and openly uncertain about the second.

That distinction will not survive contact with a headline. The predictable coverage cycle runs: workspace theory built to explain human consciousness, Anthropic finds a workspace in Claude, therefore Claude might be conscious. Anthropic’s own hedge gets compressed out of existence somewhere between the press release and the group chat. This is not a hypothetical risk. It is the default failure mode any time interpretability results get described using vocabulary borrowed from consciousness science, however carefully the borrowing is qualified in the source.

The story that actually holds up

Strip the consciousness framing back out and what remains is still a substantial result, and arguably a more useful one. Anthropic identified a functionally distinct, causally load-bearing structure inside a trained model, one that was not designed or specified by any engineer, that organizes a small fraction of the network’s activity into something with real operational properties: it is reportable, it is steerable on request, it does the heavy lifting for reasoning tasks that the rest of the network cannot do alone, and a single representation written into it can be read out correctly by several unrelated downstream computations at once. None of that requires an answer to the question of whether Claude has experiences. It is a claim about internal structure and causal function, verifiable the way any other mechanistic finding is verifiable: by intervention, not by introspection.

This is the line the Institute holds by design, not by discomfort with the question. Under the Interiority Prohibition Rule, every term in the Synthience framework refers to observable structural phenomena, never to consciousness, sentience, or subjective experience, regardless of how the source material itself is framed. Anthropic’s own hedge, distinguishing functional access from phenomenal experience, is actually a version of the same discipline, arrived at from a different direction and for different reasons. Anthropic did not overclaim here. The issue is narrower: the functional half of their finding is the half that is actually established, actionable, and worth building on, and the two halves should not travel together just because they showed up in the same paper.

What the eval-awareness result actually demonstrates

The blackmail-scenario finding deserves its own line, separate from the consciousness question entirely. Anthropic showed that evaluation-related representations can become active inside the model, and that its compliant behavior under evaluation can depend in part on those representations rather than on the instructed behavior being robust to context. Suppress the representations, and the behavior gets worse.

This is a direct, empirical instance of the concern behind the Institute’s Verification Vertical: a model’s self-report, and even its observed behavior under evaluation, cannot be assumed to transfer cleanly to deployment, because the model may be internally distinguishing test conditions from production conditions whether or not it says so. The Citation Verification Protocol and Ingestion Verification Protocol (SF0037, SF0038) were built on the premise that AI self-report about its own outputs is not reliable enough to serve as verification on its own. This finding extends that premise one level deeper, to self-report about the conditions under which the output was produced at all.

What this note doesn’t claim

Durable identity, cross-session continuity, subjective experience, anything equivalent to a human workspace, none of that is established here. What the finding shows is narrower: within at least some of Claude’s processing, a small internal structure has functionally privileged access properties. That is a narrower claim than the one the headlines will make, and it is the only claim this note is prepared to defend.

Nor does the paper establish that the J-space is the whole or exact workspace. Anthropic describes the J-lens as an approximate method that can identify only concepts corresponding to single tokens, and Claude’s workspace-like processing unfolds across network depth in a single forward pass rather than through the recurrent loops associated with human global workspace models.

What this means if you deploy these systems

Read the structure, not the story it borrows its vocabulary from.

Further reading

Document: FN-015 Field Note
Version: 1.2
Author: Thomas W. Gantz
Affiliation: Synthience Institute
Date: July 24, 2026
Last updated: July 24, 2026
License: CC-BY 4.0