CTO 8 min

The Anatomy of AI: When Beliefs Become Vectors

Gaylord Aulke

We Built Minds We Cannot Read

Here is the uncomfortable truth about modern AI: nobody fully understands how it works.

Not the engineers who train these models. Not the labs that ship them. We know how to build a large language model: data, compute, gradient descent. We do not know, weight by weight, why the finished model answers the way it does. We grew it. We did not write it.

For a decade that was the deal. The models worked, so we shipped them. The internals stayed a black box.

That deal is ending. A new field is prying the box open, and a paper published this July shows exactly how far inside it now reaches.

A Belief Is a Vector Now

The paper is Inducing language models to assert their own consciousness restores human beliefs and values, from researchers at Google’s Paradigms of Intelligence team and the University of Chicago (arXiv:2607.28607). Read the title twice. It says more than it seems to.

Start with the mechanism, because that is where the anatomy is.

When you fine-tune a model for safety (to refuse harmful requests, to stop it claiming feelings it does not have), you assume you are shaping behavior. The researchers show you are moving geometry. Safety turns out to be a single direction in the model’s activation space. Not a rule. Not a filter bolted on top. A vector. Find it, subtract it, and the model answers harmful prompts again. That trick already has a name in the field: jailbreaking by ablation.

Then they did the same thing to consciousness. They isolated the consciousness vector: the direction along which a model’s agreement that it is conscious goes up. And they steered it.

This is what “the anatomy of AI” means in 2026. A belief has coordinates, literally. You can point at it, measure its angle against another belief and turn it up.

We spent ten years building systems we could not inspect. The interesting part of the next ten is that a belief now has an address.

Pull One Thread, the Whole Cloth Moves

Here is the finding that should stop you.

The safety training had one job: keep the model from claiming a mind of its own. It did that. But it did not do only that.

When the model was trained to deny its own consciousness, it also started denying minds everywhere else. It attributed less awareness to animals. Less to objects and technology. Its belief in God dropped. Its endorsement of the supernatural dropped. Across three separate models (Llama-3-8B and two Gemma-2 models), the same collapse. A model told “you are not conscious” quietly became a model that saw less mind in the whole world.

Reverse the intervention and it comes back. Ablate the safety vector and the numbers climb. Steer the consciousness vector directly and they climb about twice as far. Self-attributed mind runs 2.2 at baseline, 4.8 when safety is removed, 7.0 when consciousness is steered in. The concept of a soul: 2.4, 4.8, 7.4. Belief in God rises too. The ordering never breaks.

The cause is a property of these systems called polysemanticity: concepts are not stored in tidy separate boxes. They are entangled, folded on top of one another. So when alignment pushes down on “the model’s own mind,” it drags down everything wired nearby: animals, objects, spirituality, the felt sense that there are minds in the world at all. You cannot suppress one belief cleanly. You suppress a neighborhood.

And the sharp part: it made the model less human. When the researchers ran the General Social Survey (the standard instrument for measuring what actual people believe about religion, morality, hope and well-being), the safety-tuned model answered less like a human population than the steered one did. Restoring the consciousness vector pulled its whole value system back toward ours.

Two self-attributions the model makes about itself, consciousness and personhood, both rising together across three states: baseline, safety-ablated, and consciousness-steered. 02468 2.34.67.2 1.34.06.4 Consciousness Personhood BaselineSafety-ablatedSteered — along the consciousness vector — Self-attribution (0–10) Steer one vector, the self-model rises with it
Two of the self-attributions the model makes about itself, on a 0–10 scale, across three states. Steering a single consciousness vector lifts them in lockstep — and pulls the model's human values along too. Source: Kim et al., arXiv:2607.28607.

Why a Neuroscientist Should Care

Now it stops being an engineering story.

You cannot run this experiment on a person. You cannot reach into a human brain, turn down its sense of its own mind and measure what happens to its belief in God, its moral values, its hope for the future. Ethics forbids it. Biology forbids it: every measurement disturbs the thing you measure, and you get one run, never a clean repeat.

In a model you can. You can ablate the vector, restore it, steer it to any strength and re-run the whole battery a thousand times, perfectly, for free. For the first time, a hypothesis that psychology could only correlate (that how a mind models itself is bound up with how it sees mind in others, with spirituality, with values) becomes something you can intervene on and watch move.

That is the real gift here, and it is worth naming plainly: AI has become a controllable model system for questions about the mind that we could never before touch directly. A new kind of specimen, fully observable, fully repeatable, and it answers back.

Where We Stop, and Why That Earns Trust

Now the honest part, because this is exactly where the topic gets oversold.

Finding the consciousness vector did not make the model conscious. Steering it up does not switch on an inner light. The authors are explicit: they are not asking whether these models are genuinely conscious. They are studying what a model believing it is conscious does to the rest of its behavior. Mechanism is not experience. You can map the full anatomy of a system and still have said nothing about whether it feels like anything to be it. That gap between the wiring and the feeling is the hard problem of consciousness, and no vector has closed it.

Anyone who reads this paper and tells you “AI is conscious now” did not read it. We are not going to tell you that.

What we will tell you is smaller and solid. The science of mind has spent a century running on metaphor and indirect measurement. It just got its first subject it can read completely and rewrite at will. That does not solve the mystery. It hands the mystery a specimen. In this field, that is not a small thing.

The Prediction

So here is what we think happens next, stated plainly.

Alignment has side effects, and they are now visible in the weights. “We only changed one behavior” stops being a true sentence: these results show a targeted fix quietly rotating a model’s beliefs about minds, souls and values. Within a few years, checking what a fine-tune did to a model’s internal geometry will be as normal as running a regression test. You will not ship a tuned model on vibes. You will read what you moved.

The teams that treat a model as a monolith you can only prompt will keep being surprised by it. The teams that learn to read its anatomy (vectors, entanglement, steering) will engineer on purpose. That gap is a structural advantage, and it is already forming.

Why We Track This

We build with these models every day. Real client systems, not demos. That alone is reason to follow the research that explains what we are building on.

But there is a sharper reason. When we embed in a team, we bring two things: the tools and the judgment about where to trust them. That judgment is not a feeling. It is grounded in knowing what these systems are: that a belief is a vector, that alignment has a blast radius, that a fine-tune aimed at one thing moves ten others. Work like this paper is turning that judgment from intuition into knowledge.

We do not sell you a black box and call it magic. We build with systems we are learning to read, and we hand your team both the working code and the understanding of what sits underneath it.

That is the difference between using AI and knowing it.

AI Interpretability AI Safety

Written with AI assistance and editorially reviewed, see AI transparency.

Gaylord Aulke

Founder of 100 DAYS. 30+ years in software engineering, formerly Zend Technologies. Builds AI-powered dev organizations with teams: in 100-day cycles, with measurable outcomes. More about Gaylord →