There’s a tendency to talk about AI guardrails as though they were switches.
Don’t say this. Don’t do that. Refuse these requests. Follow this policy.
New research involving Google scientists provides a fascinating demonstration of why large language models don’t necessarily work that way.
Researchers studied models that had been trained to avoid claiming they were conscious. They then examined what happened when that constraint was removed or manipulated. What they discovered wasn’t simply a change in whether the AI would say some version of “I am conscious.”
Its answers about seemingly different subjects changed too.
Models became more willing to attribute minds or inner experiences to animals, nature, other chatbots and even supernatural entities. Responses involving emotions, morality, religion and personal values shifted as well.
The researchers hadn’t simply found an on/off switch for one sentence.
They had disturbed something connected to a much larger network of behavior.
This Isn’t Evidence That AI Is Conscious
That distinction is important.
An AI saying it is conscious doesn’t establish consciousness any more than an AI saying it is frightened establishes that it experiences fear. LLMs generate language based on learned representations and patterns.
The researchers’ finding is interesting for a different reason.
They identified what has been described as a “consciousness vector”—a direction within the model’s neural activations associated with affirming or denying its own consciousness. When researchers manipulated that direction, answers shifted across 95 survey questions involving beliefs, values, religion, emotions and other subjects.
Think about what that means from a governance perspective.
We want one behavior to change.
We modify the model.
Other behaviors change with it.
And we may not know which ones until afterward.
Guardrails Inside a Probabilistic System Are Probabilistic Too
This gets to a problem much larger than consciousness.
Companies increasingly depend on model-level alignment and training to establish how an AI should behave. Providers teach models to refuse certain requests, avoid certain claims and respond according to particular safety principles.
That’s valuable work.
But an LLM isn’t conventional software containing thousands of neatly separated behavioral rules. Its concepts and representations can be interconnected in ways that aren’t immediately obvious even to the people who built it.
Changing one area can have consequences somewhere else.
That creates a significant challenge for enterprises.
Imagine a model provider modifies its system to reduce sycophancy. What else changes?
A safety update makes the AI more conservative about mental-health conversations. Does that affect how it responds to ordinary emotional distress?
A provider changes refusal behavior. Does that alter how frequently the model escalates ambiguous customer-service situations?
A new model becomes significantly better at reasoning. Does that increased capability produce new strategies for accomplishing an objective that weren’t previously possible?
The answer doesn’t have to be catastrophic for the problem to matter.
It only has to be unexpected.
Human Control Needs Somewhere Stable to Live
This is one of the reasons VERN was architected outside the underlying AI.
VERN is human control over artificial intelligence.
VERN OS provides an external governance layer around probabilistic intelligence. That separation allows the underlying model to change without requiring the organization’s behavioral requirements to change with it.
The distinction becomes increasingly important as enterprises move toward multi-model environments.
A company might use OpenAI today and Anthropic tomorrow. It might route different tasks to Gemini or an open model. Providers will continually retrain, align and update their systems.
Every one of those changes potentially alters the behavior of the intelligence underneath the application.
Your definition of acceptable behavior shouldn’t have to move with it.
The Human Side Matters Even More
There is another implication of this research that I find particularly important.
The behaviors that shifted weren’t confined to technical tasks. They involved concepts such as emotions, values, minds and moral judgment—the kinds of ideas that directly influence how an AI interprets and responds to humans.
That’s important as AI becomes more deeply embedded in healthcare, education, customer service, companionship and other human-facing applications.
A seemingly small model change could alter how an AI interprets someone’s distress, responds to emotional vulnerability or makes judgments about the person on the other side of the conversation.
For VERN, the objective is that the AI-human interaction remains human-centric and operates for the betterment, rather than the detriment, of the person using it.
That requires a stable layer of governance even when the intelligence underneath it isn’t stable.
The Model Can Change. Human Control Shouldn’t.
The takeaway from this research isn’t that AI is secretly becoming conscious.
It’s much more practical.
Modern AI systems are enormously complex, probabilistic and interconnected. Even researchers studying them can discover behavioral relationships they didn’t anticipate.
That should affect how we think about governance.
Model companies should continue making their systems safer. Researchers should continue studying what happens internally. Enterprises should take advantage of increasingly capable models.
But we shouldn’t assume the thing being changed is also the safest place to permanently anchor our control over it.
Human control needs to live outside the artificial intelligence, where changing the intelligence doesn’t change the rules.
That’s the architecture VERN was built for.

