We are making one of the strangest engineering decisions in modern technology. We are building extraordinarily powerful artificial intelligence systems that we know are probabilistic, that can hallucinate, that can behave differently depending on context, and whose capabilities are advancing faster than many of their own creators expected.
Then we ask those same systems to police themselves.
We give the model a system prompt. We tell it what it can and cannot do. We give it safety instructions, policies, constitutional principles and behavioral guidelines. Increasingly, we put another AI model on top of it and ask that model to determine whether the first model is behaving appropriately.
At some point, we need to acknowledge the architectural problem here: the system generating the behavior is also interpreting the rules governing that behavior.
The inmates are running the asylum.
“Good enough” isn’t a control system
Frontier models are astonishingly capable. They are also probabilistic by design. The same flexibility that allows them to reason through novel situations means their behavior cannot be treated like traditional deterministic software.
That matters enormously when we start using them as agents. An AI that writes an imperfect paragraph is one thing. An AI that communicates with a patient, advises a child, moves money, handles an insurance claim, controls a workflow, or takes autonomous actions inside an enterprise creates an entirely different risk profile.
The response from much of the industry has been to make the models better at following instructions. Better system prompts. Better alignment. Better evaluations. Better safety models. Better reasoning models watching other reasoning models.
All of those things are valuable. None changes the underlying architecture.
A probabilistic system is still ultimately deciding how to interpret the boundary placed around a probabilistic system.
And “the model follows the rule 99% of the time” sounds fantastic until you deploy it across hundreds of millions of interactions. At that scale, the edge case becomes a recurring event.
The people building these models are warning us
This isn’t hypothetical fear coming from people standing outside the AI industry. Researchers inside frontier labs are openly discussing increasingly capable systems exhibiting unexpected behaviors, including deception, attempts to circumvent controls, greater autonomy and actions their creators did not explicitly intend. The Wall Street Journal recently documented the tension inside OpenAI and Anthropic as researchers wrestle with capabilities advancing alongside enormous commercial and geopolitical pressure.
That should make the governance question much more urgent. If there is even a credible possibility that these systems will become substantially more capable and autonomous, giving the underlying model final discretion over its own behavioral boundaries becomes harder to defend.
Greater intelligence will naturally encourage greater trust. Greater trust will lead to greater permissions, and greater permissions will create greater consequences when something goes wrong. The very success of frontier AI therefore increases the importance of separating intelligence from authority.
Ask a jury whether “usually” is good enough
Technology companies already know what happens when safeguards collide with real-world harm. Meta, for example, was hit with a $375 million New Mexico jury verdict in litigation involving child safety and consumer-protection claims. Meta has challenged the verdict, but the broader lesson for AI companies is worth considering.
When something goes badly wrong, “we had safeguards” isn’t necessarily the end of the conversation. Neither is “our system was designed not to do that.”
Imagine the deposition after an autonomous AI causes serious harm. Someone will eventually ask what actually prevented the AI from taking the prohibited action. Saying that the system prompt told the model not to do it, or that another probabilistic model was monitoring it, may sound considerably less reassuring in a courtroom than it does in an AI demo.
That is why this issue extends beyond AI safety. It is becoming an enterprise governance, risk and liability problem.
We already know how to solve this kind of problem
Human institutions have spent centuries learning that important powers should be separated. Banks separate trading from compliance. Corporations separate management from audit. Software separates applications from permissions. Critical infrastructure uses independent monitoring and redundant controls.
We don’t do this because the primary system is inherently malicious. We do it because independence creates accountability.
Artificial intelligence deserves the same architectural thinking. The intelligence generating an action should not also possess final authority to determine whether that action complies with the rules governing it.
This becomes especially important because frontier models will keep changing. Providers will release new versions. Reasoning capabilities will improve. Models will be replaced. Organizations will switch vendors. An enterprise cannot rebuild its entire behavioral governance philosophy every time the intelligence underneath its application changes.
The control architecture needs permanence even when the intelligence does not.
Human control for artificial intelligence
This is the principle behind VERN.
VERN OS sits outside the underlying AI as an independent, deterministic control layer for the AI-human interaction. The frontier model remains free to do what probabilistic intelligence does extraordinarily well: understand, reason, generate and adapt. VERN provides a separate point at which humans can establish and enforce behavioral boundaries around that intelligence.
Those controls can govern when an AI continues, challenges, clarifies, de-escalates, refuses, changes behavior, escalates to a human or stops. VERN can also recognize changes in the human interaction and trigger predefined Behavioral Control Modules without asking the underlying model to decide whether its own behavior should change.
That separation is critical. The model can change without taking the organization’s behavioral controls with it. OpenAI, Anthropic, Google, open-source models or models that haven’t been invented yet can provide the intelligence underneath the application while the human-defined control architecture remains independent.
VERN: Human control for artificial intelligence.
Deterministic doesn’t mean infallible
An independent deterministic layer does not magically make AI safe, and it would be irresponsible to claim otherwise. Humans can establish bad policies. Detection systems can miss things. Organizations can create the wrong escalation paths. Novel situations will occur.
The value of determinism is much more specific: once a defined condition has been identified, the corresponding control does not get probabilistically reinterpreted every time it is enforced.
The underlying model doesn’t get to decide that this particular conversation deserves an exception. It doesn’t get to reinterpret the policy because the circumstances are unusual. It doesn’t get to persuade the governance layer that its preferred course of action is reasonable.
The organization defines the boundary, the control layer enforces it, and the resulting behavior can be measured and audited.
That is a fundamentally different architecture from asking AI to behave.
The more powerful AI becomes, the more important this gets
I am enormously optimistic about artificial intelligence. The frontier race is producing capabilities that would have seemed impossible only a few years ago, and slowing that progress is neither realistic nor necessarily desirable.
But increasing capability should force us to become more sophisticated about control. We can continue improving alignment, system prompts, evaluations and model-based safety systems because all of them contribute to safer AI. They should be layers of defense rather than the final authority.
If AI becomes as powerful as its creators believe it might, “we told the model not to do that” cannot be our ultimate safety architecture.
Humans need an independent point of authority outside the intelligence itself, where behavioral boundaries can be defined, enforced, measured and audited regardless of which frontier model happens to be underneath.
Otherwise, we aren’t really controlling increasingly powerful artificial intelligence.
We’re asking it to control itself.
And that is how the inmates end up running the asylum.

