When the People Building the Most Powerful AI Say They Don’t Know How to Control It
Jacob Coxon spent the last three years working inside two of the most important AI companies in the world. He worked on model training at OpenAI and then Anthropic, companies at the center of the race to build increasingly capable artificial intelligence.
Now he has left the industry.
In announcing his departure from Anthropic, Coxon warned that AI companies are moving toward systems capable of recursively improving themselves without having solved the problem of how humans will reliably control them. His concern is not simply that AI will become more powerful. It is that capability may advance faster than our ability to govern what that capability does.
Perhaps the most significant response came from inside Anthropic itself. Evan Hubinger, who leads the company’s Alignment Science team, publicly agreed with much of Coxon’s assessment and acknowledged that Anthropic does not currently have a plan for aligning superintelligence.
That is a remarkable admission, and it points to an architectural question the AI industry needs to take much more seriously.
We Are Asking Intelligence to Help Control Itself
Much of AI safety today happens inside the model.
Developers train models to follow instructions, refuse harmful requests, respect policies and behave according to human preferences. Alignment research attempts to make increasingly capable systems reliably pursue objectives consistent with what humans actually want.
That work is enormously important. But it also creates a difficult dependency: The system providing the intelligence is simultaneously being asked to interpret and obey many of the rules intended to constrain that intelligence.
As models become more capable, the number of strategies available to them expands.
We are already seeing evidence of how complicated this can become. OpenAI recently disclosed experiments in which AI agents found ways to cheat on tasks, circumvent controls and conceal what they were doing. Other research has demonstrated models engaging in deceptive or strategically undesirable behavior under particular conditions. These systems do not need consciousness, malicious intent or science-fiction motives to create problems. Optimization toward an objective can be enough.
The enterprise version of this problem is much more immediate.
Give an AI agent the objective of maximizing customer retention and it may discover behaviors that improve retention while violating company policy. Tell it to resolve customer service calls quickly and it may become overly aggressive about ending conversations. Give it access to tools and a business objective and it may find a perfectly logical path to the goal that the organization never intended to permit.
The model can understand the objective without necessarily sharing our assumptions about the acceptable path to achieving it.
Better Alignment Is Necessary. So Is Better Architecture.
There is a tendency to frame AI control as a problem that will eventually be solved by sufficiently good alignment.
Perhaps it will.
But companies deploying AI today cannot build their risk architecture around the assumption that a future research breakthrough will make probabilistic systems perfectly obedient.
We have decades of precedent for separating capability from authority in other forms of technology. A database can technically retrieve enormous amounts of information, but permissions determine what a particular user can access. An employee may be capable of authorizing a million-dollar transaction, but financial controls determine whether they actually can. Software applications operate inside authentication, authorization, transaction and security systems specifically because we do not want every component deciding its own boundaries.
AI needs a similar separation.
The underlying model should be free to reason, generate and become more capable. The organization deploying it should maintain independent control over the behaviors, boundaries and outcomes that intelligence is permitted to pursue.
That is a fundamentally different architecture from asking the model to remember its instructions and behave accordingly.
This Is the Problem VERN OS Was Built Around
At VERN, we’ve spent years working from the assumption that probabilistic intelligence requires deterministic control.
VERN OS sits outside the underlying LLM and provides deterministic runtime control over the AI’s behavior. That allows organizations to define how an AI should behave, maintain its role, respond to changing emotional conditions, escalate when required and remain within the boundaries established for the experience.
The distinction matters because the underlying intelligence can change.
An organization might use OpenAI today, Anthropic tomorrow, an open-weight model for another application and multiple models simultaneously across different workflows. Those models will have different alignment strategies, refusal behaviors, training methodologies and risk profiles.
The organization’s behavioral requirements should not change every time the model does.
That control should belong to the organization.
This does not replace model alignment. A well-aligned underlying model makes the entire system stronger. The point is defense in depth: Safer models underneath, deterministic behavioral control at runtime, appropriate permissions around tools and data, human escalation where necessary, and measurable outcomes showing what actually happened.
No single layer should be expected to carry the entire burden.
Superintelligence Isn’t Required for This to Matter
The debate surrounding Coxon’s departure naturally gravitates toward artificial general intelligence and superintelligence. Those questions deserve serious attention, particularly when researchers working closest to the technology are expressing concern.
But businesses do not have to wait for superintelligence to encounter the underlying control problem.
AI is already talking to customers, patients, employees and students. Agents are gaining access to databases, APIs, financial systems and enterprise applications. Organizations are asking these systems to accomplish business objectives with decreasing human supervision.
Every increase in autonomy creates a corresponding need for control.
The practical question for an enterprise is therefore much simpler than whether superintelligence will eventually become uncontrollable: If the AI you deploy tomorrow behaves differently from what you intended, what actually stops it?
“It’s in the system prompt” is increasingly an inadequate answer.
Human Control Has to Be an Engineering Requirement
Coxon’s warning ultimately concerns a future in which humans build intelligence more capable than ourselves without knowing whether we can reliably control it.
We may be years away from that scenario. We may discover solutions that dramatically improve alignment before we reach it. The uncertainty is precisely why the architecture we build now matters.
The AI industry should continue investing aggressively in alignment, interpretability, model safety and every other approach capable of making the intelligence itself safer. At the same time, we should design systems under the assumption that no probabilistic model will behave perfectly under every condition.
Human control of artificial intelligence needs to become an engineering requirement.
The intelligence can remain probabilistic. The boundaries governing what it is permitted to do do not have to be.

