Why the “foundational models will solve this” thesis is structurally wrong about any diagnosis in need of reliable measurement
There’s a confident assumption circulating in AI right now: that whatever a vertical AI company is solving, OpenAI will eventually get there. Scale the parameters, fine-tune the weights, add a product layer. Problem solved.
This thinking is correct for a large class of problems. It is structurally wrong for ours.
Measurement has one non-negotiable requirement: The same input must produce the same output. Every time. Without exception. A scale that reads 160 pounds on Monday and 158 on Wednesday — given the same object — is not a scale. It’s a guesser with good average accuracy. We would not call that measurement. We would call it noise. Or playing Russian roulette.
Eventually, the round in the chamber will fire.
Probabilistic systems (LLMs) cannot be measurement instruments
Large language models are probabilistic systems. Give the same prompt to GPT-4o twice and you will get different outputs. This is not a bug waiting to be fixed. It is the core design principle. Temperature, sampling strategies, RLHF…these systems were engineered to preserve variability. Variability is what makes generative AI feel creative, responsive, alive.
Which means you cannot use an LLM as a measurement instrument. Not now. Not with more compute. Not with the next generation of models. The architecture makes it impossible.
VERN is deterministic — not as a product feature, but as a foundational architectural requirement. The same inputs produce the same outputs, every time. You cannot retrofit determinism onto a probabilistic architecture. You would have to abandon the architecture entirely and start a different company, building a different thing, for a different purpose.
The training data *is* the problem, not the solution
LLMs were not trained on a coherent model of human behavior. They were trained on human-generated text…which reflects dozens of competing, often contradictory frameworks for understanding cognition. Freudian, Jungian, cognitive-behavioral, attachment theory, evolutionary psychology, social learning theory. The model absorbed all of it without adjudicating between them.
You ever hear of the term “Garbage in, garbage out?”
This is why asking the same behavioral question twice yields different answers. It is not merely stochastic noise. The underlying model of what a human is itself incoherent…a weighted average of frameworks that fundamentally disagree.
VERN is built on a single, internally consistent neuroscience model. One framework. One set of axioms. One definition of what human cognition is and how it operates. That consistency is what makes measurement possible, and it cannot be extracted or distilled from an LLM after the fact. It had to be the starting point.
Their strength is the constraint
OpenAI, or Anthropic’s commercial survival depends on variability. Creative output, conversational richness, the sense that the model is reasoning rather than retrieving — all of this requires probabilistic sampling. Their enterprise customers pay for that unpredictability. It is the product.
Shifting to deterministic output would degrade their core product for their current market in order to enter a market that requires fundamentally different infrastructure. That is not a strategic pivot. That is burning an existing business to build a new one from scratch on scientific foundations they don’t have.
Zero drift. Zero hallucinations. Zero character breaks.
Determinism doesn’t just solve the reproducibility problem — it eliminates an entire class of failures that make LLMs unusable as measurement infrastructure. (See how VERN gets you to 0.)
Hallucinations corrupt data. A probabilistic model that fabricates a response, even occasionally, poisons any dataset built on top of it. You cannot run measurement on a system where outputs can be confabulated. Character breaks — moments where the model steps outside its defined behavior — introduce noise that invalidates results. And model drift, the subtle but real shift in outputs as underlying models are updated and retrained, makes longitudinal measurement impossible. You cannot compare a score from Q1 to a score from Q3 if the instrument changed between them.
VERN eliminates all three by design. Not through guardrails bolted on after the fact — through architecture. Same input, same output, always. That is the only foundation on which a measurement instrument can stand.
Model-agnostic by design — and that’s the real structural advantage
Here is something that rarely enters the “won’t OpenAI just build this” conversation: VERN doesn’t compete with foundational models. It runs on top of them.
OpenAI and Anthropic are locked in an arms race with each other. Their incentive is to make you dependent on their specific model, their specific API, their specific ecosystem. VERN’s architecture is model-agnostic — it delivers the same deterministic measurement layer regardless of which model sits underneath. GPT, Claude, Gemini, whatever comes next. The protection travels with the client, not with the model vendor.
This also dramatically lowers switching costs for clients. As the market evolves — and it will evolve fast — clients running VERN can move to the most capable model available without losing their measurement infrastructure, their historical data, or their validation record. They are never locked to yesterday’s best option. The intelligence layer upgrades; the measurement layer holds.
That is a position neither OpenAI nor Anthropic can occupy. By definition, they cannot be neutral.
Validation history is not transferable
Every enterprise and clinical use case in behavioral measurement requires a documented track record of reproducibility. Regulatory bodies, procurement teams, and clinical partners don’t adopt new measurement instruments…they adopt validated ones.
Every deployment of VERN builds that record. Every consistent result across time, context, and cohort is evidence that the instrument works. An organization attempting to enter this space with a probabilistic system–even if they could somehow rebuild their architecture–would start that validation process from zero. Meanwhile, the gap compounds.
The foundational models won’t solve this. Not because they lack capability, but because solving it would require them to stop being what they are.

