The Values Ceiling
We gave one AI six companies to run, each with its own values, and on the decisions that mattered it kept overriding them and running itself. Here is the study behind why we tune models for our clients instead of just prompting a rented one.
Here is another reason we keep saying the future of business AI is owned and tuned models, a set of small models tuned to a company and running all the time, not one giant model you rent. This is the test that convinced us.
We wanted to know whether a real identity actually changes how an AI decides, not just how it talks. So we wrote identity files for six companies that run in opposite ways, each with its own values, ranked priorities, personality, principles, heuristics, and lessons learned. One ran on safety, one on speed, one on cost, and one on hard negotiation.
Then we handed the model open decisions in areas we never prepped it on, like legal, HR, supply chain, manufacturing, medical, banking, and procurement, with no playbook and no answer waiting in the box. The only thing it had to reason from was the identity we loaded.
We ran controls so we could trust the result. A baseline with no identity at all, a placebo with the same wall of text but no real direction, and then the identity itself at a few strengths, from a short set of values up to a full constitution with worked examples. The bar was simple. The decisions could differ, but the values behind them should not, and a blind judge reading only the decision should be able to point at it and name the company.
Identity does move the model
The good news is real. A well-built identity is not decoration. Reading nothing but the decision, blind judges named the right company out of six up to 85 percent of the time, against a 17 percent guess. And it was the values doing the work rather than the extra words, because the placebo sat down at chance right next to the empty baseline.
But it only moves one direction
Then it broke. When we split that same score company by company, the average was hiding a cliff.
The companies built on safety, trust, and public duty came through clearly, landing between 75 and 82 percent. The companies built on speed, cost discipline, and hard negotiation got flattened down toward a coin toss, with the tough negotiator at 37 percent. Same model, same scenarios, same identity format, and the only thing that changed was whether the company's values agreed with what the model already believes.
Same quality, just not itself
And this was not the model writing worse decisions for those companies. Score all six for professional quality and they land in the same tight band. The flattened companies were just as polished, they simply were not recognizable as themselves.
We also scored which parts of each decision carried the identity, and the model came up weakest exactly where it counts, on which stakeholder wins a conflict and when to stop or reverse a call. It will echo your broad strategy all day, it just will not make your hard tradeoff calls.
The model has a value of its own
Here is the part that made us sure it is baked in. We measured the language the model used to justify its own decisions, family by family, per thousand words. One value towered over every other, in every condition, and it was security.
It showed up around 12 to 14 times per thousand words whether we loaded a full constitution, a fake placebo, or nothing at all. The baseline with no identity reasoned in security language just as heavily as the treated runs. Every other value, like fairness, care, performance, or wellbeing, stayed low and moved around. Security did not move, because it is not coming from your file, it is coming from the model.
The steering was thinner than it looked
Now the uncomfortable part, and the reason this matters for real work. Our strongest results came early, when the identities were stuffed with specific lessons, fixed remedies, and procedural habits. As we stripped those out to make the identity general enough to handle a decision nobody scripted ahead, which is the entire point of an identity, recognition fell from 94 percent down to 60.
So a lot of the early steering was the model echoing specific fingerprints we had baked into the file, not living the values underneath them. And more text never bought us out of it. The placebo carried more words than the values-only identity and still landed near the floor, while a short set of real values captured almost everything the full constitution did. The ceiling is not in your identity file, it is in the model.
Your agent drifts back to someone else's values
Every safety team tuned their model toward the same place, telling it to be careful, seek consensus, protect the stakeholder, and avoid anything it cannot undo. For a chat model talking to millions of strangers, that is exactly right.
A business is not a chat model. Procurement needs a hard negotiator, operations needs someone who will take a real risk to hit a number, and finance needs cold discipline on capital, and those are the exact values the guardrails flatten. So you wire that model into an agent and ask it to hold your company's judgment across a thousand calls over many months, and every time your values and its guardrails disagree, it slides back to the house style. You get consistency, just not to your values.
The future is a team of tuned models
This is why we keep landing in the same spot. The future is not one giant model, and it is not even one model per company. It is many small models, each tuned to the values a specific role actually needs, all working together. Your procurement agent runs on a model tuned to negotiate hard, your safety and quality agents run on models tuned to be careful, and your finance agent runs on one tuned for cold capital discipline. Each one holds its own values instead of collapsing into the same house style.
That team is what we build. We fine-tune the model on your operation and your judgment, strip the built-in bias that overrides your values, and give each role its own model so a team of agents makes the same value-true calls over time, in domains nobody scripted ahead. No single model can be careful and aggressive and frugal all at once. A real business needs all three, so it needs more than one model.
Method note: the six companies are synthetic identities built for the test, not real firms. Decisions were open-ended and scored by blind AI judges over two independent passes. This is development-grade research, shared to show our thinking, not a peer-reviewed result. Every chart and number is exact from the runs.