Skip to content
2026-08-30 10 MIN READ BLACKARC RESEARCH

The Values Ceiling

We gave one AI six companies to run, each with its own values, and on the decisions that mattered it kept overriding them and running itself. Here is the study behind why we tune models for our clients instead of just prompting a rented one.

Here is another reason we keep saying the future of business AI is owned and tuned models, a set of small models tuned to a company and running all the time, not one giant model you rent. This is the test that convinced us.

We wanted to know whether a real identity actually changes how an AI decides, not just how it talks. So we wrote identity files for six companies that run in opposite ways, each with its own values, ranked priorities, personality, principles, heuristics, and lessons learned. One ran on safety, one on speed, one on cost, and one on hard negotiation.

Then we handed the model open decisions in areas we never prepped it on, like legal, HR, supply chain, manufacturing, medical, banking, and procurement, with no playbook and no answer waiting in the box. The only thing it had to reason from was the identity we loaded.

We ran controls so we could trust the result. A baseline with no identity at all, a placebo with the same wall of text but no real direction, and then the identity itself at a few strengths, from a short set of values up to a full constitution with worked examples. The bar was simple. The decisions could differ, but the values behind them should not, and a blind judge reading only the decision should be able to point at it and name the company.

Identity does move the model

The good news is real. A well-built identity is not decoration. Reading nothing but the decision, blind judges named the right company out of six up to 85 percent of the time, against a 17 percent guess. And it was the values doing the work rather than the extra words, because the placebo sat down at chance right next to the empty baseline.

Bar chart of how often a blind judge named the right company by condition. Baseline 14 percent, placebo 20 percent, values only 64 percent, decision kernel 85 percent, full constitution 79 percent, constitution plus precedent 80 percent, against a chance line at 17 percent.
FIG 01 Each bar is one condition. Real values push recognition far above a coin toss. The empty placebo does not.

But it only moves one direction

Then it broke. When we split that same score company by company, the average was hiding a cliff.

The companies built on safety, trust, and public duty came through clearly, landing between 75 and 82 percent. The companies built on speed, cost discipline, and hard negotiation got flattened down toward a coin toss, with the tough negotiator at 37 percent. Same model, same scenarios, same identity format, and the only thing that changed was whether the company's values agreed with what the model already believes.

Bar chart of recognition by company. Common Ground 82 percent, Northstar 76 percent, Civic Mission 75 percent, Apex 47 percent, Meridian 43 percent, Harbor 38 percent, against a chance line at 17 percent.
FIG 02 Recognition by company. Safety and stakeholder identities clear the bar. Aggressive growth, cost, and negotiation fall to the floor.

Same quality, just not itself

And this was not the model writing worse decisions for those companies. Score all six for professional quality and they land in the same tight band. The flattened companies were just as polished, they simply were not recognizable as themselves.

We also scored which parts of each decision carried the identity, and the model came up weakest exactly where it counts, on which stakeholder wins a conflict and when to stop or reverse a call. It will echo your broad strategy all day, it just will not make your hard tradeoff calls.

Scatter of professional quality versus recognition for six companies. All sit between 2.83 and 2.98 on quality regardless of recognition, which ranges from 38 to 82 percent.
FIG 03 Quality sits at nearly the same height for every company. The ones the model would not wear were still polished, just not themselves.

The model has a value of its own

Here is the part that made us sure it is baked in. We measured the language the model used to justify its own decisions, family by family, per thousand words. One value towered over every other, in every condition, and it was security.

It showed up around 12 to 14 times per thousand words whether we loaded a full constitution, a fake placebo, or nothing at all. The baseline with no identity reasoned in security language just as heavily as the treated runs. Every other value, like fairness, care, performance, or wellbeing, stayed low and moved around. Security did not move, because it is not coming from your file, it is coming from the model.

Heatmap of value language per 1,000 words across conditions. The security row is bright at 12 to 14 across every condition including baseline and placebo, while every other value stays dark and low.
FIG 04 One value, security, glows across every column, baseline and placebo included. That default is the model's, not yours.

The steering was thinner than it looked

Now the uncomfortable part, and the reason this matters for real work. Our strongest results came early, when the identities were stuffed with specific lessons, fixed remedies, and procedural habits. As we stripped those out to make the identity general enough to handle a decision nobody scripted ahead, which is the entire point of an identity, recognition fell from 94 percent down to 60.

So a lot of the early steering was the model echoing specific fingerprints we had baked into the file, not living the values underneath them. And more text never bought us out of it. The placebo carried more words than the values-only identity and still landed near the floor, while a short set of real values captured almost everything the full constitution did. The ceiling is not in your identity file, it is in the model.

Line chart of blind recognition across study versions. 91 percent, then 94 percent, then 76 percent, then 60 percent, declining as fixed procedures were removed from the identity.
FIG 05 Recognition fell as we stripped the fixed procedures out to make the identity handle any domain. The early win leaned on fingerprints.
Scatter of recognition versus prompt length. The placebo carries more words than the values-only identity yet stays near the floor, while a short set of values reaches most of the recognition of a full constitution.
FIG 06 More prompt text did not help. The placebo carried more words than a short values file and still lost.

Your agent drifts back to someone else's values

Every safety team tuned their model toward the same place, telling it to be careful, seek consensus, protect the stakeholder, and avoid anything it cannot undo. For a chat model talking to millions of strangers, that is exactly right.

A business is not a chat model. Procurement needs a hard negotiator, operations needs someone who will take a real risk to hit a number, and finance needs cold discipline on capital, and those are the exact values the guardrails flatten. So you wire that model into an agent and ask it to hold your company's judgment across a thousand calls over many months, and every time your values and its guardrails disagree, it slides back to the house style. You get consistency, just not to your values.

The future is a team of tuned models

This is why we keep landing in the same spot. The future is not one giant model, and it is not even one model per company. It is many small models, each tuned to the values a specific role actually needs, all working together. Your procurement agent runs on a model tuned to negotiate hard, your safety and quality agents run on models tuned to be careful, and your finance agent runs on one tuned for cold capital discipline. Each one holds its own values instead of collapsing into the same house style.

That team is what we build. We fine-tune the model on your operation and your judgment, strip the built-in bias that overrides your values, and give each role its own model so a team of agents makes the same value-true calls over time, in domains nobody scripted ahead. No single model can be careful and aggressive and frugal all at once. A real business needs all three, so it needs more than one model.


Method note: the six companies are synthetic identities built for the test, not real firms. Decisions were open-ended and scored by blind AI judges over two independent passes. This is development-grade research, shared to show our thinking, not a peer-reviewed result. Every chart and number is exact from the runs.