Insights

The Judgment Gap in Organizations

By Adam G·

We have spent a century building organizations that reward people for knowing, and almost no time building organizations that reward people for noticing they don’t. That asymmetry used to be affordable. It may not be much longer.

Earlier this year a research team at Scale published a benchmark called HiL-Bench. The premise is simple: take hard software and database tasks that frontier AI models already solve well, then remove a few details — the kind missing from every real specification ever written: an ambiguous requirement, a contradiction between documents, a fact nobody wrote down.

With complete information, the models solved between 75% and 89% of the tasks. Given the ability to ask a human for clarification, but left to decide for themselves when to ask, performance fell to 38% on one domain and 12% on the other.

What is striking is not the failure but its manner. The systems did not stall or flag uncertainty. They filled the gap with a confident assumption and continued. Across thousands of failure traces, one model family built elaborate work on beliefs it never examined; another recognized it was stuck and pressed on anyway.

A judgment gap, not a capability gap

The researchers named the missing thing well: a judgment gap, not a capability gap. The mechanism to ask existed. The judgment about when to use it did not.

I think most organizations have the same gap, and have had it far longer than the machines.

Consider how we select people. We ask for evidence of what they have already produced — degrees, titles, throughput, prior scope. We rarely ask for evidence that they can tell when a brief is incoherent, or when a plan assumes something that was true two years ago. The Burning Glass Institute and OneTen studied over a thousand large U.S. employers this spring and found that firms which formally dropped degree requirements saw only a two-percentage-point increase in hires holding non-degree credentials. The policy changed; the selection instinct did not. Fifty-eight percent of the prime-age American workforce has no four-year degree, and the machinery keeps reaching for the same signal.

So we assemble organizations optimized for confident execution, then act surprised when confident execution turns out to be the failure mode.

Wrong assumptions now compound at machine speed

Why does this matter more now? When execution was expensive, a confident wrong assumption travelled slowly. It passed through people, drafts, meetings, budgets — each a chance for someone to ask, quietly, “wait, what do we actually mean by this?” Execution is now cheap and fast. Wrong assumptions now compound at machine speed. The bottleneck moves upstream, to the quality of the question the work began with.

Which makes noticing a form of production — perhaps the most valuable one. And noticing does not come from expertise alone; it comes from having lived in more than one system. The person who worked in a warehouse before they worked in finance sees the supply-chain assumption the rest of the room reads past. The nurse-turned-administrator hears which part of the policy will not survive a Tuesday night. The immigrant hears the sentence in the deck that only makes sense to people who grew up here. This, to me, is the real argument for diverse organizations — not that difference is admirable, but that difference is the only reliable detector of the assumptions a homogeneous group cannot see because it shares them.

Judgment turned out to be trainable

The encouraging part of the benchmark is its last finding: judgment turned out to be trainable. A far smaller model, trained to recognize unresolvable uncertainty and act on it, improved not only its asking but its results.

If it is trainable in a model, it is cultivable in an institution. But you cannot cultivate what you never record. Which returns me to a question I keep circling: does your organization keep any record of the moment someone stopped the work to ask a better question? Or does it only remember the people who kept going?

More insights

Ready to build your AI Revenue Organization?

Book a strategy call. We’ll map the BeyondOS™ departments to deploy and the human contribution layer that makes them more valuable.

Build My AI Revenue Team