Insights
The AI Productivity Performance Paradox
What if the most valuable person on your team is the one who looks least productive?
A field experiment published in Organization Science tracked 750 knowledge workers using GPT-4. Workers with AI were 25 percent faster and produced higher-quality output on tasks within the model’s capability. But on a task just outside that capability — where the AI was confidently wrong — workers who relied on it were 19 percent less likely to reach the correct answer than those working without it. Here’s what makes this uncomfortable: the people who slowed down to verify the AI’s output, who questioned it, who caught the errors — they looked less productive by every traditional metric. Less output per hour. Fewer deliverables. Slower turnaround. They were the ones adding the most value.
The performance paradox
Harvard Business Review called this the performance paradox. Organizations still evaluate employees with pre-AI metrics — productivity, goal completion, tasks per hour — and those metrics now reward the wrong behavior. They reward uncritical speed over careful judgment. Consider what this means at scale. Gallup’s 2026 State of the Global Workplace report — 263,000 respondents across 160 countries — found that only 20 percent of employees are engaged at work. Eighty percent are either not engaged or actively disengaged. We tend to read that as a motivation problem. I wonder if it’s a measurement problem.
When organizations measure output
When organizations measure output, they get output. When they measure hours, they get hours. When they measure compliance, they get compliance. What they don’t get is thinking. They don’t get the engineer who notices a systemic flaw but says nothing because flagging problems slows her numbers. They don’t get the analyst who sees a strategic connection across departments but has no forum to share it. They don’t get the manager who could redesign a process but is too busy being measured on the old one. Every organization is full of unused intellectual capital. Not because people lack ideas — but because nothing in the system asks for them. This is the shift I keep returning to. For two centuries, organizations built their operating systems around execution. Performance was the measure of value. And it worked — when execution was scarce.
The bottleneck is contribution
But AI is making execution abundant. McKinsey’s global managing partner reported that AI saved his firm 1.5 million hours of search and synthesis in a single year. Twenty-five thousand AI agents produced 2.5 million charts in six months. The execution capacity is no longer the bottleneck. The bottleneck is contribution — the thinking, questioning, connecting, and judgment that no model can replicate. And yet our measurement systems still point in the other direction. I wonder whether the most important organizational innovation of the next decade won’t be a new technology or a new strategy. It may be a new unit of measurement. Not output. Not productivity. Not even performance. Something closer to: did this person make the organization think better? We don’t have that metric yet. But the organizations that figure it out first will have an extraordinary advantage — because they’ll be the ones who finally learn what 80 percent of their people have been waiting to contribute.