Insights
Why AI Breaks Traditional Performance Measurement
A university department I have been thinking about spent last year arguing about the wrong thing.
The argument was about cheating. Students were submitting coursework written with AI, and the faculty split into the two camps every institution ends up with: ban it and police it, or permit it and hope. Both assumed the question was integrity. It was not. It was measurement. For a century, the essay worked as an assessment instrument because it was expensive. Producing eight coherent pages was slow, and the slowness was the point. The department never wanted essays. It wanted evidence that a young person could hold an argument together under pressure. The artefact was cheap to grade and hard to fake, so it stood in for the thing that mattered. When production costs collapse, the proxy stops proxying. The artefact still arrives. It just no longer carries information about the person who submitted it. This is not an education problem. It sits in every performance review in every company right now — the deliverable arrives, and the organisation can no longer tell what the human did.
Two pieces of evidence make me think the exit
The first is a study from a Wharton-led team who ran a five-month randomised trial across ten Taipei high schools teaching Python. Every student had the same AI tutor and the same materials. The only variable was the sequence in which practice problems were assigned — fixed in one group, adjusted in the other based on how each student was actually engaging. That single design change raised final exam scores by 0.15 standard deviations, an effect the authors compare to six to nine additional months of learning. No extra instructional hours. No extra teacher workload. The tool was identical. The design around it was not. The design was the entire gain. The second is less comfortable. The OECD’s most recent Survey of Adult Skills, covering roughly 160,000 adults across 31 countries, found literacy and numeracy declining or stagnating in most of them over a decade — sharpest declines among the lowest-performing tenth, improvement at the top. Skills inequality widened within countries. Roughly one in five adults can manage only simple texts or basic arithmetic. So we are placing a technology that produces fluent output into a population whose capacity to evaluate fluent output is, for many, going the other way. That does not automatically make anyone smarter. It makes design decisions enormous.
It is a contribution instrument
What the department eventually did was small and, I think, correct. It stopped grading the artefact alone and started grading the decisions behind it. Students still submit the essay. They also submit a short record of what they tried and abandoned, which sources they rejected and why, and where they disagreed with the model. Then they spend twenty minutes defending the argument to someone who pushes back. The assignment did not get harder to write. It got harder to fake, because the object of assessment moved from the output to the reasoning behind it. Grading time went up. The faculty decided that was the price of measuring anything real. Notice what this is. It is not integrity policing. It is a contribution instrument — a way of making visible the judgment a person exercised, in a world where the product of that judgment is no longer scarce. Every organisation will need one. Your appraisal forms were built in the essay era. They measure artefacts because artefacts used to be expensive, and they will keep measuring artefacts long after that stops meaning anything, because nobody has proposed a replacement. Universities are being forced to solve this in public, on a deadline, in front of students who notice immediately when a system is pretending. The rest of us get to solve it more slowly and less honestly. I would rather learn from them than repeat them.