Insights
Designing the Handoff After AI Stops
Most organisations have decided where AI stops. Almost none have designed what happens in the second after it stops.
The handoff gets treated as routing: the system reaches its limit, a human receives the work, service continues. It is not routing. It is the point where the quality of the whole system is decided, and it is currently being designed by nobody. There is now field evidence on this. In a randomised experiment on Alibaba’s Taobao platform, workers supervised an agentic AI system handling eligible customer chats while continuing to resolve the rest themselves. Chats got shorter. Ratings for the AI-eligible chats fell substantially. What is interesting is where the human rescue worked and where it did not. When the system escalated because the problem exceeded its capability — a technical dead end — human intervention preserved service quality. When it escalated because the customer had become frustrated, intervention largely failed to recover the outcome. The reason was not skill. It was effort. In emotional escalations the workers sent fewer messages, took a smaller share of the conversation, sought less information, offered fewer solutions. They arrived at the hardest moments of their day with the least engagement. And intervening late made it worse: early entry was what sustained effort afterwards.
Read that as an organisational design finding
Read that as an organisational design finding rather than a customer service one. The system had quietly reallocated human work to consist almost entirely of the parts machines fail at — the emotionally expensive parts — without changing anything about how that work was supported, staffed, or valued. The residual became the job. Nobody designed the residual. A second study, a field experiment inside Porsche’s after-sales operation, points at the same gap from another angle. Decision-makers were given advice from a human expert, a machine learning model, or both. The combination produced better decisions — but only where people actively reconciled the conflicting inputs rather than deferring to one. And engagement in that reconciliation dropped when the decision moved from an individual to a group. Together these say something uncomfortable. The benefit of keeping humans in the loop is not automatic. It is conditional on an act of genuine engagement that most workflows do nothing to protect, and that quietly evaporates under pressure, fatigue, and group settings.
Three things follow
So the question is not whether to keep a human in the loop, but what the loop asks of them. Three things follow. Escalation should carry context, not just the case. A person entering at the worst moment with none of the history spends their first minutes reconstructing what happened, and reaches the human part of the problem already depleted. Handoffs designed for continuity of information are cheap. We rarely build them. The residual should be staffed as difficult work, not leftover work. If machines absorb the routine and leave humans the ambiguous, the emotional and the unprecedented, then the human caseload has quietly become harder per hour while headcount models still assume it became easier. That mismatch shows up as disengagement long before anyone calls it burnout. And disagreement with the system needs a legitimate route. If reconciling machine advice with your own judgment is the step that creates the value, then overriding must be an ordinary, low-cost, recorded act — not a quiet act of insubordination someone must justify.
The pattern underneath is worth watching everywhere AI enters
The pattern underneath is worth watching everywhere AI enters an organisation. We are removing the easy parts of jobs and calling it relief. What remains is the concentrated human residue: judgment, ambiguity, emotion, repair. That is the contribution we say we want. It is also, hour for hour, the most demanding work a person can do — and we are handing it over with no more support than the routine work it replaced.