The gap between what AI agents do in a study and what they do in an actual office is now well documented, and it's the most useful number in the field. Controlled studies show task productivity gains of 20% to 60%. Experiments run inside real workplaces land closer to 15% to 30%.
Neither number is bad. But the space between them is where all the practical guidance lives, because it holds everything a lab strips out — interruptions, approvals, colleagues who missed the memo, and processes that were never built to have a machine sitting in the middle of them.
What the biggest dataset shows
Microsoft's 2026 Work Trend Index pulls from trillions of anonymized Microsoft 365 productivity signals plus a survey of 20,000 AI-using workers across ten countries. Its core finding is that the value of agents shows up less as time saved on individual tasks and more as a change in what people spend their attention on. As agents pick up execution, the human job drifts toward directing the work, making the calls, and owning the result.
That's a real shift, and a demanding one. Directing work is a different skill from doing it — and plenty of jobs aren't organized to let the person doing them exercise much judgment in the first place.
Why most rollouts disappoint
The pattern in the research is consistent. Drop an agent into an existing process and you get a modest improvement. Rebuild the process around the agent and you get a big one.
That's the argument in a Harvard Data Science Review analysis of what it calls the agent-centric enterprise: the headline multiples depend on workflow redesign, not on how good the model is. Its examples are specific — a global industrial firm that cut audit reporting time by 92%, a B2B sales team that scaled its strategic insight work — and in each case the process got rebuilt rather than decorated.
Which is awkward for most teams, because redesigning a workflow takes authority individual contributors don't have. That's why so much agent adoption stops at the level of personal shortcuts.
What's working right now
Fully autonomous agents running complex knowledge work are still premature. Teams getting real results keep humans in the loop at the decision points, and they point agents at specific, repeatable processes instead of open-ended responsibilities.
Custom agents built around proprietary data and company-specific workflows beat generic tools, for a reason that's easy to miss: most of the value in a repeatable process sits in the details unique to your company, and a general-purpose assistant can't see any of them.
A working deployment tends to look like this. Pick a process that runs often, has a clear input and output, and currently eats hours of skilled attention. Map it before you automate it. Put the agent on the mechanical middle and keep the human where a wrong answer gets expensive. Then measure time saved against time spent checking the work — the number almost nobody collects.
The failure mode to watch for
The usual bad outcome isn't an agent doing something wrong. It's an agent doing something plausible that nobody checks, inside a process where checking was never anyone's job.
Every hour saved on execution has to be partly reinvested in verification. Teams that hold their gains are explicit about where that happens and who owns it. Teams that skip it post great numbers for a quarter and pay for it later, usually somewhere hard to trace back.
A realistic ambition
For one person, 15% to 30% on the tasks an agent can genuinely absorb is a solid improvement, and roughly what the evidence supports.
For an organization, the bigger multiples are real — but they sit on the far side of redesigning how the work gets done. That's an organizational project, not a software purchase.
Image: Anastasia Shuraeva, via Pexels





