Load BearingWhat actually carries the weight

01 — Technology & Work


AI Tools at Work: What Has Actually Changed, and What Has Not

There is now a substantial body of controlled research on what generative AI does to knowledge work. It is more interesting than either the promotional version or the dismissive one, and it points somewhere specific: the gains are large where tasks are structured and checkable, largest for people who were weakest at the task, and can turn negative for experts on familiar work.

A note on sources before the numbers. This subject is saturated with vendor statistics, and the figures circulating vary wildly — adoption rates from 45% to 91% depending on who is counting and what they count as use. Below, the peer-reviewed studies and central bank research are separated from survey material.

For a practical reference on employee time clock software, Monitask provides this page.

Where the gains are real

Controlled workplace studies consistently show large productivity gains in writing, customer support, software development and translation, often improving quality as well as speed.

Noy and Zhang, published in Science, found AI assistance improved writing task completion time by 40% on average — with considerable variance across workers.

Stanford's 2026 AI Index aggregates results across studies and finds the size of the gain tracks closely with how structured and measurable the work is: roughly 14–15% in customer support, 26% in software development, and much higher in marketing output — while gains shrink substantially in work requiring deeper, less-structured reasoning.

The pattern: AI adds the most value where the task has a clear right answer and an easy way to check it, and the least where judgment calls dominate.

The finding that matters most

The evidence consistently points to beginners. In customer support the split was roughly 34% improvement for novices against near-zero for experienced staff.

That is the single most useful thing in the literature, and it is nearly absent from the coverage.

It reframes what these tools are. Rather than a general accelerator, they behave more like a mechanism for distributing what good performers already know — bringing the bottom of a distribution up rather than raising everyone equally.

The implications are practical and they run in an unexpected direction:

Your least experienced people benefit most. Deployment aimed at senior staff may produce the smallest measurable effect.

Onboarding and training change shape. If a tool closes part of the gap between a new hire and an experienced one, the ramp looks different — and so does what the experienced person is for.

Career development gets harder to reason about. If the early-career work that used to build judgment is the work most easily assisted, the path to expertise changes. Nobody has good evidence on this yet.

Where it goes the other way

Tacit, expert-level or highly familiar work can see AI reduce productivity — as METR's 2025 randomised controlled trial found for experienced developers.

That result is worth sitting with. Experienced developers, on their own familiar codebases, were slower with AI assistance while believing they were faster.

Dell'Acqua and colleagues describe a "jagged technological frontier": AI improves performance on tasks within its capability and harms performance when applied beyond it. The problem is that the boundary is invisible from inside — the output looks equally confident either side of it.

The gap between perceived and measured

A survey of nearly 750 corporate executives documents a productivity paradox in which perceived productivity gains are larger than measured productivity gains, likely reflecting a delay in revenue realisation.

Gallup's February 2026 survey of 23,717 US employees found that employees who use AI frequently say it improves their productivity, but evidence that AI has fundamentally changed how work gets done across organisations remains more limited.

Self-reported gains are not the same as measured gains, and the METR finding shows they can point in opposite directions.

Time saved is not value captured

The most rigorous population-level estimate, from the Federal Reserve Bank of St. Louis, puts time savings at 5.4% of work hours — roughly 2.2 hours in a forty-hour week.

Time saved only becomes productivity if it flows back into meaningful work rather than dissolving into scattered follow-ups, tool-switching and administrative residue.

This is the same lesson every previous workplace technology taught. The hours recovered do not automatically go anywhere useful; they go wherever the organisation's defaults send them.

McKinsey found that workflow redesign is the factor most associated with financial impact from AI — and that only 21% of adopters have done it.

Why most deployments stall

MIT NANDA's research points to a lack of adaptability: most enterprise tools cannot retain feedback, learn from context or adjust to a specific workflow, so they stall in the pilot phase — only 5% of pilots reach production.

The bottleneck has shifted from technical availability to human integration — how people actually incorporate the tools into their work processes.

What has not happened

The employment picture is the one where the gap between expectation and evidence is widest.

Studies using administrative records and large surveys find little evidence of economy-wide job loss or wage decline despite rapid adoption. The Budget Lab at Yale finds no clear relationship between AI exposure and unemployment through August 2025. Humlum and Vestergaard link survey-reported ChatGPT use to Danish administrative records across eleven exposed occupations and find essentially zero effects on earnings or hours through 2024.

The executive survey likewise finds little evidence of near-term aggregate employment declines, though larger companies anticipate AI-driven workforce reductions while smaller firms expect modest gains.

Consistent with a task-based view, the analysis characterises current systems as primarily augmenting human labour rather than automating it outright.

None of this settles the long run. It does mean that confident claims about displacement happening now are not supported by the administrative data.

See automation and job displacement: how to read the evidence.

What to do with this

Deploy where the work is structured and checkable, and expect little in work dominated by judgment.

Expect the gain to concentrate among less experienced staff, and measure it that way rather than in aggregate.

Watch for the frontier. Experts on familiar work may be slower and will not notice. If output quality matters more than speed in a role, measure quality.

Redesign the workflow or accept a fraction of the benefit. This is the finding with the largest effect size in the whole literature and the one least often acted on.

Measure, do not survey. Self-reported productivity is systematically optimistic here — demonstrably so.

Treat vendor statistics with suspicion. Adoption figures spanning 45% to 91% in the same year are measuring different things, and most secondary roundups miscite the underlying studies.

The honest summary

Real, substantial, uneven. Largest where tasks are structured and where the person was weakest. Sometimes negative for experts on familiar work. Mostly not yet visible in aggregate productivity or employment statistics. And overwhelmingly dependent on whether the organisation changed how the work is done, rather than on which tool it bought.

For broader public guidance and background, consult the National Institute of Standards and Technology.