Work · Opinion
Your developers swear AI makes them faster, and the stopwatch disagrees
Developers overestimate AI speed-ups, and one engineer's month off the tools shows why. Measure merged, stable output before buying more seats or cutting staff on the promise.
The most expensive number in software right now is 20%. That is how much faster the 16 experienced open-source developers in METR's July 2025 trial believed AI tools had made them, after working through 246 tasks on their own repositories. The clock said the AI-assisted tasks took 19% longer. Before they started, the same developers had predicted a 24% speed-up. They lived through the slowdown and still came out convinced they had gained a fifth. If the people doing the work cannot feel the difference, their managers certainly cannot, and the vendors have no reason to point it out.
I read a developer's account this week of quitting AI tools for a month, and I expected the usual complaint about sloppy code. The code complaints are there. He says not one task he handed an agent could be merged unchanged, and fewer than half were solved at all. What stuck with me as someone who thinks about who pays the invoice was the arithmetic. A job he could have done by hand in 20 minutes took an agent five minutes and then took him two days to review, because he was juggling several agents across several worktrees and could not hold the context. One stalled agent burned $30 of tokens over 30 minutes, then produced the answer in 20 seconds once he told it off. He stopped reading the seven-paragraph PR descriptions the agent wrote for him. A colleague eventually caught a test that did not exercise the change it claimed to cover, from a man who had practised test-driven development for more than ten years. He describes himself at that point as "nothing but an AI shepherd".
He felt powerful the whole time. That feeling is what gets written into budget requests.
Follow the money through a typical company. The vendor sells seats and bills tokens. The engineering director reports adoption, because adoption is easy to count and looks like progress. The CFO sees a per-seat price that is small next to a salary and hears that output is up. Nobody in that chain is paid to measure how many changes shipped, stayed shipped and did not wake someone up at night. Google's DORA 2024 report estimated that a 25% rise in AI adoption was associated with a 1.5% drop in delivery throughput and a 7.2% drop in delivery stability, while most respondents said they felt more productive. Stack Overflow's 2025 survey found 84% of developers using or planning to use AI tools, yet more of them distrusted the output (46%) than trusted it (33%), and 66% named answers that are almost right as their biggest frustration. GitClear, looking at roughly 211 million changed lines, found an eightfold rise in duplicated code blocks during 2024. Someone maintains that code later, and it is usually a salaried human.
Headcount is where self-reported productivity does real damage. In February 2024 Klarna said its AI assistant was doing the work of 700 customer-service agents. By May 2025 Sebastian Siemiatkowski was telling Bloomberg that the cost focus had cost quality, and Klarna was hiring people again. Different job, same mistake: a company cut staff against a productivity claim before checking what customers actually received.
The serious reply comes from Cui, Demirer and colleagues, whose field experiments across 4,867 developers at Microsoft, Accenture and an unnamed Fortune 100 company found Copilot users completed 26.08% more pull requests. METR itself warns its sample was small, senior and working in mature codebases they knew well. I accept both points, and I think they cut the other way for operators. A pull request is exactly the unit an agent can inflate; my blogger was producing them by the dozen and merging none as written. PR counts do not see two days of review or a test that tests nothing, which is the sort of cost DORA's stability figure picks up. And the Cui study found the biggest gains among junior developers. Companies freezing graduate hiring to fund AI seats are removing the people the evidence says benefit most, while keeping seniors whom METR suggests the tools may slow down.
So before your next seat renewal, run the test the vendor will not. Take two comparable teams for a month and switch the agents off in one. Ask every developer beforehand to predict their speed-up and write it down. Then count changes merged that are not reverted within the month, incidents traced to those changes, reviewer hours and token spend per merged change. If the agent team does not beat the other on shipped, stable work, cut the seat count at renewal and take every AI-justified role reduction off the plan until someone can show you the numbers.
Prompted by One month without AI, Bustikiller's Blog.