Measuring AI’s Impact on Delivery Speed
September 15, 2026
Is AI making software development faster? Most people I talk to say “yes.” But how much? What value are you getting for your money? Are there other ways of using AI that have a better ROI?
This is part 2 of my series on quantifying AI’s impact on software development. In part 1, we looked at why assessing impact is important, and what not to do. In this part, we’re looking at how to measure AI’s impact on delivery speed. In part 3 (coming September 22nd), we’ll look at how to measure—and forecast—the unintended consequences of AI, before wrapping up the series in part 4 with an examination of business outcomes. Finally, an epilogue puts it all together with notes you can share with your CFO.
To be notified when next week’s update comes out, add my feed to your RSS reader or subscribe to my free mailing list. Details here.
What We’re Measuring
Delivery speed isn’t productivity. I’ll say it again, loud, for the folks in the back: Delivery speed isn’t productivity!
It’s a component of productivity, sure. Given two equivalent teams creating the same piece of software, the team that delivers first is the one that’s more productive.
But it’s not just about output. In my mind, software development productivity is value produced divided by money spent. It’s a strict definition, but a useful one, because it gets to the heart of why we build software: to achieve some valuable purpose. Going faster doesn’t mean you’re going in the right direction.
We’ll consider business outcomes in part 4. For now, though, we’ll tackle a simpler question: are your teams finishing their assigned work faster, regardless of whether that work is actually useful or not?
The Core Approach
First, the bad news: there’s no objective measure of software development speed. I can’t look at two teams and say, “this one is fast, and that one is slow.” That’s because none of the outputs teams create are standardized. We already talked about how lines of code and PRs are bad metrics in part 1. Well, nothing else is standardized, either.
Features? Some are big, some are small. Bugs? Some are easy, some are hard. Stories? Velocity? Story points? Completely arbitrary. They have way more to do with what’s being requested than the speed of the team.
Okay, if they’re all different sizes, couldn’t we just use estimates? Surely, if a team is constantly missing its estimates, it’s because they’re slow, right?
Not so fast! Software estimates are famously inaccurate. A team that always misses its estimates isn’t slow. It’s just over-optimistic. Teams are nearly always over-optimistic.1
1When they’re not, it’s usually because they’re playing political games with estimate padding. That isn’t good either, because then the work tends to expand to meet the estimate.
So we can’t count requests (features, bugs, stories, etc.) because they vary in size, and we can’t look at on-time delivery, because estimates are over-optimistic. Sounds like a whole lotta “can’t.” What can we do?
We can compare a team to itself. If a team’s estimates were perfectly consistent, and they went from taking 3x longer than estimated to 1.5x longer, then they’re twice as fast! Huzzah!
If only it were that simple. There’s a few snags.2
2This design was inspired by a METR study, with help from Brent Miller. AI Disclaimer: I used OpenAI’s GPT-5.6 Sol model operating in “High” thinking mode to critique my explanation, and I also used it to generate the Python programs that drew the graphs below. You can see my ChatGPT conversation here.
Estimates Aren’t Consistent
Snag the first: teams’ actual:estimate ratios aren’t actually consistent. Some are bigger, some are smaller. Plot them on a graph, and they form an “S” curve, as with this data from a real company:3
3This data is from the Star Citizen alpha 3.0 release, which publicized its internal task estimates and actuals. Interestingly, the sharp cliff at the 1.0 mark is a classic sign of work growing to meet its estimate.
If delivery speed improves, the curve will shift. Individual data points may be better or worse than the original approach, but you can see the effect in aggregate.4
4The Star Citizen data was split in half and artificially manipulated for the purpose of this example. It doesn’t represent real-world AI impact.
People Work Differently
Some people are pessimistic and estimate high. Some estimate low. Some people spend a lot of time educating others. Others work heads down as fast as possible. Some clean up messes. Others create messes.
If you do a blanket comparison of everyone to each other, you introduce a lot of noise. Noise is bad because it increases the amount of data you have to collect. To keep things clean, compare people (or teams) to themselves.
Estimates Change When People Know They’re Using AI
Third snag: people’s estimates will change according to the work they plan to do. So if they think they’re going to use AI, and they think AI makes them faster, their estimates will be smaller. Huh. Tricky.
This one’s actually pretty easy: collect the estimate before people know whether they’re going to use AI or not. Then randomly assign the task to be done with or without AI.
Process Improvements Muddy the Waters
Fourth and final snag: Are your speed increases due to introducing AI? Or the other changes you made at the same time?
AI forces you to rethink your software development processes. Leaders I talk to are streamlining request workflows and putting business experts in closer touch with developers as part of their efforts to introduce AI.
But these are the exact changes that make software development more effective in general! How much of the improvements you’re seeing are from spending thousands of dollars on tokens... and how much are just because people are talking to each other more?
Rather than comparing a chunk of “pre-AI” data to a chunk of “post-AI” data, interleave “without AI” tasks and “with AI” tasks. This reduces the influence of other changes. Update your “without AI” approach to include the same streamlined workflows as your “with AI” approach.
The Complete Measurement Approach
In other words, if you want to be rigorous about this, you need to conduct a randomized controlled trial. Specifically, a “randomized within-subject repeated-measures” trial, which reduces the number of samples you need to collect.5
5Many thanks to Brent Miller for introducing me to this technique.
It’s not as bad as it sounds! People still do their normal work.
Here’s what you do. First, establish a with AI approach and a without AI approach. (Or, more likely, “with a lot of expensive AI” and “with a sprinkling of cheap AI,” but I’ll say “with” and “without” for clarity.) Make sure they’re clearly defined and people have had a chance to practice both approaches.
The “without AI” approach will probably be close to your current approach to software development, but if you’re streamlining your “with AI” processes, be sure to incorporate those improvements into your “without AI” approach, too. Decide how you’re going to attribute rework, such as code review changes and bug fixes, back to the task that caused the rework.
Then give each person (or team) a virtual bag that’s half full of “with AI” assignments and half without. For any given task, have people record their estimate first, then randomly pull an assignment from the bag.
The main data you’ll collect is estimated time spent, actual time spent, type of task, and spending. Decide in advance if you’re tracking calendar time (including interruptions) or effort (excluding interruptions) and how you’re going to handle parallel work. Think about what other useful information you might want, such as bug and incident counts, mental energy, excess lines of code, and so forth. We’ll talk about using that data to model unintended consequences in part 3.
You’ll need the help of a statistician to do the detailed experimental design and data analysis, but broadly speaking, each person’s “with AI” data will be compared to their own “without AI” data, then aggregated together to give you the final results.
Finally, remember that speed increases aren’t the same as productivity. Pay particular attention to how the tasks you’re measuring fit into the big picture. For example, if you see a 2x improvement in coding speed, but people spend 50% of their week on other work, their overall speed improvement is only 1.33x, not 2x.6
6Work that took 40 hours now takes 30. 40 ÷ 30 = 1.33.
Is It Worth It?
This is a lot more complicated than just counting PRs and calling it good. Is it worth it?
If this were just some run-of-the-mill process improvement exercise, no, it wouldn’t be worth it. It’s too much work, and the effect sizes have to be pretty large in order for the data collection to be practical.
But this isn’t some run-of-the-mill process improvement exercise. Productivity expectations are off the charts, and so are costs. As I talk to engineering leaders, I’m seeing order-of-magnitude differences in spending. Should you spend thousands per person per month on hands-off agentic coding? Or hundreds on analysis and fancy auto-complete?
It depends on the outcomes.
You can’t pick the right approach to AI without knowing what you’re getting for your money, and with so much money on the table, you can’t afford to half-ass this. Remember, in a METR study, engineers overestimated AI’s speed benefits by nearly 50%. Without measurements, there’s no way to know what you’re actually getting.
Fortunately, it’s not as hard as it seems. With the level of impact you need to see, and enough participants, it doesn’t take many samples from each person.7 Your teams collect the data as they do their normal work. The hardest part is explicitly defining how you want people to use AI, and that’s worth doing anyway.
7The exact number will depend on the number of participants, the amount of variance in your results, and how conclusive you want the results to be.
Of course, delivery speed is only part of the picture. There’s also the question of how sustainable speed improvements are. We’ll consider unintended consequences next week, in part 3.
To be notified when next week’s update comes out, add my feed to your RSS reader or subscribe to my free mailing list. Details here.

