Your AI coding dashboard probably looks healthy. Active users: up. Suggestions accepted: up. AI-assisted commits: climbing. The rollout worked.
So why can’t you answer the question your CTO actually asked? Is AI making the engineering org faster?
The uncomfortable answer is that most AI measurement systems stop at code generation, which is the first step in a pipeline that runs through review, testing, merge, deployment, production, and maintenance. Optimizing the first step and calling it done is how you miss the bottleneck that formed three steps downstream.
The Pipeline Problem
Consider what happens when an AI tool cuts coding time by 30%. That sounds like a win. But if the resulting pull requests are larger, review takes 25% longer, and rework increases, the net effect on delivery is not obviously positive. The speed gain at coding became a queue buildup at review.
Author Emre Dundar frames this precisely: a local productivity gain is not necessarily a system productivity gain. Software engineering is a pipeline, and accelerating one stage can expose or create a bottleneck somewhere else.
The math is not hypothetical. Dundar provides a concrete example of what this looks like in practice:
| Metric | Change |
|---|---|
| Coding Time | -25% |
| PR Throughput | +30% |
| PR Pickup Time | +18% |
| Review Time | +27% |
Looking only at coding activity, AI looks highly successful. Looking at the full system, the bottleneck moved from coding to review. That is not a failure of the AI tool. It is a failure of the measurement frame.

What AI-Assisted Code Percentage Actually Tells You
The percentage of changes flagged as AI-assisted is becoming a real engineering signal, but not for the reason most people assume. It does not mean the team is proportionally more productive. What it gives you is an analytical dimension.
Once you can tag work by AI contribution, you can ask comparative questions:
- Do AI-assisted pull requests take longer to review?
- Are they larger on average?
- Do they generate more rework shortly after merge?
- Do they correlate with more or fewer failed deployments?
- Does coding time drop, and does that drop reach production as faster lead time?
AI contribution as a standalone number is just another activity metric. Joined with engineering outcomes, it becomes a diagnostic tool.
Where the Work Actually Moved
One of the persistent errors in developer productivity measurement has been treating visible artifacts as useful work: commits, lines of code, tickets closed. AI makes that assumption more dangerous because producing artifacts is getting dramatically cheaper.
The better question shifts from how much developers produced to where engineering effort moved. An AI assistant might reduce time spent on boilerplate, documentation searches, initial test writing, and repetitive refactoring. At the same time, it may increase time spent on verification, code review, debugging, architectural checking, and security review.
If 40 minutes disappear from implementation but 25 minutes appear in verification, the productivity gain is not 40 minutes. And if that verification work falls on a different developer, measuring only the original developer’s output will miss it entirely.

️ The 7-Layer AI Engineering Funnel
Instead of chasing a single AI productivity number, Dundar proposes a funnel that runs from tool access down to economic outcomes. Each layer asks a different question.
1. Exposure
Who can use AI? Licensed users, eligible developers, available tools.
2. Adoption
Who actually uses it? Active AI users, weekly usage, tool and model adoption by team.
3. Contribution
Where does AI participate in engineering work? AI-assisted changes, AI-assisted commits, AI-heavy pull requests, AI-assisted development rate.
4. Flow
What happens to development speed? Coding time, PR cycle time, PR pickup time, review time, throughput, work item cycle time.
5. Quality
What happens after code is created? Rework, reverts, defects, maintainability issues, security findings, test failures.
6. Delivery
Does the organization ship differently? Change lead time, deployment frequency, change fail rate, recovery time, deployment rework.
7. Outcome
Did something economically meaningful change? Engineering capacity, delivery predictability, customer outcomes, engineering cost, AI cost, time to market.
The further down the funnel, the closer you get to actual organizational impact. Attribution also gets harder the further down you go. That is unavoidable, not a reason to stop measuring.
Why There Probably Isn’t One AI Productivity Score
The instinct to compress everything into a single number is understandable. It is also dangerous. AI affects multiple dimensions that can move in opposite directions simultaneously. Dundar provides an example of what a real AI measurement snapshot might look like:
| Metric | Change |
|---|---|
| Coding Time | -21% |
| PR Throughput | +17% |
| Review Time | +14% |
| Rework | +9% |
| Deployment Frequency | +8% |
| Change Fail Rate | +3% |
| Developer Satisfaction | +16% |
Is AI working here? That is an interesting engineering conversation. Compressing those seven numbers into a score of 83 destroys the information needed to answer it.
Practical Rules for the Measurement System
A few principles from the framework worth applying directly:
Cohort comparisons beat company averages
AI may be highly effective in one codebase and much less useful in another. Segment by repository, type of work, team, AI contribution level, PR size, and time period. The goal is not to manufacture causality. It is to eliminate obviously misleading comparisons.
Correlation is still useful
If high-AI pull requests repeatedly show faster coding time, larger PR size, longer review time, and more rework, that pattern does not prove AI caused it. But it gives you a useful hypothesis: maybe AI-enabled workflows need smaller pull requests or stronger automated validation. You can now change the system and observe what happens. Measurement becomes an improvement loop rather than a performance judgment.
Measure teams before individuals
AI telemetry makes extremely granular individual measurement possible. That does not mean it should become a management metric. Metrics like AI-generated lines per developer, prompt count per developer, or acceptance-rate rankings create targets. Once people know a metric affects how they are evaluated, it stops behaving like neutral telemetry. Start with teams and workflows. Use deeper data diagnostically when a team needs to understand why something changed.
Every speed metric needs a counter-metric
Pair every acceleration signal with a balancing signal:
| If you measure… | Also measure… |
|---|---|
| Coding Time | Rework |
| PR Throughput | Review Time |
| PR Size | Review Load |
| Deployment Frequency | Change Fail Rate |
| AI Contribution | Quality |
| Cycle Time | Developer Experience |
| AI Cost | Delivery Outcome |
Optimization in engineering frequently transfers cost. A team can increase deployment frequency by making smaller deployments (good) or by bypassing necessary controls (not good). The number alone cannot tell you which happened.
The Minimal Dashboard Worth Building
Dundar recommends starting with nine signals rather than forty:
- Active AI Developers: are people using the tools?
- AI-Assisted Development Rate: where is AI participating in actual engineering work?
- PR Cycle Time: is work moving through development faster?
- Review Time: did the bottleneck move downstream?
- Rework Rate: are we creating additional follow-up work?
- Change Lead Time: is local acceleration reaching production?
- Change Fail Rate / Deployment Rework: are faster changes remaining reliable?
- Developer Perception: do developers feel the workflow improved?
- AI Cost: what are we paying to create those changes?
Nine signals across adoption, flow, quality, delivery, experience, and cost. That is already enough to have a substantially better conversation than your current AI dashboard allows.
When This Works and When It Does Not
This framework works best when your engineering organization has some baseline measurement maturity: you are already tracking PR cycle time, deployment frequency, and change fail rate. If you are starting from zero, the funnel gives you the right build order.
One genuine limitation: controlled experiments have shown large improvements in AI-assisted task completion speed, while real-world studies of experienced developers in mature repositories have found smaller gains, no gains, or temporary slowdowns. Those results are not contradictory. They measure different environments. A bounded programming task is different from changing a mature production system with undocumented architectural decisions, legacy dependencies, and operational constraints. Statements like “AI makes developers 30% faster” are nearly meaningless without specifying which developers, performing which tasks, in which codebases, using which tools, measured at which part of the delivery system.
AI attribution without engineering outcomes tells you what AI did. Engineering outcomes without AI context tell you what changed. The missing middle between those two, connecting AI participation to delivery impact, is where the actual measurement work now lives.



