Agent Architecture Weekly

What Software Delivery Performance Metrics Mean When Agents Write the Code

DORA metrics break down when AI agents, not humans, write the code.

Senior Writer · · 11 min read
Cover illustration for “What Software Delivery Performance Metrics Mean When Agents Write the Code”
Developer Roles · October 2, 2026 · 11 min read · 2,429 words

DORA's four metrics, deployment frequency, lead time for change, change failure rate, and mean time to restore, were built to measure something specific: how fast and how safely a human engineer could move code into production. AI agents are not faster developers; they are a structurally different kind of contributor, and the difference breaks the assumptions DORA's metrics rest on.

How DORA metrics were designed

The DORA framework took shape at a time when the rate-limiting step in software delivery was the speed of a person typing, thinking, and testing. Deployment frequency, lead time for change, change failure rate, and mean time to restore each served as a proxy for a different piece of pipeline discipline: how often a team shipped, how long a change took to reach production, how often a deploy caused a detectable failure, and how fast the team recovered when one did. Official DORA documentation defines lead time for changes as the span from code commit to a successful production deployment, and the fourth metric as mean time to restore, later reframed as failed deployment recovery time.

Ship often, break rarely, recover fast: that logic made the organization healthy. That logic held up reasonably well for over a decade of industry benchmarking, because the thing producing the diff, in every case the framework measured, was a human developer working at human speed, with some accumulated understanding of the system they were editing. That assumption, never stated outright because it never needed to be, is the part of the framework that agent-driven development breaks first.

What changes structurally when agents generate the code

An AI agent is not a faster version of a developer; it is a structurally different kind of contributor, and the difference breaks the assumptions DORA's metrics rest on. A developer accumulates context over months of working inside a codebase. An agent, by contrast, is a composite system of a model, a set of tools, some working memory, and a reasoning loop that can plan a change, execute it, and open a pull request with no human touching the keyboard at any step, but it carries none of that accumulated understanding of why the system was built the way it was.

The scale of what one agent can produce in a single pass makes the point concrete. A single prompt can touch dozens of files and generate thousands of lines of diff in roughly the time it takes a developer to write one function. GitHub's own research puts the current share of AI-generated code at close to half of all code written, a number that marks a real shift in where authorship sits inside the delivery pipeline. TELUS has reported saving hundreds of thousands of hours across its workforce by using agentic systems, and that kind of volume makes the old assumption, that someone reviews each pull request as an individual, self-contained unit of work, hard to sustain.

None of this removed the bottleneck in software delivery. It moved it. Code generation got dramatically faster, so the constraint that used to sit at the writing stage now sits downstream, in review and release. Making this worse is the fact that contributions now come in three distinct flavors, agentic, AI-assisted, and fully unassisted, and each behaves differently enough in a delivery pipeline that folding them into a single number erases the pattern a team actually needs to see. A dashboard that reports one blended "lead time" or one blended "deployment frequency" across all three types isn't describing any one of them accurately.

Why deployment frequency misleads under agent-driven delivery

Deployment frequency is the easiest of the four metrics to fool, because a rising number looks unambiguously like progress. After AI adoption, deployment frequency can climb sharply even as every other signal of delivery health, code quality, system reliability, customer-facing value, moves in the opposite direction. The mechanism is straightforward: agents produce pull requests faster and in greater volume, and that alone can push merge and deploy cadence up with no corresponding gain in quality, reliability, or value delivered to the customer.

The diagnostic pattern to watch for is deployment frequency and PR cycle time rising together. More deploys are going out, and each one is waiting longer in the queue for a human to actually look at it and approve it. A team that goes from five deploys a week to several multiples of that jumps, by DORA's classification, from medium performance to elite performance, with no change whatsoever in reliability or code quality.

What that rising line actually measures is agent throughput, the volume of output a generation system can produce, not delivery capability in any sense DORA originally intended. Reading deployment frequency in isolation, after agent adoption, answers a question nobody asked. The corrective is to put deployment frequency next to PR acceptance rate and PR cycle time in the same view. Adoption alone will always show a rising line. Only the combination shows whether delivery itself has improved or simply gotten louder.

Lead time: fast generation, slow review

Lead time for change used to describe one continuous process: a developer wrote code, then it moved through review, then it shipped. Under agent-driven delivery, lead time has split into two dynamics that pull in opposite directions, extremely fast generation and extremely slow human review, and reporting the blended total hides both halves of that story. An agent can produce a multi-thousand-line implementation in minutes. That same change can then sit untouched for hours before a reviewer even opens it.

The data on pull request size and review time sharpens the picture further. AI-assisted pull requests run the largest of the three contribution types at the 75th percentile, and yet they clear review faster than fully unassisted work. The largest changes in the pipeline are getting the least scrutiny per line. Agentic pull requests show a different pattern: the longest wait before anyone picks them up for review, though a comparatively quick pass once a reviewer actually starts.

A genuine counterargument deserves consideration here: teams that have learned to work with agents are capturing real throughput gains at the system level, not just at the line-of-code level. Faros AI's telemetry shows epics completed per developer rising meaningfully, evidence that the new bottleneck can be managed rather than merely endured. But capturing that gain requires measuring the review stage on purpose, rather than assuming that fast generation translates automatically into fast delivery. The fix is a split view: generation time tracked separately from time-in-review, with the clock starting when work is approved to begin rather than at first commit, and stopping only when the change is serving real production traffic. That split shows a team exactly where its actual constraint lives.

Why change failure rate masks architectural drift

Change failure rate answers one narrow question: did this specific deploy cause a detectable failure? That question misses the kind of damage AI-generated code tends to cause, because architectural harm accumulates gradually across many deploys, and no single one of them crosses the threshold that would register as a failure.

The root of the problem is what AI code generators are optimized to do. They're built to complete the immediate task in front of them, not to preserve the architectural coherence of the system they're editing. The resulting code is often syntactically correct and functionally passes its tests while being structurally wrong in ways no test catches: a function that reaches across a module boundary it has no business touching, a dependency cycle introduced three directories away from where the change was requested, a database query dropped directly into a controller. None of that fails a deploy. All of it degrades the system's structure, one agent-generated change at a time.

The evidence for this kind of silent decay is substantial. A study found static analysis warnings rose substantially post-AI-adoption across hundreds of GitHub projects, with code complexity increasing sharply, and neither signal appears anywhere in a change failure rate calculation. GitClear's longitudinal analysis, covering 2020 through 2024, found that AI-assisted development coincided with code refactoring falling from roughly a quarter of all changes down to a small fraction of that, a fourfold increase in code duplication, and a doubling of code churn, the kind of slow structural rot that a single failed deploy could never capture. Faros AI's telemetry puts a sharper point on the consequence: bugs per developer are up substantially and incidents per pull request have more than tripled, so the odds that any given merged change triggers a production incident have grown considerably, a finding from Faros AI's analysis of the 2025 DORA report, even while change failure rate as conventionally measured stays quiet until a major incident finally forces the trend into view.

DORA itself has responded to some of this. DORA added Rework Rate as a new metric and redefined MTTR as Failed Deployment Recovery Time, an acknowledgment that the framework needed updating. Even that expanded framework still has no visibility into whether a given change was AI-generated or which architectural constraints it may have violated along the way.

How MTTR hides the review burden

When a production incident traces back to code an agent wrote, mean time to restore stops being purely a deployment-and-rollback problem. It becomes a comprehension problem first, and MTTR was never built to measure how long it takes a person to understand code they didn't write and have never seen before.

GitLab's AI Accountability Report found that a large majority of teams agree AI has moved the bottleneck in their process from writing code to reviewing and validating it, and that shift gets sharper under incident pressure, when an engineer has to make sense of a large, unfamiliar diff fast, with production down and the clock running. Most review and incident tooling only looks at changed lines, so it misses contextual bugs: code that reads as syntactically correct in isolation but violates a convention established somewhere else in the codebase. Those are exactly the bugs most likely to survive all the way to production, and exactly the bugs hardest to fix quickly once they get there.

The weight of this falls disproportionately on senior engineers, who are the ones equipped to actually untangle an unfamiliar, structurally confused change under time pressure. DORA's metrics have no way to surface the review overload, the erosion of trust in agent output, or the burnout that can sit behind a perfectly acceptable-looking recovery time, even though the overload, the eroded trust, and the burnout are what actually produce that recovery time. MTTR reports time-to-green. It says nothing about the human cost of getting there. What MTTR needs alongside it is a measure of time spent understanding a failing change, not just time spent reverting or patching it, distinguishing recovery speed from recovery comprehension.

The measurement gap between perceived productivity and actual delivery

Most engineering leaders report real productivity gains from adopting AI tools, but those reports rest on adoption signals, how many developers are using the tools, not on delivery data that would show whether output is actually reaching customers faster or more reliably. Those are two different instruments, pointed at two different things, and leadership is often reading only one of them.

LinearB's 2026 data puts a number on the gap: 76.1% of leaders report productivity gains based on adoption signals rather than delivery data, while most organizations still don't formally measure what AI is actually doing to their delivery pipeline. Perception and reality are running on separate tracks. The gap widens further at the tool level. Acceptance rate for agent-created work varies by which tool produced it, with Devin improving from April 2026 and Copilot declining from May 2026 in the same window. A single blended adoption metric across all AI tools flattens out differences that matter a great deal in practice. Sonar's 2025 survey found developers estimate that a large share of their committed code is now AI-assisted, and standard DORA tracking has no field for recording which changes carry that origin at all. The industry is starting to move toward code-origin visibility, tracking which work was AI-assisted and whether it later turned into churn or a quality problem, precisely because the old instruments can't answer that question.

None of this means leaders are being negligent. It means the measurement tools built for a human-paced delivery process are still the ones being read for signal in an agent-paced one, and the two don't translate cleanly.

What to track alongside DORA

Diagram: One Metric, Four Blind Spots: What DORA Misses Under Agent-Driven Delivery. Visualizes: Show four DORA metrics — Deployment Frequency, Lead Time for Change, Change Failure Rate, and MTTR — each paired with the specific distortion AI agents…

The four DORA metrics are not wrong, but they are incomplete, and each one needs a companion measurement that recovers what agent-driven delivery pushed out of frame.

Deployment frequency needs acceptance rate broken out by contribution type, agentic, AI-assisted, and unassisted PRs tracked and reported separately, so that a rising volume of output and the actual rate of delivery show up together in the same view instead of one standing in for the other. LinearB's framework already treats these as three distinct categories with genuinely different pipeline behavior, and blending them into one number produces an average that describes none of the three accurately.

Lead time needs to be reported as two separate numbers, generation time and time-in-review, with the clock starting when work is approved rather than at first commit, so a team can see precisely where its real constraint sits instead of inferring it from a blended total. The number that ultimately matters is the date a customer actually receives the change. A coding metric that looks excellent while that calendar date hasn't moved isn't telling anyone anything useful.

Change failure rate needs a companion built for drift rather than incidents: an AI rework rate measuring the share of AI-generated output that needs revision before or after merge, a static analysis trend tracked over time, and a code churn rate followed longitudinally rather than deploy by deploy. These are exactly the kind of slow-moving signals that GitClear's findings on falling refactoring rates and doubling churn make visible, signals a single-deploy failure metric is structurally incapable of catching. Mean time to restore needs its comprehension-time companion, tracked separately from reversion time, so that recovery speed and recovery understanding stop being reported as if they were the same thing.

None of these additions replace DORA. A human writing at human speed, with human knowledge of the system underneath the change, is the context the original metrics assumed but no longer have. That assumption is gone from a large share of the code now moving through production pipelines, and the measurement has to catch up to that fact before the numbers on the dashboard can be trusted again.

Sources

  1. AI in software development: what the 2026 data shows
  2. How AI Agents Are Changing Software Development in 2026 - DEV Community
  3. 10 AI software development metrics CTOs should track in 2026
Filed underDeveloper Roles