Spec-Driven Development With AI Coding Agents
Specs keep AI agents aligned when code changes faster than humans can review it.

A single agent session can now touch dozens of files in one pass, produce a diff too large for any reviewer to hold in working memory, and leave the design that existed before that session only partially true afterward. This change in the unit of work is what spec-driven development has to answer for, and it did not exist in the same form when humans wrote most of the code by hand. For decades, software teams operated on an implicit contract: a developer who touched the authentication module understood the authentication module, carried its invariants in their head, and could be asked to explain a decision months later. Agentic software engineering severs that contract because the unit of work has changed shape entirely, not because agents are careless. The unit of interaction is no longer an isolated prompt answered by a human who reviews the output line by line. It is an autonomous session across an entire repository, in which the agent plans a sequence of changes, edits files, runs commands, reads the results, and iterates until it judges the task complete. Researchers at the Universidad Politécnica de Madrid, in a 2026 analysis of spec-driven development for agentic software engineering, draw a firm distinction: vibe coding and full Agentic Software Engineering are not points on a smooth continuum of the same activity. Each imposes distinct demands on how teams are structured, what artifacts they maintain, and what governance they apply, and a team that has become fluent in chat-assisted coding does not automatically become competent in agentic engineering simply because the tools look similar on the surface. The skills, the review habits, and the accountability structures that worked for one do not transfer cleanly to the other.
The Spec as Source of Truth
Spec-driven development makes the specification, not the code, the thing that governs the system. When product intent changes, the specification changes first, and the code is regenerated or revised to match it, rather than engineers patching the code directly and leaving the original design document to rot in a wiki somewhere. The code still has to run, still has to be maintained, and still matters enormously in practice, but it no longer sits at the center of the design process. The specification does that job instead, and the code is what results when an agent correctly interprets it. That reordering has a concrete shape in practice. The canonical SDD pipeline runs through a Constitution phase, where project-wide rules the agent must always obey get fixed (language choices, frameworks, testing standards, dependency constraints), followed by a Specify phase that captures user stories and acceptance criteria without touching technology choices, a Clarify phase in which the agent is made to surface ambiguities before any planning begins, a Plan phase covering architecture and data models, a Tasks phase that breaks the work into atomic, independently shippable items, an Implement phase, and an Analyze phase that closes the loop. This sequence exists to solve what the research literature calls prompt-architecture coupling: agents make architectural decisions constantly, at speed, and usually without anyone watching them do it. A specification forces those decisions into the open, so a human has to make them explicitly before the agent ever touches a file. The Madrid researchers frame specifications as the contract substrate between humans and agents, the artifact that reconstitutes, in specification-centric form, the properties that vibe coding quietly dissolved: accountability, verifiability, and the ability to hand work off between people without losing what it meant. A taxonomy published by the Federal Institute of Goias in June 2026, comparing six different development frameworks across specification, context, roles, execution, validation, and portability, reaches the same conclusion from an entirely different angle: once a framework adopts any real process at all, the isolated prompt gives way as the unit of work, replaced by persistent artifacts, explicit contracts, traceability records, and human review, the mechanisms that actually keep ambiguity down.
Multi-agent systems and the case for a shared spec
Run several agents in parallel on the same codebase without a shared specification, and each one optimizes its own slice of the work in isolation, with no way to detect that its decisions conflict with another agent's decisions three files away. One documented case from the field makes the mechanism concrete. A platform attempted a multi-agent setup with equal-status agents coordinating through shared locking, and the arrangement collapsed under its own design: agents held locks far longer than necessary, so adding more agents to the pool did not increase throughput, it dragged the whole system down to the speed of only a few. An optimistic-concurrency alternative produced the opposite failure: agents became risk-averse, steering away from the hard tasks that might trigger a conflict. The architecture that worked separates roles by function, not by status. Planner agents explore the codebase and generate tasks. Worker agents execute the tasks they are assigned without coordinating with one another, pushing changes only once the work is done. Judge agents evaluate, at the end of each cycle, whether the system should continue. A shared specification makes that separation possible by giving every agent a common reference point. Workers have no common contract to build toward without one, and a Judge has no ground truth against which to measure whether a cycle succeeded. GitHub's Agent HQ, announced at GitHub Universe in October 2025 and brought to public preview shortly after, lets teams run Claude, Codex, and Copilot on the same task simultaneously, with dedicated agents assigned to code review, test generation, security scanning, and deployment, each one specialized for its function. Coordination across that many specialized agents is coordination in name only if there is no shared document defining what correct behavior means for all of them at once. AWS's Kiro takes the same structural position from the tooling side: its design philosophy, per Grabowski, pushes deliberately against vague prompting, insisting on specs, tasks, designs, and hooks before any code gets generated.
The drift problem: why a spec written before coding is not enough
A specification that governs an agent's behavior at the start of a project and is never checked against the code again stops functioning as a contract within a matter of weeks. Grabowski's 2026 account of this failure mode is precise about the mechanism: an agent changes the code, nobody updates the specification to match, the existing tests still pass, and nothing in the pipeline ever flags the divergence. Over time the specification quietly turns into a historical record of what the system used to be, and every future agent run takes its instructions from a document that is, by then, simply inaccurate. The danger of this kind of drift is that it is invisible by construction. Tests are green. CI reports success. Nothing in the ordinary feedback loop of software development flags the gap, so nobody has a reason to go looking for it. The failures that result are architectural, so they do not announce themselves through a crash or a failed test run. One agent, tasked with adding a repository implementation, introduces a direct import from the domain layer straight into an infrastructure adapter, a textbook violation of hexagonal architecture, but it compiles cleanly and breaks nothing at runtime. Another agent, working with a context window too narrow to see the project's structural conventions, mistakes a carefully designed module hierarchy for unnecessary complexity and flattens it into a single directory. Utility functions start to quietly reference each other back and forth, and circular dependencies appear between modules. The Goias taxonomy identifies this as a structural gap across the entire field: no framework it reviewed scores strongly across all six of its dimensions, and drift between specification and code appears as a recurring risk in every one of them, paired with a companion failure of excessive trust in whatever the agent produces. A spec written before coding begins answers a question about how to start a project well. It has nothing to say about how a team discovers, six weeks later, that the system it is running no longer matches the document that was supposed to describe it.
What enforcement looks like when the spec is machine-readable and lives in the repo
An enforceable spec has to be encoded as a set of machine-readable contracts that live inside the repository next to the code they govern, not as a document in a separate tool that a human is supposed to remember to open and update. Grabowski's 2026 proposal for a Spec Growth Engine is the most developed version of this idea currently on offer, and it is built from three distinct parts. The first is a machine-readable spec graph that keeps contracts and design decisions in explicitly separate structures, rather than blending them into one undifferentiated document. The second is a Spine context assembler, whose job is to scope what any given agent sees. An agent assigned to work on a payment module does not get the entire repository tree. It receives the ownership path running from the root of the project down to the payment module specifically, plus the contracts belonging to whatever that module declares as its dependencies, a fraction of the total codebase. Without that kind of scoping, an agent working on payment code reads the whole repository by default, and a wider context window does not make an agent more disciplined, it makes the architecture harder for the agent to see clearly. Less context, chosen well, outperforms more context chosen carelessly. The third part is the drift gate, where enforcement becomes enforcement rather than aspiration, turning any divergence between the spec and the code into a blocking condition on the merge itself. Underneath the Spine and the drift gate sit three concrete, machine-readable layers of contract. Structural contracts govern which direction code is allowed to depend in, and tools such as Dependency Cruiser or ESLint's module-boundary rules check them. Type-shape contracts constrain what an API surface is allowed to look like. Behavioral contracts are enforced through the test suite itself. The Madrid researchers describe the same architecture in socio-technical terms, distinguishing the technical harness around the agent, meaning what it is permitted to know and permitted to do, from the methodological harness around the human team, meaning how objectives get defined and how outcomes get verified once the work is done. Grabowski's framework sits deliberately between two extremes: every node in the system has a spec attached to it, code and spec evolve together inside the same commit rather than in separate update cycles, and divergence between them is a blocking merge error, all without carrying the institutional weight of a heavyweight methodology like RUP or model-driven architecture.
CI as the enforcement layer that closes the loop between spec and shipped code
A compliance check that exits non-zero and stops a merge from landing does more for architectural integrity than any volume of after-the-fact design review, because it makes the spec's constraints binding at the exact moment code is trying to enter the system. The practical motivation for this is the review bottleneck agentic coding has created. AI-assisted pull requests run much larger than the ones engineers used to write by hand, per a benchmark from LinearB, and a bigger diff is simply harder to verify by eye, so it either sits in the review queue longer or gets merged without anyone looking at it closely. Faros AI's 2026 report found median PR review time climbing sharply alongside this shift, with a growing share of pull requests merging with no human review whatsoever. A human reviewer looking at a sprawling diff that touches dozens of files cannot reliably spot a single architectural boundary violation buried inside an infrastructure adapter. The scale of the diff exceeds what sustained human attention can track across a review interface, no matter how experienced the reviewer is. The drift gate inside Grabowski's Spec Growth Engine turns this from a problem of human attention into a problem of machine enforcement: the CI run checks for divergence between spec and code, exits non-zero the moment it finds a violation, and blocks the merge, so an agent's output has to satisfy the architectural contract before it is allowed to land in the main branch. The Madrid researchers describe this as the shift from informal to formal verifiability. Spec-driven development reconstitutes accountability and verifiability in a form centered on the specification itself, and CI enforcement is the mechanism that makes that verifiability operate in practice rather than remaining a principle teams only gesture toward.
The honest case against SDD
The strongest objection to the productivity case for spec-driven development is not theoretical. METR ran a randomized controlled trial in July 2025 involving sixteen experienced open-source developers completing a large batch of real issues, each expected to take roughly two hours, inside mature repositories exceeding a million lines of code. When developers were allowed to use AI tools, their tasks took longer to finish, not shorter. The gap between what they expected and what actually happened was stark: these developers expected a speedup going in, and after finishing the work, they still believed they had been faster, even though the measured data said otherwise. That perception gap bears directly on whether teams will adopt SDD's discipline. A team convinced that its agents are already saving time has little appetite for adding specification overhead and enforcement machinery on top of a workflow its members believe is already running efficiently. The counter to METR's finding is that its trial conditions may show agent use poorly constrained by any specification, not a hard ceiling on what agentic development can achieve. The study measured a prompt-and-iterate workflow, not an agent operating inside the kind of architectural contract that spec-driven development is built to impose. METR's own later data, from January 2026, complicates the picture further. Estimated time savings for METR staff using the coding agent Claude Code reached as high as 13x for some individuals, according to METR's own analysis, but METR itself calls that figure a soft upper bound, and it believes the true productivity multiplier across its staff is substantially lower, since self-reported daily estimates consistently undershot what was actually measured. Read together, the two studies point toward the same underlying conclusion: productivity gains from agentic coding are sensitive to how the workflow around the agent is structured, more than to which model or tool is doing the generating. That is the condition spec-driven development, backed by enforceable contracts in CI, is built to address directly.
Sources
- From Prompt to Process: a Process Taxonomy and Comparative Assessment of Frameworks Supporting AI Software Development Agents
- Spec-Driven Development for Agentic Software Engineering: Harnessing Human-Agent Teamwork
- The Spec Growth Engine: Spec-Anchored, Code-Coupled, Drift-Enforced Architecture for AI-Assisted Software Development


