Agent Task Decomposition at Module Boundaries
Agents need architectural boundaries defined before they write code, not discovered after.

Consider what happens when an agent is asked to add a utility function to VS Code's codebase. ContextCov's work on that production codebase documents the result in specific terms: the agent places the utility in src/vs/workbench/common/, when the project's layered architecture actually requires it to sit in src/vs/base/. The code compiles. It passes every style check. And it is wrong, in a way that neither the compiler nor the linter has any way to catch, because the rule being broken lives nowhere a machine can read it. This is the shape of the problem across agentic coding broadly in 2026. Agents now plan, execute, and iterate across multi-step workflows, touching dozens of files inside a single task, and that reach turns a boundary mistake into something serious. The agent optimized for the task directly in front of it: a function needs to exist, and the nearest caller tells it where. Nothing in that reasoning touches the layered structure the project depends on, because the agent's task representation does not include that structure. An agent cannot violate a boundary it has no way to see, and by default, it has no way to see one.
The Diff as the Worst Place to Discover a Boundary Violation
Once the mistake is made, the diff is where someone is supposed to catch it, and the diff is poorly suited to that job. Research comparing AI-assisted pull requests to human-authored ones finds them substantially larger, and defect detection is known to degrade sharply once a change crosses a well-established line-count threshold, a threshold that agent-generated PRs cross routinely. Size alone explains part of the problem. A diff shows a reviewer individual file changes, line by line, with no rendering of the module graph those changes reshape. A utility dropped into the wrong layer reads as locally correct, even to a careful reviewer working line by line, because the diff format has no way to surface cross-module topology. A 470-PR analysis bears this out directly: AI-co-authored PRs show significantly more logic and correctness issues, along with elevated security and error-handling flaws, the kind that compile cleanly and clear baseline tests without tripping anything. Reviewing a model's own output doesn't fix this either. A model asked to check its own work carries forward the same assumptions that produced the violation in the first place, so the second pass repeats the first judgment in different words. If the violation cannot be seen at the point of generation and cannot be reliably caught in the diff, the constraint has to move to before the code gets written.
Architecture No One Designed: Decomposing Tasks Without Boundaries
Pushed out past a single task, this pattern produces a codebase whose shape nobody actually chose. Architecture drift doesn't arrive as one obviously bad commit. It accumulates, one locally rational decision after another, each respecting whatever constraint was immediately in front of it and none respecting the constraints that hold the system together. Nothing in the system registers a crossed boundary as a problem, so each silent crossing makes the next one a little easier to justify. Asthana et al. (ACM CAIS 2026, the RSTD paper) trace this back to a structural habit: encoding all task logic, control flow, and output generation inside a single monolithic prompt. That habit produces brittle behavior and poor debuggability on its own terms, but it carries a second cost that matters more for architecture: the task's actual structure never gets written down anywhere outside the prompt text, so there is nothing external for the system or the engineer to inspect or enforce. When several agents run in parallel against a codebase with no shared architectural ground truth, they start contradicting each other, re-deciding questions the team already settled weeks earlier, because each one works from its own local read of the code. Versioned architecture decision records, rules files, and project memory exist to close that gap: they give every agent the same starting picture of how the system is supposed to fit together. The scale involved is already large: frontier agentic systems resolve roughly four out of five real issues on a standard coding benchmark, and still manage around two in five on a harder, contamination-resistant version of that benchmark, with a non-trivial share of code at major AI labs already written by systems operating at this level. Architectural decisions are being made implicitly, at that volume, right now. If the agent has no way to see the boundary, the engineer's job is to make the boundary visible before the agent starts, not after it finishes.
The module boundary as the primary decomposition unit
Here is the claim this piece is built on: when an engineer hands a task to an agent, the boundary of the relevant module, what it owns, what it must not touch, what it's allowed to call, should define the scope of that task from the start. It should not be something discovered after the fact by squinting at a diff. So the engineer's work has to change before the agent ever runs. Instead of writing a prompt that describes the feature to build, the engineer defines the module's scope, the call relationships it's allowed to form, and the architectural invariants it has to satisfy. The prompt becomes a secondary artifact sitting on top of a structural specification, not the other way around.
ContextCov's architectural validator shows what this looks like in practice. It builds a dependency graph of the codebase and enforces the layered structure directly: if an agent tries to create a file in a location the architecture forbids, it gets immediate feedback naming the specific rule it broke and the correct placement for the file, before that file becomes part of a pull request anyone has to review. The boundary is readable by the machine before a single line gets written, not reconstructed by a human after the fact. OpenAI's Codex moves the same way with AGENTS.md, a persistent instruction file that gives an agent project-level guidance and conventions before execution begins. Windsurf's "codemaps" do it differently: AI-generated hierarchical visual maps of large codebases, built to help developers and agents alike understand how dependencies and architecture connect across a large repository. Different mechanisms, same principle: the architecture gets established before the prompt, not inferred from the prompt's output.
The obvious objection is that this sounds slower, that writing a structural specification before every agent task adds friction a team doesn't have time for. The specification is the task. An agent handed a vague, cross-module mandate will touch a wide spread of files and hand back a diff no one can meaningfully review. An agent handed a bounded module scope produces a diff someone can actually read and judge, because the scope of what it was allowed to do was fixed before it started.
Static Decomposition at Boundaries vs. Runtime Structure
Defining the module boundary as the unit of scope is necessary, but it isn't sufficient on its own, and the RSTD research is specific about why. Static decomposition, fixing the subtask graph at design time and never letting it change, does not reliably cut retry cost. In the Kubernetes root-cause-analysis workload Asthana et al. tested, static decomposition actually drove retry cost above the monolithic baseline, by a wide margin, because a fixed sequential pipeline has to rerun several downstream subtasks whenever any single step along the way fails. Runtime-structured decomposition behaves differently: it retries only the step that actually failed, and across both workloads tested, it produced the largest reduction in retry cost of the three configurations compared. The mechanism is straightforward. A static subtask graph fixes its shape in advance, but a runtime-structured one can branch based on what an intermediate step actually finds, so what the agent discovers partway through determines the shape of the task.
This matters directly for boundary-scoped decomposition. The module boundary should set the outer scope of a task, but the execution path inside that scope needs room to adapt. If an agent hits a schema validation failure at a module interface, it should retry that interface call specifically, not restart the entire module-level task from the beginning. RSTD's broader architectural pattern, separating orchestration, state management, and LLM inference into distinct layers, with decomposition decisions pushed into executable control flow rather than left sitting in prompt text, is the natural complement to boundary-scoped planning: one layer defines where the agent is allowed to work, the other governs how it recovers when something inside that scope goes wrong.
Making boundaries machine-enforceable rather than documentation the agent ignores
Once the boundary is defined, the question becomes how to make sure an agent can't quietly step over it. The field splits into two camps on where that constraint should actually live. One camp puts architectural rules into pre-generation context: long AGENTS.md files, ARCHITECTURE.md documents, rules files an agent reads before it starts working. One side holds that piling natural-language conventions into context overloads the agent and raises its failure rate, and that the same rule, built into a tool that runs before any code gets written, gets enforced without consuming a token of context.
One version of that tool-level approach encodes substrate conventions directly: strict linters, type checkers, and runtime contract validators that enforce, for example, a clean split between a pure core/ and an impure runners/ layer. Neither the agent nor the human reviewer has to remember a rule the linter has already been taught to check. Practitioner guidance on wiring architecture tests into CI/CD argues for full adherence on a simple basis: a violation physically blocks the build from passing, so the check functions as a binary gate.
The honest synthesis is that both layers earn their place. Natural-language context, the rules files and architecture decision records, gives the agent the ground truth it needs to plan correctly. Tool-level enforcement exists for when that planning fails anyway, and it stops a violation from merging. Running a tool like gr check inside CI is the practical expression of this model: the graph of module relationships lives in the repository, versions alongside the code it describes, and exits with a non-zero status the moment a proposed change crosses a boundary it shouldn't. The merge gate carries the enforcement, and a comment left on a pull request does not.
Where the engineer's judgment now compounds most
Anthropic's 2026 Agentic Coding Trends Report states the central prediction: the value of an engineer's contribution moves toward system architecture design, agent coordination, quality evaluation, and strategic problem decomposition, with engineers orchestrating agents that write the code. The practical consequence is that human judgment needs to arrive earlier in the process, not later. Correcting course during the planning phase, while module boundaries and scope are still being defined, costs very little. Correcting course after an agent has already produced a large diff is expensive and often structurally difficult, because the mistake is now woven through dozens of files.
That judgment now compounds at two points: the plan boundary, where the module's scope and its architectural invariants get specified before any agent touches the code, and the merge boundary, where a CI gate enforces what was specified. So everything that happens between those two points can move at the speed the agent sets.
A regulatory layer has arrived to reinforce this. The EU AI Act's general application and transparency requirements under Article 50 take effect August 2, 2026, and while the high-risk system obligations under Articles 8 through 15 were pushed to December 2, 2027, a pipeline built around autonomous agents with deployment permissions that touch regulated systems may well qualify as high-risk under those same articles. So architectural traceability is no longer just a matter of engineering taste; a compliance team can ask to see it.
The deeper reason to treat this as infrastructure rather than a workflow preference is a matter of pace. Researchers estimate that the length of software benchmark tasks AI agents can complete doubles roughly every seven months. Boundary enforcement works today because an engineer reads every diff carefully, but it will not hold at next year's level of agent capability, when the diffs are larger and the tasks span more of the codebase at once. Go back to that VS Code example: an agent placing a utility in src/vs/workbench/common/ instead of src/vs/base/ is a small, recoverable mistake today, catchable because someone happened to notice. At twice the task length and twice the file count, a human reviewer can no longer catch the same kind of mistake just by reading carefully. Only a machine-readable boundary, checked automatically before the merge, can catch it.


