Agent Architecture Weekly

Multi-Agent Orchestration for Concurrent File Changes

Concurrent agent conflicts demand explicit coordination, not better merges.

Features Editor · · 10 min read
Cover illustration for “Multi-Agent Orchestration for Concurrent File Changes”
Agent Workflow Patterns · October 2, 2026 · 10 min read · 2,211 words

A multi-agent run finishes. Every file compiles. The test suite goes green. Nothing in the pull request looks wrong, and the merge goes through without a single conflict marker. That is the failure this piece is about, and it has nothing to do with whether the agents ran in parallel successfully. The dominant assumption in multi-agent coding treats the merge as the hard part, when the hard part is making sure the agents never needed a merge to resolve a disagreement they didn't know they were having. Existing multi-agent coding systems inherit a serial limit even when they look parallel on the surface: they sequence agents through phase handoffs, or they pool independent samples without any real coordination between them, and a single agent working alone abandons half of hard tasks by writing a one-file stub and exiting. The natural engineering response to that limitation is to fan agents out across files and merge their output at the end, and this supervisor/fan-out model dominates in practice because it is straightforward to build. Fan-out earns its place when tasks split cleanly into independent sub-tasks that never need to see each other's intermediate results, and parallel code review of separate files is one of the use cases where that split holds. Most real coding tasks do not split that cleanly. Modules share interfaces, call into each other, and carry structural assumptions about what the other side of a boundary will do. The codebase compiles, passes tests, and has quietly lost its architectural coherence.

How widespread concurrent agent conflict is across real repositories

Concurrent agent conflict is the default condition of any codebase where more than one agent is active at once, and the window during which agents overlap on the same code is wide enough to expose most agent-authored pull requests to it. A 2026 analysis of agent-authored pull requests across thousands of repositories found that the repositories with overlapping agent PRs account for a large majority of all agent-generated PR activity, and widening the overlap window to a single week captures nearly all of it. The AgentRoom paper's survey of related work also points to the AgenticFlict dataset, presented at AIware 2026 in Montréal, which extracted more than 336,000 fine-grained conflict regions from agent PRs across tens of thousands of repositories, establishing a conflict rate that is structural rather than incidental.

The mechanism is mechanical: it follows directly from how the workers read, write, and recombine data. Parallel execution stays safe when workers read from the same snapshot without modifying it, write into separate namespaces that never intersect, or produce results a deterministic reducer can recombine without judgment calls. It becomes unsafe the moment agents write to the same interfaces, make overlapping architectural decisions independently, or produce changes that are each individually valid but jointly inconsistent once placed side by side. None of that can be fixed by writing a better prompt. Concurrency control is a property of the coordination layer the agents run inside, not an instruction you can give an agent and trust it to hold in mind across twenty minutes of independent reasoning.

Locking and optimistic concurrency as failed coordination primitives for agents

The two concurrency control primitives any developer reaches for first, file locking and optimistic concurrency control, both fail in agent contexts, and the reasons are specific to how agents behave rather than to any flaw in the implementation. A documented internal architecture experiment makes the case cleanly. Under equal-status agents with locking, agents held their locks too long, and scaling the system from a handful of agents to twenty produced throughput no better than what two or three agents working without contention would deliver, because the coordination overhead ate the entire parallelism gain. Under optimistic concurrency control, agents turned risk-averse: lacking any hierarchy to defer to, they avoided difficult tasks altogether and confined themselves to small, safe changes that wouldn't trigger a conflict. The resolution the experiment reached was role differentiation, specifically a Planner/Worker separation, rather than isolation by itself.

Locking fails for agents because agents are not threads. A thread holds a lock briefly while performing a deterministic write and releases it an instant later. An agent holds a resource while reasoning, and reasoning is slow, context-dependent, and variable in duration in a way no scheduler can predict in advance. Lock granularity tuned for human developers, who commit in minutes and release in seconds, is far too coarse once the thing holding the lock might sit on it for the better part of an agent's working session.

Optimistic concurrency fails for the opposite reason. Agents facing conflict risk do not push through it the way a confident engineer might; they shrink their scope instead. An agent that avoids touching a contested module is quietly producing an incomplete result that will look finished and be wrong in ways nobody flagged, which is worse than a result that fails loudly. A 2026 paper on reinforcement learning for LLM-based multi-agent systems frames the deeper issue at the system level: orchestration rewards have to target system-level properties like split correctness and aggregation quality, not just whether an individual agent completed its assigned task, because naive single-agent credit assignment cannot capture the coordination decisions that decide whether the combined output holds together.

Explicit Coordination: Claims, Broadcast, and Shared Workspace

Once locking and optimistic concurrency are both ruled out, what remains is explicit, shared runtime state that lets agents see what their peers are doing before they act, not after they've already written something that conflicts. AgentRoom, published by Seonglae Cho and Donghyun Lee of Holistic AI and UC Berkeley at ICML 2026, is the clearest published architecture built on that premise: a CRDT-backed shared workspace with file-level claim semantics, an append-only broadcast log, and live agent status, all exposed as MCP tools named room_claim, room_release, room_state, room_broadcast, and room_read. room_claim(path) atomically assigns ownership of a file and rejects the call if another agent already holds that path, which closes the exact silent-overwrite failure that locking was supposed to prevent and couldn't. The broadcast log gives every agent in the room visibility into what every other agent is doing in real time, with no central coordinator needed to mediate each individual decision. The CRDT substrate merges concurrent writes at the character level as a last-resort backstop; the architecture's actual value is in stopping conflicts from reaching the merge layer at all.

The evaluation behind AgentRoom makes the stronger claim explicit: coordination is what carries the quality improvement, not parallelism and not CRDT-merge correctness. An ablation that keeps the shared substrate but strips out the MCP coordination layer falls below the performance of the full system, and the full system, claim semantics, broadcast, and all, outperforms both parallel-merge setups and sequential ChatDev-style handoffs. The effect appears even with the smallest possible room. Two agents working inside an AgentRoom suppress the lone-agent stub-and-exit failure almost entirely: a solo agent is 13.7 times more likely to abandon a hard task than an agent working inside the full coordinated system. The presence of a visible peer changes what the agent itself judges to be tractable. Coordination primitives, in other words, need to live at the file-claim and broadcast level as first-class runtime objects any agent can query at any moment, not as conventions an orchestrator enforces only after the fact.

Role Differentiation and Coordination Primitives

Coordination primitives by themselves are not sufficient. Who an agent is has to determine what it is permitted to claim, what it is obligated to broadcast, and what it must check before it acts; treating every agent as an interchangeable member of one worker pool makes coordination overhead unsustainable at scale. The Planner/Worker separation that resolved the locking and OCC failures described earlier was not a narrow implementation choice. It reflects a structural principle that not all agents should carry equal write authority over all parts of a codebase at all times. One architecture extends that principle into a full division of labor: a Coordinator agent analyzes the codebase, drafts a living specification, generates discrete tasks, and delegates them to six specialist agents, named Investigate, Implement, Verify, Critique, Debug, and Code Review, each working inside its own isolated git worktree within a task-scoped workspace, with the living spec coordinating tasks and validation across the parallel execution.

The orchestration pattern guide behind this draws the line between a supervisor and a swarm on exactly this axis. A supervisor delegates non-overlapping tasks to specialist sub-agents and synthesizes their independent results, while a swarm, the kind that can scale to 300 sub-agents across thousands of coordinated steps, requires a different coordination topology entirely because its agents are peers rather than specialists. Coordination primitives have to be designed for the role structure actually in use. In a supervisor/worker system, the supervisor's decisions function as structural constraints that every worker must read before it claims a file. In a peer swarm, broadcast and claim have to stay symmetric across all agents, but authority over cross-cutting concerns still has to be delegated explicitly to someone or something. Serial execution remains the right call for writes to the same record, for changes that touch overlapping code areas, for decisions that depend on a prior judgment call, and for irreversible external actions, and the role structure of the system is what determines which agent is positioned to even recognize that one of those conditions applies. Adding more agents to a pool without deciding who answers to whom is not a coordination strategy.

Local Structural Decisions and Global Architectural Damage

The consequences of skipping that role design go well beyond conflicting files. The worst outcome of concurrent agent work is a codebase in which every agent made a decision that looked locally reasonable, and those decisions collectively violated the architectural model the system was built on, with no error raised anywhere along the way. AI code generators optimize for finishing the task in front of them, not for preserving a system's long-term architectural coherence, and none of them can see a codebase's structure the way a developer with years inside it can. The failure modes this produces are specific and recognizable: an agent will query the database directly from a controller, reach into the billing module straight from the user module, or introduce a dependency cycle between auth and notification services, having seen similar-looking patterns somewhere in training data without ever registering that those patterns were wrong in this particular context. Every one of these decisions is individually plausible. Each passes type checks. Each passes the test suite. The violation lives at the structural level, not the syntactic one, so none of the usual automated gates catch it.

Context-window limits compound the problem as codebases scale. Once a system grows past what an agent can hold in its effective context window, the agent loses a coherent grip on system-wide invariants and dependencies, and a small error introduced in an early commit cascades into compounding failures in whatever work follows it, with no robust mechanism in place to detect or recover from the chain. Agents tend to follow whatever pattern already exists in a codebase, but following a pattern is not the same as understanding why a team chose it, what past shortcut caused the production incident that pattern was built to prevent, or what users actually do with the system as opposed to what the written requirements claim. Teams need some mechanism that surfaces a new dependency or a deviation from established approach for human review before it ships, rather than after. Steve Yegge's Beads system addresses one real dimension of this problem, the "50 First Dates" condition in which agents carry no memory between sessions and leave behind a swamp of conflicting markdown files, by storing issues as JSONL in git with hash-based IDs designed specifically to avoid merge conflicts across multi-agent workflows. Memory persistence is a floor, not a ceiling. An agent that remembers a past decision is not the same as an agent that respects an architectural constraint it was never forced to check.

Making architectural constraints machine-readable and enforceable before the merge

The only coordination primitive that reliably stops architectural drift from concurrent agents is a machine-readable constraint that fails a build before the violation ever reaches the main branch, because human review is too slow and too inattentive to catch it at the volume agents now generate. Mature projects have already started encoding architectural decisions as checkable documents instead of advisory prose. Kubernetes and Next.js both use mechanical, checkable AGENTS.md constraints, generated artifacts marked read-only, a declared single source of truth, explicit mode matrices, that both agents and CI pipelines can parse directly. An open issue on the Looming project explicitly connects AGENTS.md architecture constraints to future CI enforcement through depguard rules, which shows that the practitioner community has already identified this as the right next step rather than a speculative one. Architectural fitness functions paired with CI pipelines can catch Architecture Decision Record violations before a merge happens at all: an ADR Conformance Check runs the moment a pull request opens, blocks the merge, reports the specific violation, and separately flags any new architectural decision made without a corresponding ADR on file. The check exits non-zero, and that exit code is the primitive that actually makes the constraint enforceable rather than aspirational.

Sources

  1. AgentRoom: Concurrent Multi-Agent Coding in a CRDT-Backed Shared Workspace
  2. Reinforcement Learning for LLM-based Multi-Agent Systems through Orchestration Traces
  3. Multi-Agent Orchestration: 5 Patterns That Work in 2026
  4. AI Coding Agents in 2026: Coherence Through Orchestration, Not Autonomy

More in Agent Workflow Patterns