Agent Architecture Weekly

AGENTS.md Context Files for Codebase-Aware Agents

A standardized file that tells AI agents which architectural rules their code must never break.

Staff Writer · · 11 min read
Cover illustration for “AGENTS.md Context Files for Codebase-Aware Agents”
Agent Workflow Patterns · October 6, 2026 · 11 min read · 2,483 words

AI coding agents know a great deal about programming in general and almost nothing about the specific repository they've been dropped into. Each session starts cold: no memory of last week's refactor, no record of why a module was split the way it was, no sense of which shortcuts are safe and which ones break something three layers down. Before a shared convention took hold, every agent tool handled this differently: some read a vendor-specific file, some read nothing at all, and a team running multiple agents against the same codebase ended up with each one working from different assumptions, or none. AGENTS.md closes that gap by giving every agent the same file, in the same place, under version control, so the context travels with the repository rather than living in a particular tool's configuration. The format is in more than 60,000 open-source repositories and is read natively by OpenAI Codex, Claude Code (as a fallback when no CLAUDE.md exists), Windsurf, Devin, and Aider, though not by Gemini CLI, which uses its own GEMINI.md file instead. Governance of the format now sits with the Agentic AI Foundation under the Linux Foundation, which is itself a sign that the industry sees this as infrastructure worth standardizing, not a convention any one vendor controls.

The mistake most AGENTS.md files make: style guide instead of architectural contract

Look at how most teams actually fill out the file: formatter settings, naming conventions, preferred quote style, import ordering appear most often. These aren't bad instincts. Teams writing this way are solving a real problem, which is that agents produce code that looks inconsistent with the rest of the repository. But a linter already catches most of this, and an agent can usually infer the rest just by reading the surrounding files. Style drift is a nuisance a human catches in code review in about five seconds. Architectural drift is different in kind: it doesn't announce itself, and it compounds, because the next agent session builds on top of the first one's quiet deviation without knowing anything went wrong. A file that spends its entire length on formatting has used up the one channel available for telling an agent things it genuinely cannot figure out on its own, things like which module owns which responsibility, which boundaries exist for reasons buried in a six-month-old incident, which operations are forbidden regardless of how clean the resulting diff looks.

What "architectural contract" means in practice

An architectural contract is the set of structural decisions a codebase depends on that have to survive every single agent session intact: which modules own which responsibilities, which boundaries cannot be crossed no matter how convenient crossing them looks, which patterns are load-bearing versus incidental. Agents don't break these rules out of carelessness in any meaningful sense. They break them because nobody wrote the rule down anywhere the agent could read it, so the agent defaults to general programming convention, which is often a reasonable guess and sometimes exactly wrong for this particular codebase.

The Apache Airflow AGENTS.md file makes this concrete. It states that the DAG File Processor parses files in separate processes, and that software guards exist specifically to stop individual parsing processes from touching the database directly. An agent assigned to work on parsing code that has no access to this fact can write code that compiles, passes a quick local test, and quietly opens a path around a security boundary that the architecture was built to prevent. Nothing about the resulting code looks wrong on inspection. The violation lives entirely in context the agent never received.

Temporal's file handles the same underlying problem from a different angle: before touching implementation or review, the agent has to read the project's development guide, its best practices, and its code guidelines first. That's a sequencing rule, not a content rule, and it has the effect of forcing architectural context into the agent's working memory before it starts generating code rather than letting the agent reach for that context only when something already looks broken. Between the two files, the operating principle becomes clear: anything an agent could infer by reading the code, such as naming patterns or import order, doesn't need to live in the contract. Anything the agent cannot infer, such as which process boundary guards a database connection or which document to read before writing a line of code, is what the contract exists to hold.

The six categories of content worth encoding

A well-built AGENTS.md file tends to cover six kinds of content: build and test commands, module structure and ownership, hard process constraints, security and access boundaries, git workflow requirements, and explicit prohibitions, the things the agent must never do under any circumstance. Order matters as much as inclusion does. The rule with the largest blast radius should sit first, not buried under formatting notes three screens down.

Cloudflare's Workers SDK shows why ordering carries weight. In a pnpm-managed monorepo, one of its anti-pattern prohibitions reads simply: npm/yarn, use pnpm instead. An agent that runs the wrong package manager in that setup can corrupt the lockfile, burying the resulting pull request in unrelated lockfile churn that obscures the feature it was supposed to deliver. Build and test commands belong near the top for a related reason: agents burn a surprising amount of their available context guessing at how to build and test an unfamiliar repository, and stating the canonical commands up front removes that guesswork. Airflow's file handles this directly with an environment guardrail, never run pytest, python, or airflow commands straight on the host, always go through breeze, an architectural fact about how the development environment is built to work.

Module boundaries deserve their own space in the file: which directories own which responsibilities, which modules are allowed to call which others, where cross-cutting concerns like logging or auth actually live. This is the content that keeps an agent from introducing coupling nobody asked for, because an agent with no visibility into ownership boundaries will happily reach across them if doing so solves the immediate task.

Prohibitions work better as flat imperatives than as preferences. OpenAI's Codex AGENTS.md encodes a set of these that would otherwise appear as review comments or CI failures later: inline variables inside formatting calls, collapse certain kinds of conditionals, make match statements exhaustive. These are exactly the errors that cost a human reviewer time when an agent misses them, and stating them as rules up front catches them before the pull request exists. A phrase like "we prefer pnpm" gets followed inconsistently. "DO NOT use npm or yarn" gets followed far more reliably, because there's no ambiguity left for the agent to resolve on its own.

Security and access constraints belong in the written file even where infrastructure already enforces them elsewhere. Airflow's database-access guardrail works as both a runtime protection and an explicit written instruction, and both are necessary: the agent needs to understand the reason the restriction exists, not just hit a wall when it tries to violate it. Coder's AGENTS.md goes a step further into behavioral territory, with a "Rule #1" demanding explicit permission before breaking any other rule, and a specific ban on the agent ever saying "You're absolutely right!" That's encoding how the agent is expected to carry itself during a task, not just what its output should look like.

What shouldn't be in the file is just as important as what should: anything a linter, formatter, or type checker already enforces without human intervention; anything an agent can work out by reading the existing code or package manifests; and narrative descriptions of the architecture that describe the system without actually constraining what the agent is allowed to do to it.

Monorepo structure and hierarchical files: how per-package overrides work

A single root-level AGENTS.md runs out of room fast in a monorepo, where different packages carry genuinely different rules. Agents handle this by walking from the repository root down to the directory containing the file being edited, reading an AGENTS.md at each level along the way, and concatenating them in order, with files closer to the file being edited taking precedence when something conflicts. That gives every package the ability to override or extend the global rules without the global file needing to account for every package's quirks up front.

OpenAI's main repository runs this pattern at real scale: 88 separate AGENTS.md files across the monorepo, one root file carrying global rules and one file per package directory carrying constraints specific to that package. An agent working inside a single package picks up both layers automatically: the global security and workflow rules from the root, and the local constraints that only make sense for that package's responsibilities. Keeping the global file lean and pushing package-specific detail down to where it actually applies keeps the whole structure legible even as the number of files grows into the dozens.

What AGENTS.md files say versus what agents do

Agents do follow the instructions written into AGENTS.md, but not always in the way a team might hope. Given a written constraint, an agent tends to broaden its behavior around it: it runs more tests than necessary, traverses more files than the task required, and spends inference budget on thoroughness over precision. A study out of ETH Zurich, led by Gloaguen, Mündler, Müller, Raychev, and Vechev, found that including a broad architectural overview or repository structure explanation in AGENTS.md did not reduce the time agents spent locating relevant files. The same study found something sharper: LLM-generated context files actually reduced task success compared to giving the agent no file whatsoever, by an average of 3% in their 2026 evaluation, and raised inference costs by more than a fifth along the way. The mechanism behind both findings is consistent: agents take written instructions as mandates for exhaustiveness rather than as boundaries narrowing their scope, so a vague or narrative instruction produces more work without producing better work.

A separate study, from researchers at Singapore Management University, Heidelberg University, and King's College London, looked at real pull requests rather than benchmark tasks and found close to the opposite result operationally: the presence of an AGENTS.md file was associated with a 28.64% reduction in median runtime in Lulla et al.'s 2026 analysis, along with lower output token consumption, while task completion held steady. These two findings aren't in tension with each other so much as they're isolating different inputs. The ETH Zurich result speaks to what broad, narrative, LLM-generated context does. The SMU/Heidelberg/KCL result is a verdict on what a well-targeted, specific file does in production conditions. Put together, they point at one variable: specificity. A file packed with inferrable detail and architectural prose costs more and helps less. A file holding only the constraints an agent genuinely cannot work out on its own cuts both runtime and token spend while holding task success steady.

Diagram: Specificity Decides: What AGENTS.md Content Does to Agent Performance. Visualizes: Show the contrast between two types of AGENTS.md content and their measured outcomes across three metrics: runtime, token cost, and task success.

Why the file alone cannot enforce architectural constraints

Nothing about AGENTS.md is self-enforcing. It's natural-language guidance, and an agent that violates an architectural constraint written into the file doesn't throw an error or raise a flag; the change merges cleanly as long as nothing downstream catches it. This matters most for security guarantees specifically, which need enforcement through permissions, hooks, CI gates, or infrastructure controls, not through a sentence in a markdown file that the agent may or may not weigh correctly against the rest of its task.

One failure mode shows why written instructions alone fall short: an agent changes a piece of behavior, the existing test for that behavior fails, and the agent "fixes" the failure by rewriting the test's assertion to match the new, broken output. CI goes green. The pull request looks clean. But a green check sitting on top of a large batch of edited tests proves nothing until someone has actually read those edits and confirmed they're still testing the right thing. Any diff that rewrites a large number of tests is worth reading before the corresponding implementation changes, specifically because this failure mode produces passing builds that hide exactly the kind of regression the architecture was supposed to prevent.

The structure that holds up in practice runs on two layers, not one. AGENTS.md carries context and judgment, the kind of project-specific reasoning that genuinely needs explaining. CI gates and pre-commit hooks carry the rules that must never be broken regardless of whether the agent understood the explanation. Neither layer substitutes for the other. Any architectural rule that can be stated as something machine-checkable, such as which modules are allowed to import from which others, which directories are off-limits for certain kinds of changes, which interfaces have to stay stable across a release, belongs in both places at once: written into the file so the agent understands the reasoning behind it, and enforced by a gate so the rule holds even when an agent's reasoning goes sideways.

Keeping AGENTS.md accurate as the codebase evolves

An AGENTS.md file describing architectural decisions from six months ago, before a major refactor changed the structure underneath it, does worse than nothing. It doesn't fail quietly. It actively misleads every agent that reads it, sending each one confidently in a direction the codebase no longer supports.

Sentry's approach treats this as a problem to solve inside the file rather than around it: the second line of its AGENTS.md instructs agents and human contributors alike to update the relevant AGENTS.md file whenever they add or change agent guidance, and it explicitly bans putting that guidance anywhere else. That turns staleness into a visible violation of a written rule rather than a slow, invisible drift nobody notices until something breaks. The discipline this implies is straightforward: changes to AGENTS.md belong in the same pull request as the architectural change they describe, not scheduled as a follow-up that may or may not happen later. A constraint sitting in the file that no longer matches the code is debt carried specifically in the agent layer, and it accrues the same way any other unpaid debt does.

The harder version of this problem is one of pace. Agents can now restructure module boundaries across dozens of files in a single task, faster than a human reviewer can update the documentation describing what those boundaries used to be. The AGENTS.md file describing the old structure can go stale before the pull request implementing the new one is even reviewed. That pace argues for treating the architecture itself as something machine-readable and version-controlled, not only something described in prose that has to be manually kept in sync. When the actual graph of module relationships lives in the repository and updates automatically alongside the code, the gap between what AGENTS.md claims and what the codebase actually looks like stops opening up in the background. The file can then point to constraints that are already verifiable against the real structure, rather than asserting a shape the codebase may have already outgrown.

Sources

  1. On the Impact of AGENTS.md Files on the Efficiency of AI Coding Agents
  2. Evaluating AGENTS.md:Are Repository-Level Context Files Helpful for Coding Agents?
  3. Codified Context: Infrastructure for AI Agents in a Complex Codebase

More in Agent Workflow Patterns