Skip to content
< samuelsantana.dev />
Back to the BlogThe /context Memory files list: two CLAUDE.md files loaded inside the context window, and AGENTS.md struck through outside it, not loaded.

Context engineering in practice: what the agent doesn't load doesn't exist

Samuel Santana
Published on September 20, 2026
IASoftware Architecture

On August 19 I found out that the rules I had written for all my repositories had never applied to a single session. They lived in an AGENTS.md in the folder that holds my projects, titled "Codex Custom Instructions for Developers". I don't use Codex. I use Claude Code, and Claude Code at that point read CLAUDE.md. The file was on disk, well written, and outside the context of every session.

Anthropic defines context engineering as the set of strategies for curating and maintaining the optimal set of tokens during model inference, "including all the other information that may land there outside of the prompts" (Effective context engineering for AI agents, September 2025). My July post on tokens, context, skills and agents introduces the pieces. This one is about what I do with them and, above all, about what broke.

Everything here happened in Claude Code between July and September 2026, and every mechanism I cite comes from the documentation or the changelog of the time.

What wasn't loaded doesn't exist

The Claude Code memory page, as archived on August 18, settled the question in one sentence: "Claude Code reads CLAUDE.md, not AGENTS.md." For anyone who already had an AGENTS.md for other agents, it recommended creating a CLAUDE.md that imports it.

The odd part is that I had read the opposite in another official document. The agents guide that ships inside the Next.js 16.2.10 package, at node_modules/next/dist/docs/01-app/02-guides/ai-agents.md, said: "Most AI coding agents — including Claude Code, Cursor, GitHub Copilot, and others — automatically read AGENTS.md when they start a session" (the file, straight from the published package). A few lines further down, the same page explained that create-next-app also generates a CLAUDE.md containing @AGENTS.md, so Claude Code users get the same instructions. If Claude Code read AGENTS.md on its own, the import would be pointless. The import was the right part; the sentence wasn't.

In my Next.js projects this never bit me, because since July 7 the vertex-web .claude/CLAUDE.md imported the file:

# Vertex Web - System Context & AI Agent Rules

@../AGENTS.md

Nobody imported the AGENTS.md in the parent folder. On August 19 it became D:\github\CLAUDE.md. Since Claude Code loads CLAUDE.md from the working directory and every directory above it, a file in the parent folder enters every session opened inside any repository.

Two days ago, on September 18, version 2.1.277 started reading AGENTS.md: "in a project with no CLAUDE.md, Claude Code reads AGENTS.md instead". The condition is what matters. According to the documentation from that same day, if there is a CLAUDE.md, .claude/CLAUDE.md or CLAUDE.local.md in the working directory or above it, Claude Code reads those and "ignores every AGENTS.md". To read both, you change the Project instructions setting to claude-md-and-agents-md. The sessions I opened in vertex-web, which had its own CLAUDE.md, would still have gone without that AGENTS.md, even on the release from two days ago. And today, with a CLAUDE.md in the parent folder itself, every AGENTS.md under it is ignored under the default setting.

The lesson isn't "use the right file name". It is to swap the question "which file does this tool read?", whose answer depends on version, configuration and which documentation you happened to read, for "which files made it into this session?". The documentation already gave the command in August: /context, under Memory files. If the file isn't listed there, the model doesn't see it.

Loading is not overriding

The same page has a second detail, and it changes how you write instruction files at more than one level. In Claude Code, the CLAUDE.md files it finds "are concatenated into context rather than overriding each other", from the filesystem root down to the working directory. The agents.md convention says something else: in monorepos, "the closest one takes precedence". Similar files, different semantics.

With concatenation, both files reach the model, and the documentation warns what happens when they disagree: "if two rules contradict each other, Claude may pick one arbitrarily". So the precedence I want is written down, at both ends. The parent-folder file is in Portuguese; this line says that each repository's own CLAUDE.md wins on conflict:

- Cada repo tem seu próprio `CLAUDE.md` com as convenções específicas — **ele vence este arquivo**
  em caso de conflito.

And at the top of the repository's file:

> `D:\github\CLAUDE.md` applies too. Where the two disagree, this file wins.

It's crude: the model reads both, and the text itself says which one wins. Better still would be no contradiction at all, and the documentation recommends reviewing the files periodically; until the next review, the written rule is what there is.

A document that exists can also not exist

On August 21, a session working on the blog's prerendering justified a choice with "otherwise Google indexes the preview". The repository's own docs/rendering-strategies.md had already tested and refuted that argument: Vercel previews sit behind SSO and respond with X-Robots-Tag: noindex. The document was there. It just wasn't in context, because the docs/ folder doesn't load by itself; CLAUDE.md does. My notes from that day record the mistake the right way: the argument was written without reading the document the repository already had.

The answer to that isn't "read the docs more". It's moving the conclusion into the file that always loads, with the reason attached. This is what vertex-web's CLAUDE.md says today:

- **Previews are behind Vercel SSO.** `curl` gets a 302 to `vercel.com/sso-api`. [...] They also carry
  `X-Robots-Tag: noindex`, which is why "previews would get indexed" is never a valid argument here.

This is the hybrid model the Anthropic article describes for Claude Code itself: "CLAUDE.md files are naively dropped into context up front, while primitives like glob and grep allow it to navigate its environment and retrieve files just-in-time". What goes up front is what the agent can't afford not to know; the rest stays one pointer away. In my setup, CLAUDE.md carries conclusions and rules that are expensive to break, docs/ carries the measurements and the history, and CLAUDE.md says which document to read before touching what. One vertex-web section is titled "Rendering — read docs/rendering-strategies.md before touching a route", and right below it comes the short version of the conclusions, in case nobody opens the document.

Why not put everything in the file that always loads

Because context isn't free. The same Anthropic article calls the problem context rot: "As the number of tokens in the context window increases, the model's ability to accurately recall information from that context decreases". The model has an "attention budget", and every token spends a little of it.

The public evidence points the same way. Chroma's technical report (July 2025) tested 18 models: performance "varies significantly as input length changes, even on simple tasks", and on LongMemEval every model did significantly better with the focused prompt than with the full one. Earlier, Lost in the Middle (Liu et al., TACL) showed performance dropping when the relevant information sits in the middle of a long context.

The Claude Code documentation turns this into a working rule: aim for under 200 lines per CLAUDE.md, because "longer files consume more context and reduce adherence". Splitting the file into @imports organizes it but saves nothing, since imported files also load at launch.

One caveat: these studies measure long inputs in general, not instruction files, and I have no adherence measurement for my own CLAUDE.md. What I take from them is the direction: every line in the always-loaded file competes with the task for attention.

A good rule carries its reason and its proof

vertex-web's CLAUDE.md opens with a section called "Rules that will bite": the rules whose violation has already cost something, each with its reason and its evidence. The first one:

- **Merging does not publish.** Vercel is not set to auto-promote here; after a merge to `main`,
  Samuel promotes the deployment in the dashboard by hand. The Vercel MCP connector cannot do it —
  verified: it lists, inspects, triggers deploys and changes protection, and does not promote.
  "Merged" is not "live".

Another, in the rendering section, says useCurrentUser() has three states, not two, and ends with the cost of forgetting: "it has happened twice, in the header and nearly again in the comments".

The reason isn't decoration. The documentation says CLAUDE.md content "is delivered as a user message after the system prompt" and that there is no "guarantee of strict compliance". A rule without a reason is fragile in two ways: it gets applied as ritual where it doesn't fit, or dropped as soon as a plausible argument shows up. With the reason, the model can tell whether the case at hand fits. So can I, weeks later.

The documentation asks for instructions "concrete enough to verify". The "verified: it lists, inspects..." is exactly that: it says which test was done, so nobody has to redo it.

And whatever must hold no matter what shouldn't depend on context. Claude Code treats these files "as context, not enforced configuration"; to block an action regardless of what the model decides, the documentation points to a hook. The same principle applies outside the agent: in vertex-web, secrets in a commit are stopped by a git hook, secretlint running over staged files through husky and lint-staged. No instruction does that job.

Context files are infrastructure

On August 19 I decided that in the three active repositories, CLAUDE.md and the docs/ folder would become gitignored and untracked. Agent instructions and plans live on disk, not on GitHub. The decision had two consequences: one immediate, one delayed.

The immediate one was dead pointers. In one repository, twenty references pointed at files that were leaving the public tree, three of them visible in the published Storybook, including a link to a DESIGN.md on GitHub that would have returned 404. Each was rewritten to carry its own reasoning, and the rule went into all three CLAUDE.md files: a comment that says "see docs/x.md" is a dead pointer for anyone reading the repository on GitHub; the why lives at the point of use.

The delayed one showed up on August 21: CLAUDE.md had vanished from disk in all three repositories. No mystery. Git had stopped being the backup for those files, and nothing replaced it. I rebuilt vertex-web's from what was verifiable in the repository, and the new file opens by saying so:

> **This file was rebuilt on 21/08/2026 from what is verifiable in the repository.** The original was
> gitignored (PR #79) and then lost from disk — a clone does not restore what git does not track, and
> nothing noticed until a session went looking for it. Anything the old file recorded that is *not*
> derivable from the code is gone; what follows was checked against the source, the build output and
> the CI config, not remembered.

What couldn't be rebuilt is precisely what matters most: whatever can't be derived from the code. The documentation gets there from the other side: the /doctor check "cuts content Claude can derive from the codebase" and keeps "pitfalls, rationale, and conventions that differ from tool defaults". The valuable content of a context file is the part with no other copy. If it leaves git, it needs a backup somewhere else; that day, this went down as an open decision.

Context files age too

The parent-folder CLAUDE.md said four repositories "ainda têm AGENTS.md — não foram tocados" (still have AGENTS.md, untouched). Git says otherwise: in three of them (vertex-api, vela-core and vela-ui), the file was removed on August 19, in commits made between 4:43 and 4:45 p.m., Brasília time The sentence was still there on September 7, when a session loaded it exactly as written. It was only corrected on September 18, and the correction came with a date: "Conferido em 18/09/2026" (checked on September 18, 2026).

The date is the part worth copying. It tells whoever reads it, person or model, how old the fact is. Claude Code does the same with auto memory: it records the write time in a modified frontmatter field, because "the timestamp shows how current the fact is".

The handoff: the note is a hypothesis, the system is the fact

"Each Claude Code session begins with a fresh context window", the documentation says. Between sessions, only what is written crosses over. In my case, besides the CLAUDE.md files, a GAPS.md crosses over, with "▶ RETOMAR AQUI" ("resume here") sections, newest on top. The first section of the parent-folder CLAUDE.md says to read it before anything else. It is far too large to enter context whole, and it doesn't need to: the instruction points at one section. It's a homemade version of what the Anthropic article calls structured note-taking. The skeleton, simplified and translated:

## ▶ RETOMAR AQUI — updated on August 21, 2026

**Session on <workstream>.** What was asked and what actually happened.

### ✅ Delivered
| PR | What | State (merged, published, waiting) |

### 🔍 Assumptions the code overturned

### ⏳ Pending, and on whom

Verified on <date>: <how, with which command or header, not "assumed">

The most useful rule in this arrangement isn't in GAPS.md. It's in CLAUDE.md, right below the instruction to read it. In Portuguese in the original, it says: check before executing; on August 19, two items listed as open were already done (the domain on Vercel and the domain verified in Resend); an item written by the previous session is a hypothesis about the past, not the current state.

**Conferir antes de executar.** Em 19/08 duas pendências listadas como abertas já estavam feitas —
o domínio na Vercel e o domínio verificado no Resend. Um item escrito pela sessão anterior é uma
hipótese sobre o passado, não o estado atual.

Two days later the same pattern came from a plan. The code overturned both premises of the blog prerendering plan: removing cookies() did not make the post route static, because generateStaticParams was missing, and the admin panel, which "didn't need translation", was translated on purpose in all three languages. That day's note summed it up: "a plan written at night is a hypothesis about the past, and the repo wins".

What made the notes more useful over time was saying how each thing was verified: "verified in the Vercel build log and by HTTP header, not by the local route table". A claim that says how it was checked can be checked again in seconds.

Memory, skills and tools: what enters only when needed

Claude Code's auto memory keeps an index, MEMORY.md, whose first 200 lines (or 25 KB) enter every session, plus one file per memory, read on demand. There are four types, user, feedback, project and reference, recorded in the frontmatter. In mine, the correction and project memories carry the fact plus two more lines: Why and How to apply.

An example from September 7. When I authorize a risky action, I usually attach a verification condition; that day it was "check in the plan that it only creates the role and the policy". The Terraform plan showed two creations and two changes, and one of the extra changes would revert a production variable to a sample value, which would have taken down the site's admin access. The session stopped and reported. The memory that stayed records the rule, the case and what to do when the condition fails: don't complete the action, not even "just the safe part" without saying so. Without the why, the next session would treat the condition as a box to tick.

Procedures go into skills. vertex-web's, .claude/skills/verify/SKILL.md, has existed since July 10: how to bring the site up with vertex-api and Postgres, the data-seeding gotchas, and a trap that cost me time, namely that searching the response for a message's text gives false positives, because every page embeds the full next-intl catalog in its RSC payload. It's a recipe, not a fact for every session, and the documentation draws the same line: if an entry "is a multi-step procedure or only matters for one part of the codebase, move it to a skill or a path-scoped rule instead". Until it's used, a skill takes up only its name and description in context, which Anthropic calls progressive disclosure.

Tools are context too. Twice, an MCP connector added to the account mid-session didn't show up in it, only in the next one: Resend's in August and AWS's on August 31. The documentation doesn't address that case explicitly, so I treat it as observed behavior, not a product rule. Even so, it became a line in CLAUDE.md, and sessions start by checking which tools loaded.

And there is what must never enter, because everything that passes through context stays in the transcript. Connection strings and API keys are never pasted into the chat or printed in tool output, and creating an API key through the connector is forbidden: the tool's response would carry the secret with it.

How I check today

None of this needs new tooling, only checking instead of assuming:

  • At session start, /context: is every file I expect listed under Memory files? If the project depends on AGENTS.md, also look at Project instructions in /config.
  • Ignored instruction: did the file load? Does another instruction contradict it? Is it specific enough to be verified?
  • Dead pointers in every public repository, before opening a PR:
# Tracked files that cite documents that aren't on GitHub
git grep -nE "see docs/|CLAUDE\.md|AGENTS\.md" -- . ':!*.md' ':!.gitignore'
  • Every claim in a note with a date and a method: "checked on", with the command or header that proved it.
  • What must always hold goes into a hook or CI, not into CLAUDE.md.

Why this matters

The model only follows what it sees, and it sees it as a user message, with no guarantee of obedience. That is why context engineering, day to day, is less about prompt prose and more about operations: what loads, in what order, who wins on conflict, what stays out until needed, what must never enter, and how old each fact is.

None of the failures I described was about model capability: a file that didn't load, a document nobody consulted, a file that vanished from disk, a sentence that aged, and pending items that were already done. They were all about context, and all of them would have been caught by someone who checked instead of assuming. That someone can be the agent itself, as long as the context tells it to check.

References

Comments

Loading comments...