Filesystem-based memory is the quiet default of production LLM agents, yet researchers had largely passed over how it should be organized, searched, or maintained. A 2026 paper by Zhou et al. supplies the first formal vocabulary for this design space, describing three agent roles around a single memory filesystem and measuring what organization does and does not buy. This post is a desk-research breakdown of that paper, paired with the file conventions that Anthropic, Letta, Basic Memory, and Hermes have already shipped. If you have read our agent memory systems production guide, treat this as the research paper behind the practice.

Why the default is a directory tree

The default memory store for LLM agents is a plain directory tree of markdown files, a convention so widespread that the authors of arXiv 2607.26637 call it the unstudied baseline of agent memory. The files are familiar: AGENTS.md for repository-level agent instructions, Cursor rules for editor-scoped context, Claude’s memory tool writing into a hosted /memories directory, and the Codex CLI reading markdown guidance files.

The appeal is practical. A directory tree is inspectable — you can open MEMORY.md or SKILL.md in any editor and see exactly what the agent “knows.” It is portable — copy the folder, move the agent. It versions cleanly with git, giving you history, diffs, and rollback for free. Anthropic’s engineering guidance on effective context engineering and on harness design for long-running apps both lean on lightweight identifiers (file paths) plus on-demand loading rather than stuffing everything into the context window. Our earlier production guide to agent memory systems reached the same practical conclusion before the paper formalized it.

What Zhou et al. add is the observation that this default — SKILL.md, MEMORY.md, NOTES.md, CLAUDE.md, and friends — had become the production norm through accumulation, not through study. Before this paper the memory shape, the tool harness, and the agent strengths had not been varied systematically to see which choices actually matter — the gap the paper fills.

The file conventions in one place

  • MEMORY.md / NOTES.md — the agent’s running declarative store.
  • SKILL.md — distilled procedural knowledge, one file per skill.
  • CLAUDE.md / AGENTS.md — project-scoped instructions loaded at session start.
  • /memories — the sandboxed directory exposed by Claude’s memory tool.

Three roles around one store

Zhou et al. describe three cooperating roles wrapped around a single memory filesystem: a management agent that integrates and organizes incoming content, a search agent that answers queries with cited sources, and an execution agent whose task trajectories get distilled into skills. The paper is explicit that this unifies declarative memory and procedural skills in one store (arXiv 2607.26637).

This single-store framing contrasts with architectures that split memory into separate subsystems. Letta grew out of the MemGPT line of work and treats memory as a first-class managed object inside the agent runtime, while Hermes exposes memory, skills, and context files as distinct but coexisting features. Honcho takes yet another route, modeling user and session state as a pluggable provider rather than files.

The paper’s contribution is to say: whichever runtime you use, the functional decomposition is management, search, and execution — and all three can read and write the same tree. That matters because a management agent that organizes well makes the search agent’s job cheaper, and an execution agent that writes good skill files makes future runs cheaper still. The roles are separable in implementation but coupled in effect.

Why one store beats split stores

Split architectures force you to decide, for every new fact, whether it is “declarative” or “procedural” before writing it. A single filesystem defers that decision: a note can start as NOTES.md scratch and later be distilled into a SKILL.md by the management agent. The paper’s growth study exploits exactly this flexibility — content flows in one door and can be reorganized, merged, or promoted without changing stores.

What organization actually buys

Organization buys search economy, not answer accuracy: Zhou et al. find that organized memory stores roughly halve retrieval cost where the material is large, while no agent configuration converts organization itself into better answers (arXiv 2607.26637). “Search economy” is the paper’s term for how many tokens and tool calls the search agent spends to find what it needs.

That distinction matters because most memory marketing conflates the two. Benchmarks like LoCoMo and LongMemEval measure whether the agent answers long-horizon questions correctly; the paper’s result says organization is orthogonal to that axis. A well-tended hierarchy makes retrieval cheaper, not smarter — the accuracy ceiling is set by what is in the store and by the search agent’s reasoning, not by the folder structure.

Practically, this gives you a budgeting rule. If your agent’s context is dominated by memory dumps and retrieval is eating your token bill, organization is the lever — halving retrieval cost is a real, bankable win. If your agent is already retrieving cheaply but answering badly, reorganizing files will not save you; you need better content, better search prompts, or techniques like Self-RAG that make retrieval decisions more selective. Our comparison of agent memory systems in 2026 makes the same cost-versus-accuracy split when evaluating vendors.

Where it breaks: organizational erosion

Organizational erosion is the paper’s second headline finding: as the store grows, organization decays for every management agent except the strongest, and stale or duplicated notes accumulate faster than they get cleaned (arXiv 2607.26637). Memory systems rot by default.

This connects to a known failure mode. MemGPT introduced self-editing memory precisely because unmanaged context degrades; A-MEM builds dynamic linking and evolution of memory notes because static notes go stale. Zhou et al. quantify the erosion dynamic: weaker management agents can create structure initially but cannot maintain it under continued growth, so the store drifts toward a de facto verbatim dump — losing the search-economy gains that motivated organizing in the first place.

The operational lesson is that memory maintenance is an ongoing job, not a one-time schema. You need either a strong management agent (the paper’s only stable configuration), periodic compaction passes, or an external process that prunes and merges. Anthropic’s guidance on structured note-taking and compaction treats compaction as a recurring context-management step, which is the same insight applied at the session scale. If you skip it, the directory tree quietly reverts to a junk drawer — and your retrieval costs climb back with it.

The toolset reshapes the store

The tool harness reshapes the memory store as strongly as swapping the underlying model: Zhou et al. report that changing the tool set alone produces store-shape differences on the order of changing models entirely (arXiv 2607.26637). The tools an agent has determine what kind of memory it can even express.

The paper contrasts two harnesses: a sandboxed shell, where the agent can grep, find, and rewrite files freely, versus a constrained memory-tool interface. Claude’s memory tool (memory_20250818) exposes six file operations — view, create, str_replace, insert, delete, and rename — against a /memories directory the developer hosts, as documented in Anthropic’s memory tool reference and their context engineering cookbook. The shell is expressive but error-prone; the function interface is safe but narrows what the agent writes and how.

This has direct cost implications beyond memory. We audited how tool schemas inflate context bills in our MCP token waste audit, and the same logic applies here: every tool definition costs tokens every turn, and every tool’s affordances bias the agent’s writing behavior. If you connect memory through MCP servers — for example, filesystem servers from the reference server collection — you are making a store-shape decision, not just a plumbing decision. Choose the toolset for the memory shape you want, then keep the model constant while you tune it.

Memory shapes compared

Memory shape is the paper’s central design axis, and the choice among agent-organized hierarchies, verbatim dumps, chunk retrieval, tool-managed directories, and git-backed filesystems trades organization cost against retrieval cost and accuracy. The table below maps the three shapes the paper studies directly (agent-organized hierarchy, verbatim dump, chunk retrieval) alongside two ecosystem shapes (tool-managed directories, git-backed MemFS); cells for the ecosystem rows are the author’s synthesis, not paper measurements.

Memory shape Retrieval mode Organization cost Retrieval cost effect Accuracy effect Best fit
Agent-organized hierarchy Path-guided file reads High, ongoing (erosion risk) Roughly halved at scale Neutral per the paper Long-lived agents with a strong management agent
Verbatim dump Full-context or grep scan None High and growing Neutral (lookups get noisier) Short-lived agents, small stores
Chunk retrieval / vector index Embedding similarity search Medium (index maintenance) Low per query Benchmark-dependent (LoCoMo, LongMemEval) High-volume corpora, semantic lookup
Tool-managed /memories directory Six file ops on demand Low-medium Moderate Neutral Claude-style agents with hosted sandboxes
Git-backed MemFS File reads plus versioned diffs Medium (commit discipline) Moderate Neutral Teams needing audit trails and rollback

Agent-organized hierarchy is the paper’s strongest configuration: the management agent builds a taxonomy, and the search economy roughly halves where material is large — but only a strong management agent holds that structure as the store grows (arXiv 2607.26637). Verbatim dump is the paper’s baseline: zero organization cost, rising retrieval cost, and the shape erosion drifts toward.

Chunk retrieval is what Mem0 and Zep productize — embedding indexes over extracted memory items, with vendor-side maintenance absorbing the organization cost. Tool-managed /memories directories, per Anthropic’s memory tool, sit between hierarchy and dump: structured enough for cheap lookup, loose enough to erode. Git-backed MemFS, as Letta implements, adds versioning on top, which our 2026 memory comparison flags as the main differentiator for regulated environments. Basic Memory occupies a related niche, building agent memory on local markdown with a knowledge-graph structure over it.

Patterns you can implement today

Four implementable patterns come straight from the paper and the shipping tools: just-in-time file loading, structured note-taking, scheduled compaction, and explicit file tools. Anthropic’s guidance is that agents should keep lightweight identifiers — file paths — in context and load full content only when needed (effective context engineering), which is the cheapest search-economy win available.

Structured note-taking means the agent writes findings into named files with consistent frontmatter rather than appending prose to one giant log; the paper’s management role exists to enforce exactly this. Compaction — summarizing and pruning older notes on a schedule — is the countermeasure to organizational erosion, and Letta’s MemFS makes it auditable by committing every change to git. Explicit file tools, whether Claude’s six memory operations or a shell sandbox, matter because the toolset shapes the store as strongly as the model does.

Concrete starting points: stand up a /memories directory with a MEMORY.md index and per-topic files; give your agent Hermes-style context files for project-scoped knowledge; adopt Basic Memory if you want markdown plus a graph layer; or run Letta if you want the runtime to manage memory for you. Pair these with the prompt-level patterns in our context engineering patterns for 2026.

FAQ: filesystem agent memory

Is filesystem memory just putting markdown files in a repo?

No — the files are only the substrate; the design lives in the three agent roles and the tool harness around them. Zhou et al. show that the same directory behaves very differently depending on whether a management agent organizes it, which file operations the agent can perform, and how the search agent retrieves (arXiv 2607.26637). A repo of markdown without those roles is closer to the paper’s verbatim-dump baseline than to a managed memory system.

What does the paper mean by “search economy”?

Search economy is the token and tool-call cost the search agent pays to find relevant material in the store. The paper measures it directly and finds that organized stores roughly halve retrieval cost where material is large, making organization a cost lever rather than a quality lever (arXiv 2607.26637). Where accuracy benchmarks such as LongMemEval score whether answers are correct, search economy scores what the lookup costs in tokens and tool calls.

Does better file organization improve answer accuracy?

No — the paper explicitly finds that no agent configuration converts organization into better answers; organization’s measurable benefit is cost reduction, not accuracy (arXiv 2607.26637). Accuracy on long-horizon question sets like LoCoMo depends on what is stored and how the search agent reasons. Treat organization as an infrastructure investment that lowers bills, and pursue accuracy through content quality and search prompting instead.

Do I need a memory tool, or can agents use plain shell file commands?

You can use either, but the choice reshapes the store. The paper finds the tool harness alone changes memory shape as strongly as swapping models: a sandboxed shell permits free-form reorganization, while Claude’s memory tool constrains agents to six operations — view, create, str_replace, insert, delete, rename — on a hosted /memories directory (Anthropic docs). Shells are expressive but error-prone; memory tools are safe but narrower.

How strongly does the toolset shape the store compared with the model?

Roughly equally, per the paper: changing the tool set alone reshapes the memory store about as strongly as swapping the model entirely (arXiv 2607.26637). That means tuning your file tools, your MCP server choices (modelcontextprotocol.io), and your harness is a first-order design decision, not a second-order detail after model selection.

The Bottom Line

Filesystem memory won by default, and Zhou et al. finally give the default a design vocabulary: three roles around one store, organization that halves retrieval cost but never lifts accuracy, erosion that only strong management agents resist, and toolsets that shape the store as much as models do. The verdict from this desk research: adopt file-based memory deliberately — budget for ongoing organization, pick your tool harness for the store shape you want, and stop expecting folder structure to fix bad answers. Start with just-in-time loading and compaction, then graduate to a managed hierarchy only if your token bills demand it.

How This Guide Was Built

This guide is desk research, not a benchmark: no hands-on testing, no local query counts, no measured author-machine results. Findings are attributed to arXiv 2607.26637 (also available as HTML, PDF, and on Hugging Face papers) or to the vendor documentation and papers cited inline — Anthropic’s engineering notes, Letta’s docs, Hermes’s guides, and the related-memory papers MemGPT, A-MEM, LoCoMo, LongMemEval, and Self-RAG. Product claims reflect vendor documentation as of October 2026, not independent verification. Discussion and prior arena entries live on the NiteAgent arena.

← Back to all posts