When Chroma engineers tested 18 large language models (LLM) across eight lengths, scores dropped as the prompts grew — even on simple tasks [1]. Long context, it turns out, is a workspace, not a hard drive.

The big picture:

For two years, the pitch was simple: a giant context window solves the AI memory problem.

Vendors told builders to paste the whole chat history, let the model sort it out, and stop paying for a separate database.

The data says otherwise. On LongMemEval, a test built to check recall across sessions, popular chat bots and long-context models lost 30% of their accuracy on 500 memory questions [2].

Vendors now admit the limits. Anthropic calls context “a finite resource” and tells subagents to hand back short 1,000-to-2,000-token summaries instead of raw chat logs [5].

What strikes me is that the fix is a design choice, not a bigger model. The teams I see surviving long tasks choose exactly which tier of memory holds each fact. One team kept a flat history past 12,000 tokens, and its logic drifted; splitting that state into four tiers cut their overhead by 68%.

By the numbers

  • 18 LLMs — Context rot: A 2025 report tested 18 models and found performance drops as input grows, even on basic tasks [1].
  • 30% — Recall gap: Chat bots and long-context models lose 30% accuracy when asked to recall facts across long chats [2].
  • 26% — Extraction boost: A memory layer that pulls out facts and brings them back gives a 26% gain in scored answers over a full-context baseline, and speeds up the slowest calls by 91% [3].
  • 74.0% — Files beat tools: A file-backed agent on a small model hit 74.0% on a memory test, beating the 68.5% scored by a leading memory tool [4].

What I’d watch:

The builders closest to this problem have stopped asking how big the window is and started asking which tier holds each fact. Here is what I am watching next.

  • Paging out early: Teams use the live prompt as a scratchpad, clean it up between loops, and move lasting facts to a database they can search. This layered setup is cheaper than a giant window.
  • Summaries as rules: Anthropic notes a worker agent may burn tens of thousands of tokens exploring, but should return only a short summary [5]. I read that as a strict rule, not just a tip.
  • Files before tools: Letta scored 74.0% on a memory test using plain files, with no special search tool [4]. This shows an agent’s built-in rules matter more than the memory product you buy.
  • Where exact facts go: Builders are sending money, dates, and IDs to fixed databases, not vector search engines. Search by meaning is for fuzzy recall, not for the exact truth.

The catch

My read is that the memory-tool market is more messy than the press releases admit.

The two strongest claims in this space fight each other. A recent paper claims a 26% gain from a new memory layer [3], while Letta shows a plain file system beats that same type of tool [4].

Both tests measure bounded tasks, not messy live setups. Every new database adds a write path, a delay, and a fresh place for the system to break.

At a glance

  • The Big Shift: New tests prove a giant context window does not fix AI memory. On one test, 18 models got worse as input grew [1], and another test saw a 30% recall drop across long chats [2].
  • Why It Matters: If recall falls as history grows, where an agent keeps its facts decides if a long task finishes. This is a design choice that is cheaper to get right than to buy.
  • What I’d Watch:
  • Early paging: How teams shrink the live prompt between loops instead of letting a flat history grow.
  • Summary contracts: How worker agents return short summaries rather than full chat logs.
  • Files vs. tools: How plain file systems keep beating niche memory products on the same tests.
  • Exact-fact stores: How money, dates, and IDs move out of vector search into fixed databases.
  • The Catch: The two best memory tests contradict each other, both rely on bounded tasks, and every added tier brings new write paths and failure points.

Related reading

Sources

[1] Chroma, “Context Rot: How Increasing Input Tokens Impacts LLM Performance,” research.trychroma.com. https://research.trychroma.com/context-rot [2] Wu et al., “LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory,” arXiv:2410.10813 (v2 2025-03-04). https://arxiv.org/abs/2410.10813 [3] Chhikara et al., “Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory,” arXiv:2504.19413 (2025-04-28). https://arxiv.org/abs/2504.19413 [4] Letta, “Benchmarking AI Agent Memory: Is a Filesystem All You Need?” Letta Research Blog (2025-08-12). https://www.letta.com/blog/benchmarking-ai-agent-memory [5] Anthropic Engineering, “Effective context engineering for AI agents” (2025-09-29). https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents