LLM context engineering

Context engineering is the discipline of choosing which tokens occupy a language model’s context window at each inference step. Where prompt engineering asked how to phrase instructions, context engineering asks what configuration of the whole context state — system instructions, tools, memory, retrieved data, message history — is most likely to produce the desired behavior, and it asks that question again on every turn of an agent loop.1

The constraint: context is a finite resource

Transformer attention pairs every token with every other token, so n tokens cost n² pairwise relationships; models are also trained on distributions where shorter sequences dominate, leaving fewer parameters specialized for context-wide dependencies. The result is a performance gradient rather than a hard cliff: recall and long-range reasoning degrade as context grows, a failure mode the literature calls context rot.2

Liu et al. measured the gradient directly. In multi-document QA and synthetic key-value retrieval, moving the relevant passage through the input produces a U-shaped performance curve: models handle information at the very beginning (primacy) or end (recency) of the window well, and degrade sharply in the middle. In the worst cases GPT-3.5-Turbo’s mid-context accuracy fell below its closed-book accuracy with no documents at all (56.1%), drops exceeded 20 points at 20–30 documents, and some models failed exact-match retrieval of a UUID pair among 140–300 distractors. Encoder-decoder models were position-robust only within their training-time sequence length, and extended-context variants were not reliably better than their base models at actually using the extra window.3

The engineering consequence both sources draw: context must be treated as a budget, and good context engineering finds the smallest set of high-signal tokens that maximizes the likelihood of the desired outcome.45

Two channels put content into context

Unconditional injection — the system prompt and other always-on instructions — places content in the window on every turn, no relevance test applied. Every turn pays its attention cost, though prompt caching makes its dollar cost near-zero after the first turn.

Query-conditioned retrieval admits content because it matches the current turn: embedding-based memory recall, RAG over documents, or “just-in-time” loading where the agent follows lightweight identifiers (paths, queries, links) and pulls data in with tools. The 2024 agent-memory survey taxonomizes the retrieval side as three operations — memory writing (extracting and compressing observations), memory management (summarizing, merging, forgetting, reflection), and memory reading (recency truncation or top-K embedding retrieval) — fed by three sources (inside-trial, cross-trial, and external knowledge) and stored in textual or parametric form.6 Anthropic reports production agents converging on a hybrid: a small file of standing instructions injected up front, everything else retrieved just-in-time with primitives like glob and grep.7

MemGPT is the purest statement of the retrieval architecture: virtual context management, in which the LLM pages data between a small “main context” and external stores through its own function calls, with memory-pressure warnings delivered as interrupts — the operating-system memory hierarchy re-implemented in prompts.8

The trigger problem: what retrieval cannot surface

Retrieval admits content by relevance to a query — the current turn’s text, or a query the agent formulates. A behavioral rule whose trigger condition lives in the surrounding session state rather than in anything anyone says presents no lexical material for a matcher to hit. A rule of the form “when the session is a voice call, format responses this way” will surface at random in text chats where it does not apply and stay buried during voice calls where it does, because the voice-call marker never appears in the query stream. This follows from retrieval requiring a query at all; it cannot be tuned away inside any particular retriever.

Two allocation implications:

  1. Unconditional rules and trigger-in-context rules belong in the injected channel. Persona, style, modality-conditional formatting, and hard constraints must be evaluated by the model against session state every turn, which only the always-on prompt guarantees. Keep them brief — injected tokens spend the attention budget.
  2. Accumulating, topically organized content belongs in retrieval. Facts, preferences, project state, and prior decisions grow without bound and are needed only when relevant; recall-by-topic serves them well and keeps them out of the budget until then.

(Generalized from a 2026-09-25 design discussion about whether Hermes’ voice-mode formatting rules should migrate from the SOUL.md system prompt into Mnemosyne canonical memory. The reasoning is implementation-independent: the memory plugin’s prefetch matches canonical memories against query tokens, exactly the channel mismatch described above.)

Long-horizon continuity

When a task outgrows any window, three patterns recur. Compaction summarizes a nearly-full conversation and restarts with the summary — tune for maximum recall first, then prune. Structured note-taking (agentic memory) has the agent persist notes outside the window and re-read them after resets, as Claude’s Pokémon run demonstrated across thousands of game steps. Sub-agent architectures isolate exploration in focused context windows that return 1–2K-token distillations to a coordinating agent.9 All three are injection-rebuilding processes: they manufacture the small token set that the next window will unconditionally contain, which makes them the continuity mechanism that respects the same budget the rest of the discipline optimizes.

Open questions

  • Non-lexical triggers. Can retrieval key on structured session metadata (channel, modality, time) as a first-class field rather than text similarity, without collapsing back into hand-written rule engines?
  • Prompt size versus attention. The U-curve was measured on 2023 models; how it scales with current long-context training, and whether a large always-on prompt measurably taxes mid-context retrieval, is unquantified.
  • Memory evaluation. The survey notes memory modules lack standardized evaluation — benchmarks are task-level, so whether a given memory system actually helps is under-instrumented.10

Connections

  • attention-economy — the same scarcity logic at two substrates: human attention harvested by feeds, model attention spent token by token
  • discontinuous-self — bundle-theory selfhood implemented in engineering: an agent’s continuity across sessions is precisely what its memory architecture constructs
  • price-of-intelligence — token curation is cost curation while context stays metered, even as quality-adjusted prices fall
  • llm-inference-provider-landscape — the market layer that meters the window this page curates

Sources

Footnotes

  1. Prithvi Rajasekaran, Ethan Dixon, Carly Ryan, and Jeremy Hadfield (Anthropic Applied AI) 2025 — Effective context engineering for AI agents ↩

  2. Prithvi Rajasekaran, Ethan Dixon, Carly Ryan, and Jeremy Hadfield (Anthropic Applied AI) 2025 — Effective context engineering for AI agents ↩

  3. Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, Percy Liang 2023 — Lost in the Middle: How Language Models Use Long Contexts ↩

  4. Prithvi Rajasekaran, Ethan Dixon, Carly Ryan, and Jeremy Hadfield (Anthropic Applied AI) 2025 — Effective context engineering for AI agents ↩

  5. Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, Percy Liang 2023 — Lost in the Middle: How Language Models Use Long Contexts ↩

  6. Zeyu Zhang, Xiaohe Bo, Chen Ma, Rui Li, Xu Chen, Quanyu Dai, Jieming Zhu, Zhenhua Dong, Ji-Rong Wen 2024 — A Survey on the Memory Mechanism of Large Language Model based Agents ↩

  7. Prithvi Rajasekaran, Ethan Dixon, Carly Ryan, and Jeremy Hadfield (Anthropic Applied AI) 2025 — Effective context engineering for AI agents ↩

  8. Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Ion Stoica, Joseph E. Gonzalez 2023 — MemGPT: Towards LLMs as Operating Systems ↩

  9. Prithvi Rajasekaran, Ethan Dixon, Carly Ryan, and Jeremy Hadfield (Anthropic Applied AI) 2025 — Effective context engineering for AI agents ↩

  10. Zeyu Zhang, Xiaohe Bo, Chen Ma, Rui Li, Xu Chen, Quanyu Dai, Jieming Zhu, Zhenhua Dong, Ji-Rong Wen 2024 — A Survey on the Memory Mechanism of Large Language Model based Agents ↩