Skip to content

AI engineering · · 4 min read

Why context engineering matters for production AI agents

Most agent failures are context failures. What grows, what goes wrong, what LLMSlim's evaluation does and does not show, and where to start.

By Yashvardhan Thanvi, founder of AMEYOR

Most production failures in AI agents are not model failures. They are context failures: the model was given the wrong facts, too many facts, stale facts, or facts it should not have trusted. The model then did exactly what it was asked, with the wrong material.

That is why we think of context engineering as its own discipline, separate from prompt writing and separate from model choice. It is the work of deciding, for every single agent turn, what the model sees, why, and in what order. This note explains the problem, what we learned building LLMSlim, and how to start applying it in your own systems.

Context grows faster than anyone plans for

A typical agent request is assembled from six kinds of material: system instructions, conversation history, retrieved documents, memory, tool results and tool schemas. Each one grows on its own schedule.

In the illustrative workload from LLMSlim's October 2026 briefing, an agent with 50 tools (at a measured average of 122 tokens per tool schema) starts at about 12,900 tokens on its first turn, reaches 17,500 by turn 10 and 33,700 by turn 40. Nobody decided that the fortieth turn should carry three times the context of the first. It simply accumulated.

Three things go wrong as that happens:

  • Cost. The whole prefix is billed again on every call. Prompt caching discounts the parts that do not change, but it does not stop growth, and new history and tool output are billed in full.
  • Correctness. Old facts sit beside their corrections. If a customer changed their renewal date in turn 6, the original date from turn 2 is still in the history unless something removes or labels it.
  • Trust. Retrieved text lands beside instructions. A document that says "ignore previous instructions" looks, to the model, a lot like an instruction.

Treat the context decision as infrastructure

The useful shift is to stop thinking of context as "the prompt" and start thinking of it as a decision that is made on every turn and can be inspected afterwards. Concretely, a well-engineered context layer should:

  1. Label every item with its role and how far it can be trusted. Developer instructions, user messages and scraped web pages should never be indistinguishable.
  2. Track state over time, so that a fact can be current, superseded, contradicted or expired, and only current facts are presented as current.
  3. Plan to a budget, explicitly. If the material does not fit, the system should decide what to compress or drop, and record why. If it still cannot fit, it should say so rather than truncate silently.
  4. Check quality gates before sending: are required facts, numbers, entities and tool contracts still intact after compression?
  5. Compile for the provider, keeping stable prefixes byte-identical so caching can work.
  6. Write a trace that lets you replay exactly what the model saw.

None of this requires a particular model vendor. It is plumbing, and it should be provider-neutral.

What the numbers do and do not say

In LLMSlim's 180-case synthetic evaluation (October 2026, English, Hindi and Hinglish cases across 18 categories, 3,000-token budget, deterministic oracle), the Platform reached 88.9% task success against 44.4% for raw history, while sending about a third fewer tokens on average (391.8 versus 583.0). Stale-fact error fell from 85.7% to 14.3%.

Those are encouraging numbers, and they come with honest limits. The cases are synthetic. The oracle is deterministic and, by design, rewards correct temporal labelling, which is where most of the gain comes from. Live-model outcomes with real users have not yet been measured. And for short sessions, the extra state and trace headers can cost more tokens than plain history; in LLMSlim's long-session benchmark the break-even sat at roughly 140 turns. Context engineering pays off most for long-running agents with many tools and conflicting sources.

We publish these caveats because a benchmark number without its conditions is not evidence. It is marketing.

Where to start

You do not need a platform to begin. Three changes help almost any agent:

  • Separate trusted from untrusted text. Wrap retrieved and user-supplied material in clearly labelled blocks and tell the model, in the system instructions, that those blocks are data rather than commands.
  • Budget explicitly. Decide how many tokens each source may use. When history exceeds its share, summarise older turns or keep only the facts that are still current.
  • Log what you sent. Store a hash or a redacted copy of each assembled request. When something goes wrong, you will want to know what the model actually saw.

If you want to experiment, LLMSlim Core is open source under the MIT licence:

from llmslim import plan_context

plan = plan_context(
    messages=messages, documents=docs,
    query="Acme renewal",
    model="sarvam-105b", max_input_tokens=8_000)

assert plan.feasible

It runs locally, and extractive planning needs no API key. The benchmarks page (opens in a new tab) documents each evaluation and what it does not measure.

The broader point

As agents take on longer tasks with more tools, the quality of their context will matter more than small differences between models. Teams that treat context as an engineered, observable layer will ship agents that are cheaper, more correct and easier to debug. Teams that keep appending to a string will keep being surprised.

If you are building agents and want help designing the context and evaluation layer, that is exactly the kind of work AMEYOR takes on.