Case study · AI context infrastructure · 2026
LLMSlim
The context layer for production AI agents: what every agent sees, decided, checked and on the record.
- Role
- Product, architecture, research, engineering
- Status
- Core: open source (MIT) · Platform: private beta
$ pip install llmslimThe problem
Context grows with every turn, tool and document.
Agents resend their whole prefix on every call. Old facts sit beside their corrections, retrieved text lands beside instructions, and nobody can say afterwards what the model actually saw.
- Turn 112,900
- Turn 1017,500
- Turn 4033,700
Illustrative from the briefing: 50 tools at a measured average of 122 tokens each (375 MCP schemas); history grows about 300 tokens a turn.
- Cost
- The whole prefix is billed again on every call. Caching discounts reuse, not growth.
- Correctness
- 85.7% stale-fact error with raw history in the 180-case synthetic benchmark.
- Trust
- Unlabelled baselines failed 100% of injection probes in the same benchmark.
The product
It owns the decision between your data and your model.
Every source is tagged with a role and a trust level, current facts are separated from superseded ones, the budget is planned explicitly, quality gates are checked, and the request is compiled for the provider with a trace you can replay.
Capabilities
One architecture, two products.
LLMSlim Core
Open source · MIT licence · runs locally- Open source
Compression engine
Extractive, rewrite and hybrid strategies. Extractive runs locally with no API key.
- Open source
Adaptive Context Planner
Divides a fixed token budget across competing context and records why. Reports INFEASIBLE rather than truncating silently.
- Open source
Agent Context Runtime
Prepares each agent turn as a checked, traceable unit with 10 quality gates.
- Open source
Cache-aware preparation
Keeps request prefixes byte-stable for provider prompt caching. Savings are offline estimates.
- Open source
Tool schemas & MCP
Normalises and fingerprints tool definitions; LLMSlim never executes tools.
- Open source
CLI and Studio
Inspect any agent turn in the browser or from the terminal.
LLMSlim Platform
Private beta · by invitation · runs in your cloud- Private beta
Temporal state
Each fact is active, stale, superseded, contradicted, uncertain or expired.
- Private beta
Trust firewall & provenance
Labels where context came from and how far it can be trusted.
- Private beta
Provider-neutral Context IR
One context decision compiled for OpenAI, Anthropic or Gemini.
- Private beta
Traces and replay
Immutable traces; exact replay reproduced 100% across 15 benchmark runs.
- Experimental
Context Intelligence
Learned policies run in shadow mode; no production uplift is claimed.
from llmslim import plan_context
plan = plan_context(
messages=messages, documents=docs,
query="Acme renewal",
model="sarvam-105b", max_input_tokens=8_000)
assert plan.feasibleExample adapted from the LLMSlim briefing. Local by default; no API key for extractive planning.
Provider compatibility
| Target | Level | In the code today |
|---|---|---|
| OpenAI | Maintained adapter | Responses builder, Agents SDK input filter |
| Anthropic | Maintained adapter | Messages builder with cache breakpoints |
| Google Gemini | Maintained adapter | generate_content and CachedContent builders |
| Sarvam AI | Maintained integration | Official-SDK provider; prompt caching not yet verified |
| Self-hosted (vLLM) | Maintained builder | Chat builder and Transformers KV continuation |
| MCP | Maintained adapter | Catalogs over Streamable HTTP and stdio |
Request builders are execution-free: they produce the provider payload and the host sends it. Source: repositories, verified 9 Oct 2026.
Evidence
Doubled task success, with a third fewer tokens.
On a synthetic benchmark, which is the right place to start and the wrong place to stop. Live-model outcomes with design partners are the next milestone, and no live-model improvement is claimed yet.
- task success, LLMSlim Platform
- 88.9%task success, LLMSlim Platformvs 44.4% for raw history. 180-case synthetic evaluation, deterministic oracle, Oct 2026.
- fewer tokens
- ~33%fewer tokensMean 391.8 vs 583.0 tokens against raw context, same 180 synthetic cases.
- automated tests
- 908automated tests680 Core, 228 Platform; over 90% branch coverage. Per release reports, 9 Oct 2026.
- releases in 16 weeks
- 11releases in 16 weeksJune to October 2026, per PyPI and GitHub releases.
- Raw history44.4%
- Truncation33.3%
- RAG only44.4%
- Core (prior)50.0%
- LLMSlim Core83.3%
- LLMSlim Platform88.9%
LLMSlim Platform benchmark, October 2026. Synthetic cases in English, Hindi and Hinglish across 18 categories, 3,000-token budget, deterministic oracle. Most of the gain comes from temporal labelling, which this oracle rewards by design. These are not production customer outcomes; live-model results are pending.
| Metric | Raw history | Core | Platform |
|---|---|---|---|
| Stale-fact error | 85.7% | 85.7% | 14.3% |
| Wrong context | 29.4% | 28.6% | 5.9% |
| Mean tokens | 583 | 429 | 392 |
Long sessions: in an in-memory benchmark, Platform context levelled off near 2,000 tokens, 83% fewer than raw history at 1,000 turns. Below about 140 turns, Platform's state and trace headers cost more than short raw history; the value case is long-running agents.
In their words
What early users said.
“got 50% fewer tokens on my rag pipeline with no quality drop on my eval set, super easy drop-in replacement”
Unedited public comments on LLMSlim's first Product Hunt launch (v0.2.0, 16 July 2026). See them on Product Hunt ↗
In the product
Studio and the published limits.


Timeline
Compression library to context platform in sixteen weeks.
- 16 Jun 2026
First release on PyPI
- Jul 2026
Entity rules, rewrite and hybrid modes
- Aug 2026
Provenance locks and safe tool schemas
- 12 Sep 2026
MCP catalogs, new site and Studio
- 20 Sep 2026
Adaptive Context Planner
- 30 Sep 2026
Agent runtime and cache planning
- 2 Oct 2026
Private Platform control plane
- 9 Oct 2026
Context Intelligence beta
Ecosystem & status
Startup programs, not investors.
- OpenAI for Startups
- Claude for Startups
- Sarvam Startup Program
- Zoho for Startups
- MongoDB for Startups
- Auth0 for Startups
- Zendesk for Startups
- Mixpanel for Startups
- Sentry for Startups
- Descope Hello World
- Pulumi for Startups
LLMSlim has been accepted into these programs, which offer cloud and API credits, software licences, technical resources and community access. None is an investor, lender, customer or funder, and membership implies no technical endorsement or commercial partnership.
Commercial status, October 2026
Core is free under the MIT licence. Platform is a private beta, by invitation. LLMSlim has no paying customers or revenue today, and no paid plan is live.
Building agents that need better context?
Ask about LLMSlim Platform's private beta and enterprise deployment, or hire AMEYOR to design the context and evaluation layer of your own AI product.
Or book a free 30-minute call (opens in a new tab) with the founder.
Prefer email? hello@ameyor.in

