Skip to content

Case study · AI context infrastructure · 2026

LLMSlim

The context layer for production AI agents: what every agent sees, decided, checked and on the record.

Role
Product, architecture, research, engineering
Status
Core: open source (MIT) · Platform: private beta
$ pip install llmslim
Official LLMSlim demo (1 min 46 s, with sound): prompt compression, growing agent context, the Python API, the terminal and Studio. Figures shown come from LLMSlim's published benchmarks.

The problem

Context grows with every turn, tool and document.

Agents resend their whole prefix on every call. Old facts sit beside their corrections, retrieved text lands beside instructions, and nobody can say afterwards what the model actually saw.

Context per request, illustrative agent
  • Turn 112,900
  • Turn 1017,500
  • Turn 4033,700

Illustrative from the briefing: 50 tools at a measured average of 122 tokens each (375 MCP schemas); history grows about 300 tokens a turn.

Cost
The whole prefix is billed again on every call. Caching discounts reuse, not growth.
Correctness
85.7% stale-fact error with raw history in the 180-case synthetic benchmark.
Trust
Unlabelled baselines failed 100% of injection probes in the same benchmark.

The product

It owns the decision between your data and your model.

Every source is tagged with a role and a trust level, current facts are separated from superseded ones, the budget is planned explicitly, quality gates are checked, and the request is compiled for the provider with a trace you can replay.

How LLMSlim sits between an application and its modelSix context sources flow into LLMSlim, which labels trust, tracks state, plans to a token budget, checks quality gates, compiles for the provider and writes a trace. The host application then calls the model API. LLMSlim never executes tools or model calls.InstructionsHistoryDocumentsMemoryTool resultsTool schemasLLMSlimOWNS THE CONTEXT DECISION01Label trust02Track state03Plan to budget04Check gates05Compile06TraceModel APIcalled by the host:OpenAI, Anthropic, Gemini,Sarvam, vLLM
Diagram drawn by AMEYOR from the LLMSlim product briefing (Oct 2026). LLMSlim prepares context; it never executes tools or model calls.

Capabilities

One architecture, two products.

LLMSlim Core

Open source · MIT licence · runs locally
  • Compression engine

    Extractive, rewrite and hybrid strategies. Extractive runs locally with no API key.

    Open source
  • Adaptive Context Planner

    Divides a fixed token budget across competing context and records why. Reports INFEASIBLE rather than truncating silently.

    Open source
  • Agent Context Runtime

    Prepares each agent turn as a checked, traceable unit with 10 quality gates.

    Open source
  • Cache-aware preparation

    Keeps request prefixes byte-stable for provider prompt caching. Savings are offline estimates.

    Open source
  • Tool schemas & MCP

    Normalises and fingerprints tool definitions; LLMSlim never executes tools.

    Open source
  • CLI and Studio

    Inspect any agent turn in the browser or from the terminal.

    Open source

LLMSlim Platform

Private beta · by invitation · runs in your cloud
  • Temporal state

    Each fact is active, stale, superseded, contradicted, uncertain or expired.

    Private beta
  • Trust firewall & provenance

    Labels where context came from and how far it can be trusted.

    Private beta
  • Provider-neutral Context IR

    One context decision compiled for OpenAI, Anthropic or Gemini.

    Private beta
  • Traces and replay

    Immutable traces; exact replay reproduced 100% across 15 benchmark runs.

    Private beta
  • Context Intelligence

    Learned policies run in shadow mode; no production uplift is claimed.

    Experimental
plan.py
from llmslim import plan_context

plan = plan_context(
    messages=messages, documents=docs,
    query="Acme renewal",
    model="sarvam-105b", max_input_tokens=8_000)

assert plan.feasible

Example adapted from the LLMSlim briefing. Local by default; no API key for extractive planning.

Provider compatibility

LLMSlim provider support levels
TargetLevel
OpenAIMaintained adapter
AnthropicMaintained adapter
Google GeminiMaintained adapter
Sarvam AIMaintained integration
Self-hosted (vLLM)Maintained builder
MCPMaintained adapter

Request builders are execution-free: they produce the provider payload and the host sends it. Source: repositories, verified 9 Oct 2026.

Evidence

Doubled task success, with a third fewer tokens.

On a synthetic benchmark, which is the right place to start and the wrong place to stop. Live-model outcomes with design partners are the next milestone, and no live-model improvement is claimed yet.

task success, LLMSlim Platform
88.9%task success, LLMSlim Platformvs 44.4% for raw history. 180-case synthetic evaluation, deterministic oracle, Oct 2026.
fewer tokens
~33%fewer tokensMean 391.8 vs 583.0 tokens against raw context, same 180 synthetic cases.
automated tests
908automated tests680 Core, 228 Platform; over 90% branch coverage. Per release reports, 9 Oct 2026.
releases in 16 weeks
11releases in 16 weeksJune to October 2026, per PyPI and GitHub releases.
Task success on 180 synthetic cases
  • Raw history44.4%
  • Truncation33.3%
  • RAG only44.4%
  • Core (prior)50.0%
  • LLMSlim Core83.3%
  • LLMSlim Platform88.9%

LLMSlim Platform benchmark, October 2026. Synthetic cases in English, Hindi and Hinglish across 18 categories, 3,000-token budget, deterministic oracle. Most of the gain comes from temporal labelling, which this oracle rewards by design. These are not production customer outcomes; live-model results are pending.

Selected metrics, same 180 synthetic cases
MetricRaw historyCorePlatform
Stale-fact error85.7%85.7%14.3%
Wrong context29.4%28.6%5.9%
Mean tokens583429392

Long sessions: in an in-memory benchmark, Platform context levelled off near 2,000 tokens, 83% fewer than raw history at 1,000 turns. Below about 140 turns, Platform's state and trace headers cost more than short raw history; the value case is long-running agents.

In their words

What early users said.

“got 50% fewer tokens on my rag pipeline with no quality drop on my eval set, super easy drop-in replacement”

Their own result, reported on Product HuntJul 2026Source ↗ (comment by Naz on Product Hunt, opens in a new tab)

Unedited public comments on LLMSlim's first Product Hunt launch (v0.2.0, 16 July 2026). See them on Product Hunt ↗

In the product

LLMSlim Studio showing the Agent Context Runtime form with target model, objective, token budget and quality floor inputs.
LLMSlim Studio, Agent Context Runtime. Unedited capture, 10 Oct 2026.
LLMSlim benchmarks page describing evaluations and their limitations.
llmslim.app/benchmarks. Unedited capture, 10 Oct 2026.

Timeline

Compression library to context platform in sixteen weeks.

  1. 16 Jun 2026

    First release on PyPI

  2. Jul 2026

    Entity rules, rewrite and hybrid modes

  3. Aug 2026

    Provenance locks and safe tool schemas

  4. 12 Sep 2026

    MCP catalogs, new site and Studio

  5. 20 Sep 2026

    Adaptive Context Planner

  6. 30 Sep 2026

    Agent runtime and cache planning

  7. 2 Oct 2026

    Private Platform control plane

  8. 9 Oct 2026

    Context Intelligence beta

Ecosystem & status

Startup programs, not investors.

  • OpenAI for Startups
  • Claude for Startups
  • Sarvam Startup Program
  • Zoho for Startups
  • MongoDB for Startups
  • Auth0 for Startups
  • Zendesk for Startups
  • Mixpanel for Startups
  • Sentry for Startups
  • Descope Hello World
  • Pulumi for Startups

LLMSlim has been accepted into these programs, which offer cloud and API credits, software licences, technical resources and community access. None is an investor, lender, customer or funder, and membership implies no technical endorsement or commercial partnership.

Commercial status, October 2026

Core is free under the MIT licence. Platform is a private beta, by invitation. LLMSlim has no paying customers or revenue today, and no paid plan is live.

Taking on new projects

Building agents that need better context?

Ask about LLMSlim Platform's private beta and enterprise deployment, or hire AMEYOR to design the context and evaluation layer of your own AI product.

Or book a free 30-minute call (opens in a new tab) with the founder.

Prefer email? hello@ameyor.in