60–95% fewer tokens in your agent loops, same answers. Meet Headroom.
Headroom is an open-source context compression layer that cuts AI agent token usage by 60-95% while preserving accuracy, available as a proxy, library, or MCP server.

The Hidden Cost of AI Agents: Context Bloat
AI coding agents have transformed how developers debug, refactor, and navigate codebases. But there is a catch that every heavy user has felt: the token bill. Not because models are priced too high per token, but because agents consume an astonishing volume of them. A single SRE debugging session can burn over 65,000 tokens just in context — with every file read, log dump, and tool output piling into the prompt window.
Enter Headroom, a new open-source context compression layer that tackles this problem head-on. Created by Tejas Chopra, Headroom sits between your agent and the LLM, intercepting everything the agent reads — tool outputs, log files, RAG chunks, search results, even conversation history — and compresses it before the model ever sees it. The results are striking: 60–95% fewer tokens, with the same (or better) answers.
The Numbers That Matter
Headroom's developers benchmarked it across real agent workloads with consistent results:
- Code search (100 results): 17,765 → 1,408 tokens — a 92% reduction
- SRE incident debugging: 65,694 → 5,118 tokens — a 92% reduction
- GitHub issue triage: 54,174 → 14,761 tokens — a 73% reduction
- Codebase exploration: 78,502 → 41,254 tokens — a 47% reduction
Accuracy on standard benchmarks (GSM8K, TruthfulQA, SQuAD v2, and the Berkeley Function Calling Leaderboard) remained intact. In several cases, scores actually improved slightly — suggesting that cleaner, less noisy context helps the model focus on what matters.
How Headroom Works Under the Hood
Headroom doesn't use a single compression algorithm. Instead, it routes content through a stack of specialized compressors, each tuned for a specific data type:
- SmartCrusher handles JSON, nested objects, and arrays of dictionaries — common in API responses and tool outputs.
- CodeCompressor uses AST-aware compression for Python, JavaScript, Go, Rust, Java, and C++, preserving structure while eliminating redundancy.
- Kompress-base is a custom Hugging Face model trained specifically on agentic traces, optimized for prose and mixed content.
- CacheAligner stabilizes prompt prefixes so that Anthropic and OpenAI KV caches actually hit, providing an additional latency and cost benefit.
Critically, Headroom implements CCR (reversible compression). All original content is cached locally and the LLM can retrieve it on demand if needed. Nothing is destroyed — compression is lossless at the structural level, and the model can always request the full original.
Drop-In Deployment: Zero Code Changes
Headroom offers three deployment modes, each suited to different use cases:
Proxy mode is the most immediately useful. Running headroom proxy --port 8787 creates a local proxy that you point any existing tool at. No code changes, no library integration — just set your API endpoint to localhost and compression starts immediately. It works with any language and any OpenAI-compatible client.
Wrap mode is even simpler for Agent CLI users. The command headroom wrap claude automatically wraps Claude Code, routing all its traffic through Headroom. The same one-command approach works for Codex, Cursor, Aider, and Copilot CLI.
Library mode provides Python and TypeScript APIs — compress(messages) — for direct integration into applications built with LangChain, Agno, or the Vercel AI SDK. Native middleware integrations are available for these frameworks, eliminating the need for a proxy altogether.
Beyond Compression: Cross-Agent Memory
Headroom goes beyond token compression. It includes a cross-agent memory store that shares context across Claude, Codex, and Gemini sessions with automatic deduplication. Additionally, the headroom learn feature mines past failed agent sessions and writes corrections back to your CLAUDE.md or AGENTS.md files — effectively learning from mistakes over time.
For users on Opus-class models, enabling HEADROOM_OUTPUT_SHAPER=1 trims verbose model output as well, which on 5x output pricing adds up quickly.
Getting Started
Headroom is available via pip: pip install "headroom-ai[all]". From there, running headroom wrap claude takes under five minutes to see first results. The project is fully open source under a permissive license and is available on GitHub.
For developers who aren't yet burning significant tokens on agent context, the message is simple: bookmark it — you will be. As AI agents become more integrated into daily development workflows, tools like Headroom that optimize the cost and efficiency of those workflows will become essential infrastructure.


