Headroom

media MCP Server

Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.

VerifiedFresh
mediamedia
35 views74,013 stars5,715 forksApache-2.0

Why This Matters

Discovered via github-topic:mcp and last synced 1d ago.

VerifiedFresh
Source
github-topic:mcp
Stars
74,013
Last synced
1d ago
Install
Check source

Install

Install instructions not detected yet

Check the source repository for the latest setup steps.

View source instructions
37
Tools
0
Resources
0
Prompts
Standard I/O
Transport

Available Tools (37)

No

No

LiteLLM

`litellm.callbacks = [HeadroomCallback()]`

Provider-native

No

BFCL

Tools

TruthfulQA

Factual

Scope

Deploy

GSM8K

Math

QA

100

Multi-agent

`SharedContext().put / .get`

codex

grok

LangChain

`HeadroomChatModel(your_llm)`

Workload

Before

Agent

`headroom wrap`

Agno

`HeadroomAgnoModel(your_model)`

Baseline

Headroom

OpenClaw

✅

Codex

✅

Strands

[Strands guide](https://docs.headroomlabs.ai/docs/strands)

Benchmark

Category

Aider

✅

Cursor

Manual setup

Yes

Yes

copilot

opencode` in one command - **MCP server** — `headroom_compress`, `headroom_retrieve`, `headroom_stats` for any MCP client - **Cross-agent memory** — shared store across Claude, Codex, Gemini, auto-dedup - **`headroom learn`** — mines failed sessions, writes corrections to `CLAUDE.md` / `AGENTS.md` - **Output token reduction** — trims what the model *writes back* (not just what you send): drops ceremony/restated code and skips deep "thinking" on routine steps. See [Output token reduction](#output-token-reduction-cut-what-the-model-writes-back). - **Reversible (CCR)** — originals are cached for retrieval on demand ## How it works (30 seconds) ``` Your agent / app (Claude Code, Cursor, Codex, LangChain, Agno, Strands, your own code…) │ prompts · tool outputs · logs · RAG results · files ▼ ┌────────────────────────────────────────────────────┐ │ Headroom (runs locally — your data stays here) │ │ ──────────────────────────────────────────────── │ │ CacheAligner → ContentRouter → CCR │ │ ├─ SmartCrusher (JSON) │ │ ├─ CodeCompressor (AST) │ │ └─ Kompress-base (text, HF) │ │ │ │ Cross-agent memory · headroom learn · MCP │ └────────────────────────────────────────────────────┘ │ compressed prompt + retrieval tool ▼ LLM provider (Anthropic · OpenAI · Bedrock · …) ``` - **ContentRouter** — detects content type, selects the right compressor - **SmartCrusher / CodeCompressor / Kompress-base** — compress JSON, AST, or prose - **CacheAligner** — stabilizes prefixes so provider KV caches actually hit - **CCR** — stores originals locally; LLM calls `headroom_retrieve` if it needs them → [Architecture](https://headroom-docs.vercel.app/docs/architecture) · [CCR reversible compression](https://headroom-docs.vercel.app/docs/ccr) · [Kompress-v2-base model card](https://huggingface.co/chopratejas/kompress-v2-base) ## Get started (60 seconds) ```bash # 1 — Install pip install "headroom-ai[all]" # Python npm install headroom-ai # Node / TypeScript # 2 — Pick your mode headroom wrap claude # wrap a coding agent headroom proxy --port 8787 # drop-in proxy, zero code changes # or: from headroom import compress # inline library # 3 — See the savings headroom perf headroom dashboard # live savings dashboard (proxy must be running) ``` Granular extras: `[proxy]`, `[mcp]`, `[ml]`, `[code]`, `[memory]`, `[relevance]`, `[image]`, `[agno]`, `[langchain]`, `[evals]`, `[pytorch-mps]` (Apple-GPU memory-embedder offload — set `HEADROOM_EMBEDDER_RUNTIME=pytorch_mps`). Requires **Python 3.10+**. ## Proof **Savings on real agent workloads:**

Before

After

OpenCode

✅

aider

opencode

continue

goose

openclaw

vibe` in one command; undo with `headroom unwrap <tool>` - **MCP server** — `headroom_compress`, `headroom_retrieve`, `headroom_stats` for any MCP client - **Cross-agent memory** — shared store across Claude, Codex, Gemini, auto-dedup - **`headroom learn`** — mines failed sessions, writes corrections to `CLAUDE.local.md` (default, gitignored) or `CLAUDE.md` / `AGENTS.md` / `GEMINI.md` - **Output token reduction** — trims what the model *writes back* (not just what you send): drops ceremony/restated code and skips deep "thinking" on routine steps. See [Output token reduction](#output-token-reduction-cut-what-the-model-writes-back). - **Reversible (CCR)** — originals are cached for retrieval on demand ## How it works (30 seconds) ``` Your agent / app (Claude Code, Cursor, Codex, LangChain, Agno, Strands, your own code…) │ prompts · tool outputs · logs · RAG results · files ▼ ┌────────────────────────────────────────────────────┐ │ Headroom (runs locally — your data stays here) │ │ ──────────────────────────────────────────────── │ │ CacheAligner → ContentRouter → CCR │ │ ├─ SmartCrusher (JSON) │ │ ├─ CodeCompressor (AST) │ │ └─ Kompress-v2-base (text, HF) │ │ │ │ Cross-agent memory · headroom learn · MCP │ └────────────────────────────────────────────────────┘ │ compressed prompt + retrieval tool ▼ LLM provider (Anthropic · OpenAI · Bedrock · …) ``` - **ContentRouter** — detects content type, selects the right compressor - **SmartCrusher / CodeCompressor / Kompress-v2-base** — compress JSON, AST, or prose - **CacheAligner** — stabilizes prefixes so provider KV caches actually hit - **CCR** — stores originals locally; LLM calls `headroom_retrieve` if it needs them → [Architecture](https://headroom-docs.vercel.app/docs/architecture) · [CCR reversible compression](https://headroom-docs.vercel.app/docs/ccr) · [Kompress-v2-base model card](https://huggingface.co/chopratejas/kompress-v2-base) ## Get started (60 seconds) ```bash # 1 — Install pip install "headroom-ai[all]" # Python npm install headroom-ai # Node / TypeScript # 2 — Pick your mode headroom wrap claude # wrap a coding agent headroom proxy --port 8787 # drop-in proxy, zero code changes # or: from headroom import compress # inline library # 3 — Verify setup and see the savings headroom doctor # health check — confirms routing is working headroom perf headroom dashboard # live savings dashboard (proxy must be running) ``` Granular extras: `[proxy]`, `[mcp]`, `[ml]`, `[code]`, `[memory]`, `[vector]` (optional HNSW backend — needs a C++ toolchain, not in `[all]`), `[relevance]`, `[image]`, `[agno]`, `[langchain]`, `[evals]`, `[pytorch-mps]` (Apple-GPU memory-embedder offload — set `HEADROOM_EMBEDDER_RUNTIME=pytorch_mps`). Requires **Python 3.10+**. ## Proof **Savings on real agent workloads:**

Cline

✅

Continue

✅

Goose

✅

OpenHands

✅

cursor

aider

cline

continue

openhands

openclaw

omp

zcode` in one command; undo with `headroom unwrap <tool>`. - **MCP server** — `headroom_compress`, `headroom_retrieve`, `headroom_stats` for any MCP client. - **Cross-agent memory** — one shared store across Claude, Codex, Gemini and Grok, with automatic dedup. - **`headroom learn`** — mines failed sessions and writes corrections to `CLAUDE.local.md` (default, gitignored), `CLAUDE.md`, `AGENTS.md`, `GEMINI.md` or `GROK.md`. - **Output token reduction** — trims what the model *writes back*, not only what you send. See [below](#output-token-reduction). - **Reversible (CCR)** — originals are cached locally and retrieved on demand. ## How it works ``` Your agent / app (Claude Code, Cursor, Codex, LangChain, Agno, Strands, your own code…) │ prompts · tool outputs · logs · RAG results · files ▼ ┌────────────────────────────────────────────────────┐ │ Headroom (runs locally — your data stays here) │ │ ──────────────────────────────────────────────── │ │ CacheAligner → ContentRouter → CCR │ │ ├─ SmartCrusher (JSON) │ │ ├─ CodeCompressor (AST) │ │ └─ Kompress-v2-base (text, HF) │ │ │ │ Cross-agent memory · headroom learn · MCP │ └────────────────────────────────────────────────────┘ │ compressed prompt + retrieval tool ▼ LLM provider (Anthropic · OpenAI · Bedrock · …) ``` - **ContentRouter** detects the content type and selects a compressor for it. - **SmartCrusher / CodeCompressor / Kompress-v2-base** handle JSON, source code and prose respectively. - **CacheAligner** flags volatile content that would bust a provider KV-cache prefix. It never rewrites prompts. - **CCR** stores originals locally so the model can call `headroom_retrieve` when it needs the full text. → [Architecture](https://docs.headroomlabs.ai/docs/architecture) · [CCR](https://docs.headroomlabs.ai/docs/ccr) · [Kompress-v2-base model card](https://huggingface.co/chopratejas/kompress-v2-base) ## Get started (60 seconds) ```bash # 1 — Install uv tool install --python 3.13 "headroom-ai[all]" # CLI in a self-contained env pip install "headroom-ai[all]" # Python — ships the `headroom` CLI npm install headroom-ai # TypeScript SDK only — no CLI # 2 — Pick a mode headroom deploy # turnkey local deployment + agent config headroom wrap claude # wrap a coding agent headroom proxy --port 8787 # drop-in proxy, zero code changes # or: from headroom import compress # inline library # 3 — Check it and watch the savings headroom doctor # health check — confirms routing works headroom perf headroom dashboard # live savings (proxy must be running) ``` Inline, in Python: ```python from headroom import compress from openai import OpenAI messages = [{"role": "user", "content": "Analyze these results"}] result = compress(messages, model="gpt-4o") client = OpenAI() response = client.chat.completions.create(model="gpt-4o", messages=result.messages) print(f"Saved {result.tokens_saved} tokens ({result.compression_ratio:.0%})") ``` Launch a wrapped agent session each time, so the setup runs. `headroom wrap` starts a local proxy, installs **[Serena](https://github.com/oraios/serena)** for semantic code navigation, and launches the agent configured to route through Headroom. Serena is registered at user scope (for Claude Code, in `~/.claude.json`), so it stays available in your other projects until you run `headroom unwrap`. Skip it with `--code-memory none`. The `headroom` CLI ships only in the PyPI package. The npm `headroom-ai` package is the TypeScript SDK — a library you import (`import { compress } from 'headroom-ai'`) — and provides no `headroom` command. ## Proof Four scenarios built from real MCP server output formats, measured with the provider tokenizer and the shipped `compress()`. Seeded and offline, so you get the same numbers we did: ```bash uv run python benchmarks/index_proof_table.py --seed 20260902 ```

ZCode

✅