media MCP Server
Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Discovered via github-topic:mcp and last synced 1d ago.
Install instructions not detected yet
Check the source repository for the latest setup steps.
No
`litellm.callbacks = [HeadroomCallback()]`
No
Tools
Factual
Deploy
Math
100
`SharedContext().put / .get`
grok
`HeadroomChatModel(your_llm)`
Before
`headroom wrap`
`HeadroomAgnoModel(your_model)`
Headroom
✅
✅
[Strands guide](https://docs.headroomlabs.ai/docs/strands)
Category
✅
Manual setup
Yes
opencode` in one command - **MCP server** — `headroom_compress`, `headroom_retrieve`, `headroom_stats` for any MCP client - **Cross-agent memory** — shared store across Claude, Codex, Gemini, auto-dedup - **`headroom learn`** — mines failed sessions, writes corrections to `CLAUDE.md` / `AGENTS.md` - **Output token reduction** — trims what the model *writes back* (not just what you send): drops ceremony/restated code and skips deep "thinking" on routine steps. See [Output token reduction](#output-token-reduction-cut-what-the-model-writes-back). - **Reversible (CCR)** — originals are cached for retrieval on demand ## How it works (30 seconds) ``` Your agent / app (Claude Code, Cursor, Codex, LangChain, Agno, Strands, your own code…) │ prompts · tool outputs · logs · RAG results · files ▼ ┌────────────────────────────────────────────────────┐ │ Headroom (runs locally — your data stays here) │ │ ──────────────────────────────────────────────── │ │ CacheAligner → ContentRouter → CCR │ │ ├─ SmartCrusher (JSON) │ │ ├─ CodeCompressor (AST) │ │ └─ Kompress-base (text, HF) │ │ │ │ Cross-agent memory · headroom learn · MCP │ └────────────────────────────────────────────────────┘ │ compressed prompt + retrieval tool ▼ LLM provider (Anthropic · OpenAI · Bedrock · …) ``` - **ContentRouter** — detects content type, selects the right compressor - **SmartCrusher / CodeCompressor / Kompress-base** — compress JSON, AST, or prose - **CacheAligner** — stabilizes prefixes so provider KV caches actually hit - **CCR** — stores originals locally; LLM calls `headroom_retrieve` if it needs them → [Architecture](https://headroom-docs.vercel.app/docs/architecture) · [CCR reversible compression](https://headroom-docs.vercel.app/docs/ccr) · [Kompress-v2-base model card](https://huggingface.co/chopratejas/kompress-v2-base) ## Get started (60 seconds) ```bash # 1 — Install pip install "headroom-ai[all]" # Python npm install headroom-ai # Node / TypeScript # 2 — Pick your mode headroom wrap claude # wrap a coding agent headroom proxy --port 8787 # drop-in proxy, zero code changes # or: from headroom import compress # inline library # 3 — See the savings headroom perf headroom dashboard # live savings dashboard (proxy must be running) ``` Granular extras: `[proxy]`, `[mcp]`, `[ml]`, `[code]`, `[memory]`, `[relevance]`, `[image]`, `[agno]`, `[langchain]`, `[evals]`, `[pytorch-mps]` (Apple-GPU memory-embedder offload — set `HEADROOM_EMBEDDER_RUNTIME=pytorch_mps`). Requires **Python 3.10+**. ## Proof **Savings on real agent workloads:**
After
✅
opencode
goose
vibe` in one command; undo with `headroom unwrap <tool>` - **MCP server** — `headroom_compress`, `headroom_retrieve`, `headroom_stats` for any MCP client - **Cross-agent memory** — shared store across Claude, Codex, Gemini, auto-dedup - **`headroom learn`** — mines failed sessions, writes corrections to `CLAUDE.local.md` (default, gitignored) or `CLAUDE.md` / `AGENTS.md` / `GEMINI.md` - **Output token reduction** — trims what the model *writes back* (not just what you send): drops ceremony/restated code and skips deep "thinking" on routine steps. See [Output token reduction](#output-token-reduction-cut-what-the-model-writes-back). - **Reversible (CCR)** — originals are cached for retrieval on demand ## How it works (30 seconds) ``` Your agent / app (Claude Code, Cursor, Codex, LangChain, Agno, Strands, your own code…) │ prompts · tool outputs · logs · RAG results · files ▼ ┌────────────────────────────────────────────────────┐ │ Headroom (runs locally — your data stays here) │ │ ──────────────────────────────────────────────── │ │ CacheAligner → ContentRouter → CCR │ │ ├─ SmartCrusher (JSON) │ │ ├─ CodeCompressor (AST) │ │ └─ Kompress-v2-base (text, HF) │ │ │ │ Cross-agent memory · headroom learn · MCP │ └────────────────────────────────────────────────────┘ │ compressed prompt + retrieval tool ▼ LLM provider (Anthropic · OpenAI · Bedrock · …) ``` - **ContentRouter** — detects content type, selects the right compressor - **SmartCrusher / CodeCompressor / Kompress-v2-base** — compress JSON, AST, or prose - **CacheAligner** — stabilizes prefixes so provider KV caches actually hit - **CCR** — stores originals locally; LLM calls `headroom_retrieve` if it needs them → [Architecture](https://headroom-docs.vercel.app/docs/architecture) · [CCR reversible compression](https://headroom-docs.vercel.app/docs/ccr) · [Kompress-v2-base model card](https://huggingface.co/chopratejas/kompress-v2-base) ## Get started (60 seconds) ```bash # 1 — Install pip install "headroom-ai[all]" # Python npm install headroom-ai # Node / TypeScript # 2 — Pick your mode headroom wrap claude # wrap a coding agent headroom proxy --port 8787 # drop-in proxy, zero code changes # or: from headroom import compress # inline library # 3 — Verify setup and see the savings headroom doctor # health check — confirms routing is working headroom perf headroom dashboard # live savings dashboard (proxy must be running) ``` Granular extras: `[proxy]`, `[mcp]`, `[ml]`, `[code]`, `[memory]`, `[vector]` (optional HNSW backend — needs a C++ toolchain, not in `[all]`), `[relevance]`, `[image]`, `[agno]`, `[langchain]`, `[evals]`, `[pytorch-mps]` (Apple-GPU memory-embedder offload — set `HEADROOM_EMBEDDER_RUNTIME=pytorch_mps`). Requires **Python 3.10+**. ## Proof **Savings on real agent workloads:**
✅
✅
✅
✅
aider
continue
openclaw
zcode` in one command; undo with `headroom unwrap <tool>`. - **MCP server** — `headroom_compress`, `headroom_retrieve`, `headroom_stats` for any MCP client. - **Cross-agent memory** — one shared store across Claude, Codex, Gemini and Grok, with automatic dedup. - **`headroom learn`** — mines failed sessions and writes corrections to `CLAUDE.local.md` (default, gitignored), `CLAUDE.md`, `AGENTS.md`, `GEMINI.md` or `GROK.md`. - **Output token reduction** — trims what the model *writes back*, not only what you send. See [below](#output-token-reduction). - **Reversible (CCR)** — originals are cached locally and retrieved on demand. ## How it works ``` Your agent / app (Claude Code, Cursor, Codex, LangChain, Agno, Strands, your own code…) │ prompts · tool outputs · logs · RAG results · files ▼ ┌────────────────────────────────────────────────────┐ │ Headroom (runs locally — your data stays here) │ │ ──────────────────────────────────────────────── │ │ CacheAligner → ContentRouter → CCR │ │ ├─ SmartCrusher (JSON) │ │ ├─ CodeCompressor (AST) │ │ └─ Kompress-v2-base (text, HF) │ │ │ │ Cross-agent memory · headroom learn · MCP │ └────────────────────────────────────────────────────┘ │ compressed prompt + retrieval tool ▼ LLM provider (Anthropic · OpenAI · Bedrock · …) ``` - **ContentRouter** detects the content type and selects a compressor for it. - **SmartCrusher / CodeCompressor / Kompress-v2-base** handle JSON, source code and prose respectively. - **CacheAligner** flags volatile content that would bust a provider KV-cache prefix. It never rewrites prompts. - **CCR** stores originals locally so the model can call `headroom_retrieve` when it needs the full text. → [Architecture](https://docs.headroomlabs.ai/docs/architecture) · [CCR](https://docs.headroomlabs.ai/docs/ccr) · [Kompress-v2-base model card](https://huggingface.co/chopratejas/kompress-v2-base) ## Get started (60 seconds) ```bash # 1 — Install uv tool install --python 3.13 "headroom-ai[all]" # CLI in a self-contained env pip install "headroom-ai[all]" # Python — ships the `headroom` CLI npm install headroom-ai # TypeScript SDK only — no CLI # 2 — Pick a mode headroom deploy # turnkey local deployment + agent config headroom wrap claude # wrap a coding agent headroom proxy --port 8787 # drop-in proxy, zero code changes # or: from headroom import compress # inline library # 3 — Check it and watch the savings headroom doctor # health check — confirms routing works headroom perf headroom dashboard # live savings (proxy must be running) ``` Inline, in Python: ```python from headroom import compress from openai import OpenAI messages = [{"role": "user", "content": "Analyze these results"}] result = compress(messages, model="gpt-4o") client = OpenAI() response = client.chat.completions.create(model="gpt-4o", messages=result.messages) print(f"Saved {result.tokens_saved} tokens ({result.compression_ratio:.0%})") ``` Launch a wrapped agent session each time, so the setup runs. `headroom wrap` starts a local proxy, installs **[Serena](https://github.com/oraios/serena)** for semantic code navigation, and launches the agent configured to route through Headroom. Serena is registered at user scope (for Claude Code, in `~/.claude.json`), so it stays available in your other projects until you run `headroom unwrap`. Skip it with `--code-memory none`. The `headroom` CLI ships only in the PyPI package. The npm `headroom-ai` package is the TypeScript SDK — a library you import (`import { compress } from 'headroom-ai'`) — and provides no `headroom` command. ## Proof Four scenarios built from real MCP server output formats, measured with the provider tokenizer and the shipped `compress()`. Seeded and offline, so you get the same numbers we did: ```bash uv run python benchmarks/index_proof_table.py --seed 20260902 ```
✅
The cheapest AI media API on the market. Generate images (Flux), music (AceStep), speech with voice cloning, transcribe video/audio, OCR, video generation, background removal, upscale, style transfer, and prompt enhancement — all through one unified API. Free $5 credit on signup.
Write professional press releases that get media attention and coverage
Learn how to use the deAPI AI Media Suite (Community) Claude skill. Complete guide with installation instructions and examples.
Learn how to use the Press Release Writer Claude skill. Complete guide with installation instructions and examples.