media MCP Server
MCP server that prevents LLM vision downscaling, tiles large images & screenshots so Claude, GPT-4o, and Gemini see every detail
Discovered via unknown and last synced 3mo ago.
1. Install the package
npx -y image-tiler-mcp-server
2. Add to claude_desktop_config.json
{
"mcpServers": {
"image-tiler-mcp-server": {
"command": "npx",
"args": [
"image-tiler-mcp-server"
]
}
}
}Config file location: ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) / %APPDATA%\Claude\claude_desktop_config.json (Windows)
string
number
string
Format
Default tile
number
string
Description
Tile page to return (0 = first 5, 1 = next 5, etc.)
`url` (required), `focus` (optional)
`claude`
Without tiling
string
string
Auto-downscaled to ~1,568 x 827 (~3.7% of pixels survive)
Downscaled to ~1,456 x 768 (~3% of pixels survive)
Type
string
number
1120
string
boolean
Analyze each tile using image stats and return content classification (blank, low-detail, mixed, high-detail) plus `stdDev` and `entropy` values per tile
string
Arguments
258
string
string
number
number
number
string
number
boolean
Max dimension in px (0 to disable, or 256-65536). Values 1-255 are silently clamped to 256. Pre-downscales the image so its longest side fits within this value before tiling.
string
Example prompt
string
Whether to emulate a mobile device. When true, defaults `viewportWidth` to 390, `deviceScaleFactor` to 2, and sets a mobile user agent.
number
`filePath` (required), `preset` (optional), `focus` (optional)
All supported vision model presets with tile sizes, min/max bounds, and per-tile token rates.
Additional delay in ms after page load (max 30000)
> **OpenAI note:** The `openai` config targets the GPT-4o / o-series vision pipeline (512px tile patches). GPT-4.1 uses a fundamentally different pipeline (32x32 pixel patches) and is not currently supported. It would require a separate model config with a different calculation approach. > **Gemini 3 note:** Gemini 3 uses a fixed token budget per image (1,120 tokens regardless of dimensions). Tiling increases total token cost but preserves fine detail. For cases where detail isn't critical, consider sending a single image instead. ## Why Tile? You screenshot a full page, paste it into Claude, and Claude **crushes it to a thumbnail**. Any image with a long edge over 1,568 pixels gets auto-downscaled to fit within ~1.15 megapixels. A 3,600 x 20,220px full-page capture becomes ~279 x 1,568, losing over 99% of its pixels before the model even sees it. GPT-4o is more forgiving but still destructive: it scales your image to fit within 2,048px, then scales the shortest side down to 768px, *then* tiles internally. An 8,192px-wide NASA panorama becomes ~1,456 x 768 before GPT-4o's own tiling even begins. Gemini 1.5/2.0 handles large images natively at 768px tiles without downscaling. Gemini 3, however, caps each image at a fixed token budget (1,120 tokens) regardless of size. Tiling gives each piece its own budget. Each tile stays within the model's sweet spot, so the LLM processes it at full resolution. ### What Happens Without Tiling Using `assets/portrait.png` (3,600 x 20,220, a full-page National Geographic capture) as an example:
boolean
The cheapest AI media API on the market. Generate images (Flux), music (AceStep), speech with voice cloning, transcribe video/audio, OCR, video generation, background removal, upscale, style transfer, and prompt enhancement — all through one unified API. Free $5 credit on signup.
Write professional press releases that get media attention and coverage
Learn how to use the deAPI AI Media Suite (Community) Claude skill. Complete guide with installation instructions and examples.
Learn how to use the Press Release Writer Claude skill. Complete guide with installation instructions and examples.