image-tiler-mcp-server

media MCP Server

MCP server that prevents LLM vision downscaling, tiles large images & screenshots so Claude, GPT-4o, and Gemini see every detail

Install Ready
mediamedia
6 views1 stars1 forksv3.1.2MIT

Why This Matters

Discovered via unknown and last synced 3mo ago.

Install Ready
Source
unknown
Stars
1
Last synced
3mo ago
Install
Instructions detected

Install

1. Install the package

npx -y image-tiler-mcp-server

2. Add to claude_desktop_config.json

{
  "mcpServers": {
    "image-tiler-mcp-server": {
      "command": "npx",
      "args": [
        "image-tiler-mcp-server"
      ]
    }
  }
}

Config file location: ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) / %APPDATA%\Claude\claude_desktop_config.json (Windows)

45
Tools
0
Resources
0
Prompts
Standard I/O
Transport

Available Tools (45)

preset

string

tileSize

number

format

string

URI

Format

Preset

Default tile

deviceScaleFactor

number

tilesDir

string

Default

Description

0

Tile page to return (0 = first 5, 1 = next 5, etc.)

capture-and-analyze

`url` (required), `focus` (optional)

1568px

`claude`

Model

Without tiling

url

string

waitUntil

string

Claude

Auto-downscaled to ~1,568 x 827 (~3.7% of pixels survive)

GPT-4o

Downscaled to ~1,456 x 768 (~3% of pixels survive)

Parameter

Type

userAgent

string

start

number

1536px

1120

dataUrl

string

mobile

boolean

true

Analyze each tile using image stats and return content classification (blank, low-detail, mixed, high-detail) plus `stdDev` and `entropy` values per tile

outputDir

string

Prompt

Arguments

768px

258

filePath

string

screenshotPath

string

delay

number

end

number

maxDimension

number

sourceUrl

string

viewportWidth

number

skipBlankTiles

boolean

10000

Max dimension in px (0 to disable, or 256-65536). Values 1-255 are silently clamped to 256. Pre-downscales the image so its longest side fits within this value before tiling.

model

string

What

Example prompt

imageBase64

string

false

Whether to emulate a mobile device. When true, defaults `viewportWidth` to 390, `deviceScaleFactor` to 2, and sets a mobile user agent.

page

number

tile-and-analyze

`filePath` (required), `preset` (optional), `focus` (optional)

JSON

All supported vision model presets with tile sizes, min/max bounds, and per-tile token rates.

3000

Additional delay in ms after page load (max 30000)

gemini3

> **OpenAI note:** The `openai` config targets the GPT-4o / o-series vision pipeline (512px tile patches). GPT-4.1 uses a fundamentally different pipeline (32x32 pixel patches) and is not currently supported. It would require a separate model config with a different calculation approach. > **Gemini 3 note:** Gemini 3 uses a fixed token budget per image (1,120 tokens regardless of dimensions). Tiling increases total token cost but preserves fine detail. For cases where detail isn't critical, consider sending a single image instead. ## Why Tile? You screenshot a full page, paste it into Claude, and Claude **crushes it to a thumbnail**. Any image with a long edge over 1,568 pixels gets auto-downscaled to fit within ~1.15 megapixels. A 3,600 x 20,220px full-page capture becomes ~279 x 1,568, losing over 99% of its pixels before the model even sees it. GPT-4o is more forgiving but still destructive: it scales your image to fit within 2,048px, then scales the shortest side down to 768px, *then* tiles internally. An 8,192px-wide NASA panorama becomes ~1,456 x 768 before GPT-4o's own tiling even begins. Gemini 1.5/2.0 handles large images natively at 768px tiles without downscaling. Gemini 3, however, caps each image at a fixed token budget (1,120 tokens) regardless of size. Tiling gives each piece its own budget. Each tile stays within the model's sweet spot, so the LLM processes it at full resolution. ### What Happens Without Tiling Using `assets/portrait.png` (3,600 x 20,220, a full-page National Geographic capture) as an example:

includeMetadata

boolean