media MCP Server
Windows-native MCP server for local audio transcription — GPU accelerated via Vulkan, works with Claude Desktop
Discovered via github-topic:model-context-protocol and last synced 3mo ago.
1. Install the package
npx whisper-windows-mcp
2. Add to claude_desktop_config.json
{
"mcpServers": {
"whisper-windows-mcp": {
"command": "npx",
"args": [
"whisper-windows-mcp"
]
}
}
}Config file location: ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) / %APPDATA%\Claude\claude_desktop_config.json (Windows)
Best for
Size
`text` (default), `timestamps`, `json`, or `srt`
Quick tests
Run as detached job — returns a job ID immediately. Use `check_progress` to monitor. Recommended for files over 10 minutes.
Batch ID returned by `start_batch`
Also generate an English translation `.en.srt`. Only applies when source is not English.
Path to model .bin file (required)
Excellent
Better
Path to file (required)
Description
CPU thread override
CPU thread count override
For folders: `duration` (default), `name`, or `size`
Time
Sampling temperature 0.0–1.0. Default 0.0 (deterministic). Higher values reduce hallucination on noisy audio.
Stereo speaker diarization — requires stereo audio with speakers on separate channels.
**Planned.** When set to `true`, tool responses return metadata only — no transcript text is returned to Claude's API. For regulated or confidential content. See [PRIVACY.md](PRIVACY.md).
Prior context string — improves accuracy for domain-specific vocabulary or speaker names. Example: `"Names: Keemstar, DramaAlert."`
Path to Silero VAD model .bin. Strips silence before transcription — reduces hallucinations on noisy files.
Save transcript as .txt next to the source file
Parallel processor count. Default 1.
GPU device index for multi-GPU systems. Default 0.
Job ID returned by `transcribe_audio`
Which file to process (1-based). Omit to list files first.
Path to a single file or folder (required)
Path to whisper-cli.exe (required)
Language code or `auto` to detect. Default: `en`
Candidate sequences evaluated. Default 5.
Path to folder (required)
Beam search width. Higher = more accurate, slower. Default 5.
Process duration in milliseconds from offset.
Multilingual, CPU-friendly
One word per timestamped segment. Useful for clip alignment.
Include subfolders
Re-enable context conditioning between segments. Default false.
Start offset in milliseconds.
Formats
Max segment length in characters.
Description
Model filename (e.g. `ggml-large-v3-turbo.bin`) or full path. Must be a `.bin` file in the configured models directory.
Path to ffmpeg if not in system PATH
— plain text, no time codes
— structured JSON (blocking mode only)
— SubRip subtitle file saved next to source
— LRC lyrics/karaoke format saved next to source
— WebVTT subtitle file saved next to source
— timestamped segments, e.g. `[00:00:01.230 --> 00:00:04.560] Hello world` (default)
— CSV with timestamps saved next to source
The cheapest AI media API on the market. Generate images (Flux), music (AceStep), speech with voice cloning, transcribe video/audio, OCR, video generation, background removal, upscale, style transfer, and prompt enhancement — all through one unified API. Free $5 credit on signup.
Write professional press releases that get media attention and coverage
Learn how to use the deAPI AI Media Suite (Community) Claude skill. Complete guide with installation instructions and examples.
Learn how to use the Press Release Writer Claude skill. Complete guide with installation instructions and examples.