whisper-windows-mcp

media MCP Server

Windows-native MCP server for local audio transcription — GPU accelerated via Vulkan, works with Claude Desktop

VerifiedInstall Ready
mediamedia
9 views0 stars0 forksv2.3.0NOASSERTION

Why This Matters

Discovered via github-topic:model-context-protocol and last synced 3mo ago.

VerifiedInstall Ready
Source
github-topic:model-context-protocol
Stars
0
Last synced
3mo ago
Install
Instructions detected

Install

1. Install the package

npx whisper-windows-mcp

2. Add to claude_desktop_config.json

{
  "mcpServers": {
    "whisper-windows-mcp": {
      "command": "npx",
      "args": [
        "whisper-windows-mcp"
      ]
    }
  }
}

Config file location: ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) / %APPDATA%\Claude\claude_desktop_config.json (Windows)

50
Tools
0
Resources
0
Prompts
Standard I/O
Transport

Available Tools (50)

Accuracy

Best for

Model

Size

output_format

`text` (default), `timestamps`, `json`, or `srt`

Basic

Quick tests

background

Run as detached job — returns a job ID immediately. Use `check_progress` to monitor. Recommended for files over 10 minutes.

batch_id

Batch ID returned by `start_batch`

translate_to_english

Also generate an English translation `.en.srt`. Only applies when source is not English.

WHISPER_MODEL

Path to model .bin file (required)

Fast

Excellent

Moderate

Better

file_path

Path to file (required)

Variable

Description

threads

CPU thread override

WHISPER_THREADS

CPU thread count override

sort_by

For folders: `duration` (default), `name`, or `size`

Hardware

Time

temperature

Sampling temperature 0.0–1.0. Default 0.0 (deterministic). Higher values reduce hallucination on noisy audio.

diarize

Stereo speaker diarization — requires stereo audio with speakers on separate channels.

WHISPER_PRIVACY_MODE

**Planned.** When set to `true`, tool responses return metadata only — no transcript text is returned to Claude's API. For regulated or confidential content. See [PRIVACY.md](PRIVACY.md).

prompt

Prior context string — improves accuracy for domain-specific vocabulary or speaker names. Example: `"Names: Keemstar, DramaAlert."`

vad_model

Path to Silero VAD model .bin. Strips silence before transcription — reduces hallucinations on noisy files.

save_to_file

Save transcript as .txt next to the source file

processors

Parallel processor count. Default 1.

gpu_device

GPU device index for multi-GPU systems. Default 0.

job_id

Job ID returned by `transcribe_audio`

file_index

Which file to process (1-based). Omit to list files first.

path

Path to a single file or folder (required)

WHISPER_CLI_PATH

Path to whisper-cli.exe (required)

language

Language code or `auto` to detect. Default: `en`

best_of

Candidate sequences evaluated. Default 5.

folder_path

Path to folder (required)

beam_size

Beam search width. Higher = more accurate, slower. Default 5.

duration

Process duration in milliseconds from offset.

Excellent

Multilingual, CPU-friendly

word_timestamps

One word per timestamped segment. Useful for clip alignment.

recursive

Include subfolders

condition_on_prev_text

Re-enable context conditioning between segments. Default false.

offset_t

Start offset in milliseconds.

Type

Formats

max_segment_length

Max segment length in characters.

Parameter

Description

model_name

Model filename (e.g. `ggml-large-v3-turbo.bin`) or full path. Must be a `.bin` file in the configured models directory.

FFMPEG_PATH

Path to ffmpeg if not in system PATH

text

— plain text, no time codes

json

— structured JSON (blocking mode only)

srt

— SubRip subtitle file saved next to source

lrc

— LRC lyrics/karaoke format saved next to source

vtt

— WebVTT subtitle file saved next to source

timestamps

— timestamped segments, e.g. `[00:00:01.230 --> 00:00:04.560] Hello world` (default)

csv

— CSV with timestamps saved next to source