media MCP Server
Transcribe audio into text. Agentic AI supported through MCP Server.
Discovered via github-topic:mcp-server and last synced 3mo ago.
1. Install the package
uvx --from audio-transcriber audio-transcriber-mcp
2. Add to claude_desktop_config.json
{
"mcpServers": {
"audio-transcriber": {
"command": "npx",
"args": [
"audio-transcriber"
]
}
}
}Config file location: ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) / %APPDATA%\Claude\claude_desktop_config.json (Windows)
Path to the Eunomia security guardrail policies JSON file.
~1 GB
Transcribes audio from a provided file or by recording from the microphone.
Default / Example
~5 GB
Parameters
~10 GB
MCP                   *Version: 0.14.0* ## Overview Transcribe your .wav .mp4 .mp3 .flac files to text or record your own audio! This repository is actively maintained - Contributions are welcome! Contribution Opportunities: - Support new models Wrapped around [OpenAI Whisper](https://pypi.org/project/openai-whisper) ## MCP ## MCP Tools
~1 GB
~2 GB
Boolean flag for enabling internal audio processing tools.
Agent                   *Version: 0.33.0* > **Documentation** — Installation, deployment, and usage across the CLI, Python API, > MCP server, and A2A agent are maintained in the > [official documentation](https://knuckles-team.github.io/audio-transcriber/). --- ## Overview **Audio Transcriber** is a production-grade Agent and Model Context Protocol (MCP) server designed to interface directly with Transcribe your .wav .mp4 .mp3 .flac files to text or record your own audio!. --- ## Key Features - **Consolidated Action-Routed MCP Tools:** Minimizes token overhead and eliminates tool bloat in LLM contexts by grouping methods into optimized, togglable tool modules. - **Enterprise-Grade Security:** Comprehensive support for Eunomia policies, OIDC token delegation, and granular execution context tracking. - **Integrated Graph Agent:** Built-in Pydantic AI agent supporting the Agent Control Protocol (ACP) and standard Web interfaces (AG-UI). - **Native Telemetry & Tracing:** Out-of-the-box OpenTelemetry exports and native Langfuse tracing. --- ## CLI or API This agent wraps the Transcribe your .wav .mp4 .mp3 .flac files to text or record your own audio! API. You can interact with it programmatically or via its integrated execution entrypoints. Detailed instructions on how to use the underlying API wrappers, extended schema bindings, and developer SDK references are maintained in [docs/index.md](docs/index.md). --- ## MCP This server utilizes dynamic Action-Routed tools to optimize token overhead and maximize IDE compatibility. ### Available MCP Tools
Functionality
Eunomia guardrail deployment type (e.g., `none`, `embedded`, `remote`).
Contents
Toggle the audio processing tool module.
Standard OpenAI Whisper model to use for local transcription (e.g., `base`, `tiny`, `small`).
Security authentication type to apply (e.g., `jwt`, `none`).
`True`
OpenTelemetry collector endpoint for exporting traces.
The cheapest AI media API on the market. Generate images (Flux), music (AceStep), speech with voice cloning, transcribe video/audio, OCR, video generation, background removal, upscale, style transfer, and prompt enhancement — all through one unified API. Free $5 credit on signup.
Write professional press releases that get media attention and coverage
Learn how to use the deAPI AI Media Suite (Community) Claude skill. Complete guide with installation instructions and examples.
Learn how to use the Press Release Writer Claude skill. Complete guide with installation instructions and examples.