transdex-mcp
MCP server for whisper-based transcription and translation, supporting local stdio and remote HTTP transports with file workflow safety.
README
<h1 align="center">transdex-mcp</h1> <h3 align="center">Whisper-first transcription and translation tools over MCP</h3>
<p align="center"> Transdex exposes local stdio and remote Streamable HTTP transports from one codebase. MCP clients can plan safe file workflows, run whisper.cpp transcription, preserve subtitle timing, and retrieve bounded result files. </p>
<p align="center"> <img alt="Node.js 20 or newer" src="https://img.shields.io/badge/Node.js-%3E%3D20-339933"> <img alt="MCP stdio and Streamable HTTP" src="https://img.shields.io/badge/MCP-stdio_%2B_HTTP-5A67D8"> <img alt="Project status pre-release" src="https://img.shields.io/badge/status-pre--release-orange"> </p>
Contents
- What Is This?
- Client Modes
- Architecture
- Remote Tool Contract
- Quick Start
- Connect an MCP Client
- Configuration
- Security and Privacy
- Local Workflow
- Project Structure
- Testing
- Troubleshooting
- Roadmap
- License
What Is This?
transdex-mcp is a standalone MCP codebase for transcription and translation workflows. It has its own Git history, package identity, configuration path, model cache, server entry points, and deployment policy.
The default transcription profile is:
| Layer | Default |
|---|---|
| ASR provider | whisper.cpp |
| Model | large-v3-turbo-q5_0 |
| Language | Automatic detection |
| Output | Transcript plus optional SRT |
| Execution | Bounded asynchronous queue |
| Local transport | MCP over stdio |
| Remote transport | MCP Streamable HTTP |
This repository is a pre-release implementation. It is not ready for public directory submission or untrusted multi-instance deployment.
Client Modes
MCP is the product boundary. Client-specific behavior is kept at the connection edge.
| Client | Recommended connection | Primary input style |
|---|---|---|
| Codex CLI, app, or IDE | Local stdio for workspace files; Streamable HTTP for a deployed service | MCP roots and local paths, or remote tool arguments |
| ChatGPT | Streamable HTTP | Top-level remote file references supplied by the host |
| Other MCP clients | stdio or Streamable HTTP | Depends on client capabilities |
Codex and ChatGPT both support MCP servers. The repository does not require a custom widget; clients can call the tools directly. See OpenAI's MCP documentation for current client configuration behavior.
Architecture
MCP client
|
+-- Local stdio
| +-- Workspace roots
| +-- Plan / preview / confirm
| +-- Translation and transcription jobs
|
+-- Remote Streamable HTTP
+-- Validate and stage a remote media reference
+-- Enforce byte, duration, queue, and session limits
+-- Run whisper.cpp asynchronously
+-- Return opaque signed result links
Both transports use the same validated JobSpec, safe sidecar writer,
transcription providers, subtitle preservation rules, and cancellation path.
Remote Tool Contract
| Tool | Effect |
|---|---|
| start_transcription | Stages one remote file and enqueues Whisper inference |
| get_transcription_job | Reads status and bounded progress |
| get_transcription_result | Promotes completed outputs and returns resource links |
| cancel_transcription_job | Stops queued or running inference |
Compatible OpenAI hosts can inject a top-level file parameter declared through
openai/fileParams:
{
"download_url": "https://temporary.example/media",
"file_id": "file_example",
"mime_type": "video/mp4",
"file_name": "recording.mp4"
}
The server downloads the bytes during start_transcription; it never stores a
temporary URL for a background worker. Nested file parameters are intentionally
unsupported.
Quick Start
1. Install Node dependencies
npm ci
Node.js 20 or newer is required.
2. Build whisper.cpp
git clone https://github.com/ggml-org/whisper.cpp.git "$HOME/.local/share/whisper.cpp"
cmake -S "$HOME/.local/share/whisper.cpp" \
-B "$HOME/.local/share/whisper.cpp/build" \
-DCMAKE_BUILD_TYPE=Release
cmake --build "$HOME/.local/share/whisper.cpp/build" -j 4
Pin a reviewed whisper.cpp revision for deployment rather than building an unpinned branch.
3. Provision models
cd "$HOME/.local/share/whisper.cpp"
bash ./models/download-ggml-model.sh large-v3-turbo-q5_0
bash ./models/download-vad-model.sh silero-v6.2.0
Model weights and the whisper.cpp binary are not bundled. Remote deployments should provision them while building the service image and keep request-time auto-download disabled.
4. Configure the runtime
export TRANSDEX_WHISPER_CPP_BIN="$HOME/.local/share/whisper.cpp/build/bin/whisper-cli"
export TRANSDEX_WHISPER_CPP_MODEL_DIR="$HOME/.local/share/whisper.cpp/models"
export TRANSDEX_WHISPER_CPP_MODEL=large-v3-turbo-q5_0
export TRANSDEX_WHISPER_CPP_VAD=1
export TRANSDEX_WHISPER_CPP_VAD_MODEL="$HOME/.local/share/whisper.cpp/models/ggml-silero-v6.2.0.bin"
export TRANSDEX_WHISPER_CPP_AUTO_DOWNLOAD=0
Use .env.example as the remote deployment checklist. The application does not load dotenv files automatically.
5. Start a transport
Local stdio:
npm run mcp
Local HTTP development server:
npm start
The default development endpoints are:
http://127.0.0.1:8787/mcp
http://127.0.0.1:8787/healthz
Connect an MCP Client
Codex with local stdio
Use an absolute checkout path:
codex mcp add transdex -- node /absolute/path/to/transdex-mcp/src/mcp/main.js
codex mcp list
The checked-in .mcp.json contains the equivalent repository-local definition for clients that read project MCP configuration.
Codex with Streamable HTTP
Add the deployed endpoint to ~/.codex/config.toml:
[mcp_servers.transdex]
url = "https://transdex.example.com/mcp"
bearer_token_env_var = "TRANSDEX_REMOTE_TOKEN"
tool_timeout_sec = 1800
Other remote MCP clients
Expose /mcp through HTTPS, set TRANSDEX_PUBLIC_BASE_URL and
TRANSDEX_HTTP_ALLOWED_HOSTS, and configure one authentication mode. A static
bearer token is only appropriate for one trusted tenant. Multi-user deployments
must validate OAuth at a trusted reverse proxy and forward a stable principal
header after stripping client-supplied copies.
Configuration
HTTP and session settings
| Variable | Default | Purpose |
|---|---|---|
| TRANSDEX_HTTP_HOST | 127.0.0.1 | Bind address |
| TRANSDEX_HTTP_PORT | 8787 | HTTP port |
| TRANSDEX_PUBLIC_BASE_URL | Local URL after listen | Origin used for artifact links |
| TRANSDEX_HTTP_ALLOWED_HOSTS | Empty | Required accepted Host values for a public endpoint |
| TRANSDEX_HTTP_BEARER_TOKEN | Empty | Static single-tenant staging token |
| TRANSDEX_SINGLE_TENANT | 0 | Required acknowledgement for static bearer mode |
| TRANSDEX_TRUSTED_AUTH_PROXY | 0 | Enable trusted proxy principal binding |
| TRANSDEX_AUTH_PRINCIPAL_HEADER | Empty | Stable principal header set by the trusted proxy |
| TRANSDEX_ALLOW_UNAUTHENTICATED | 0 | Unsafe isolated-development override |
| TRANSDEX_PUBLIC_WORKSPACE_ROOT | Random private local root; required remotely | Real service-owned root with mode 0700 |
| TRANSDEX_MAX_UPLOAD_BYTES | 536870912 | Maximum staged upload size |
| TRANSDEX_MAX_SESSION_BYTES | 1073741824 | Aggregate staged bytes per session |
| TRANSDEX_MAX_SESSION_STAGING_QUEUE | 2 | Pending staging requests per session |
| TRANSDEX_FILE_DOWNLOAD_TIMEOUT_MS | 60000 | Whole remote download deadline |
| TRANSDEX_MAX_MEDIA_DURATION_SECONDS | 14400 | Maximum inspected media duration |
| TRANSDEX_MAX_DECODED_AUDIO_BYTES | Derived from duration | Maximum 16 kHz mono PCM expansion |
| TRANSDEX_MAX_ACTIVE_JOBS | 1 | Concurrent Whisper jobs per session |
| TRANSDEX_MAX_GLOBAL_JOBS | 1 | Concurrent Whisper jobs for the service |
| TRANSDEX_MAX_GLOBAL_QUEUED_JOBS | 8 | Jobs waiting for a global Whisper slot |
| TRANSDEX_MAX_GLOBAL_STAGING | 2 | Concurrent remote downloads |
| TRANSDEX_MAX_GLOBAL_STAGING_QUEUE | 8 | Downloads waiting for a staging slot |
| TRANSDEX_MAX_SESSIONS | 8 | Concurrent MCP session cap |
| TRANSDEX_MAX_SESSIONS_PER_PRINCIPAL | 2 | Session cap for one trusted principal |
| TRANSDEX_DOWNLOAD_TTL_MS | 900000 | Signed artifact URL lifetime |
| TRANSDEX_DOWNLOAD_SECRET | Random per process | HMAC secret; at least 32 bytes when configured |
| TRANSDEX_FILE_HOSTS | Empty | Required exact or wildcard remote file host allowlist |
| TRANSDEX_ALLOW_ANY_PUBLIC_FILE_HOST | 0 | Unsafe host-allowlist override |
Whisper settings
| Variable | Default | Purpose |
|---|---|---|
| TRANSDEX_WHISPER_CPP_MODEL | large-v3-turbo-q5_0 | Default model |
| TRANSDEX_PUBLIC_WHISPER_MODELS | large-v3-turbo-q5_0 | Models exposed to remote callers |
| TRANSDEX_WHISPER_CPP_MODEL_DIR | User cache | Pre-provisioned GGML directory |
| TRANSDEX_WHISPER_CPP_BIN | whisper-cli on PATH | Binary path |
| TRANSDEX_WHISPER_CPP_LANGUAGE | auto | Whisper language code |
| TRANSDEX_WHISPER_CPP_THREADS | Up to 4 | Worker threads |
| TRANSDEX_WHISPER_CPP_PROCESSORS | 1 | Parallel processors |
| TRANSDEX_WHISPER_CPP_ACCELERATION | auto | auto, cpu, or gpu |
| TRANSDEX_WHISPER_CPP_DEVICE | Empty | Optional GPU device number |
| TRANSDEX_WHISPER_CPP_VAD | Off unless model is set | Enable Silero VAD |
| TRANSDEX_WHISPER_CPP_VAD_MODEL | Empty | VAD GGML path |
| TRANSDEX_WHISPER_CPP_TIMEOUT_MS | 1800000 | Inference timeout |
| TRANSDEX_WHISPER_CPP_AUTO_DOWNLOAD | 0 remotely | Model auto-download switch |
Remote callers can select only models in TRANSDEX_PUBLIC_WHISPER_MODELS and
cannot provide arbitrary local paths. Media ceilings and the auto-download
switch are parsed strictly at startup.
Security and Privacy
Implemented safeguards include:
- HTTPS-only remote media URLs with hostname allowlisting
- DNS rejection for private, loopback, link-local, and reserved destinations
- DNS-pinned HTTPS connections and redirect revalidation
- Whole-download deadlines, streamed byte limits, and atomic staging
- ffprobe duration checks plus ffmpeg decoded-PCM limits
- Local-only ffmpeg/ffprobe protocol and demuxer allowlists
- Minimal child-process environments that omit application secrets
- Private per-session workspaces with ownership and symlink checks
- Bounded per-session and service-wide queues
- Per-principal session and artifact quotas in trusted-proxy mode
- Stable public errors without server-local paths
- Opaque signed artifact links
- Process-group cancellation and graceful shutdown
Current limitations:
- Jobs and artifact metadata are in memory and local to one process.
- Static bearer authentication is single-tenant, not public user OAuth.
- Trusted principal mode depends on a correctly configured reverse proxy.
- The reverse proxy must enforce request and bandwidth rate limits.
- ffmpeg, ffprobe, and whisper.cpp still require a low-privilege, no-network worker boundary with OS or container resource limits.
- Remote translation is not yet adapted to the direct remote-file flow.
Do not expose this pre-release directly to untrusted users or deploy multiple replicas until worker isolation, edge rate limiting, durable shared storage, and authenticated tenant ownership are in place.
Local Workflow
The local stdio server retains the plan-preview-confirm workflow for text, Markdown, subtitle, directory, and media jobs.
npm run mcp
The compatibility CLI remains available for direct terminal use:
npm run cli -- --help
Project Structure
src/
|-- mcp-http/ # Remote Streamable HTTP tools and artifact service
|-- mcp/ # Local stdio planning and job workflow
|-- providers/ # whisper.cpp and optional cloud providers
|-- jobs/ # JobSpec runner and safe artifact writing
|-- media/ # ffmpeg preparation and segment merging
|-- safety/ # Path and exclusive-write guards
+-- workers/ # Translation and subtitle preservation
test/
|-- mcp-http.test.js
|-- mcp-http-server.test.js
|-- remote-file-staging.test.js
|-- mcp.test.js
+-- whisper-public-profile.test.js
Testing
Run the complete test suite:
npm test
The suite mocks whisper-cli. Before deployment, run a real smoke test with
the exact pinned whisper.cpp binary, model, VAD file, ffmpeg version, and target
server architecture.
Troubleshooting
whisper.cpp model is missing
Pre-provision the selected model in TRANSDEX_WHISPER_CPP_MODEL_DIR. Keep
request-time auto-download disabled for remote deployments.
VAD is enabled but no model is configured
Set TRANSDEX_WHISPER_CPP_VAD_MODEL to the downloaded Silero model or disable
VAD.
A remote file is rejected
Check that download_url uses HTTPS, resolves only to public addresses, has a
supported media extension, and matches TRANSDEX_FILE_HOSTS.
An MCP client cannot connect
For stdio, verify the absolute Node and repository paths. For HTTP, verify the
HTTPS /mcp URL, reverse proxy streaming, Host allowlist, and authentication
policy. In Codex, use /mcp or codex mcp list to inspect configured servers.
A job remains queued
Inspect both the per-session queue and the service-wide gate. Increase global concurrency only after benchmarking CPU, memory, disk expansion, and cancellation behavior.
Roadmap
- Adapt host-assisted subtitle translation to the remote file flow
- Add an optional MCP progress and download widget
- Replace in-memory jobs with a durable queue
- Move artifacts to tenant-scoped object storage
- Add native OAuth metadata and durable per-user accounting
- Add model and binary digest pinning
- Add a real whisper.cpp integration smoke test
- Add container and deployment manifests
- Complete privacy-policy and directory-submission review
License
This repository currently has no chosen redistribution license. package.json
is marked private and UNLICENSED. Do not publish the npm package or redistribute
the repository until the owner chooses a license and reviews third-party
notices.
whisper.cpp, Whisper model weights, ffmpeg, and optional cloud providers have their own licenses and distribution requirements.
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。