transcriber-mcp

transcriber-mcp

Enables transcription and speaker diarization of audio files, interviews, and YouTube URLs, producing speaker-attributed transcripts with timestamps. Supports multiple backends (local Whisper, OpenAI API) and output formats (txt, vtt, srt, json).

Category
访问服务器

README

dialogue-transcriber

Transcribe conversations and find out who said what.

CI PyPI License: Apache-2.0 Python

Point it at an interview, panel discussion, meeting recording, or YouTube URL and get back a transcript where every line is attributed to a speaker — plus a web UI to inspect the speaker clusters, listen to any segment, and fix labels by hand.

The review UI: speaker clusters, waveform, timeline, and searchable transcript

How it works

audio  ──►  transcribe  ──►  segment  ──►  extract_clips  ──►  embed  ──►  cluster
              (Whisper)        (sentence-       (ffmpeg)       (TitaNet)    (UMAP +
                               level)                                       KMeans +
                                                                            silhouette)

Whisper produces word-level timestamps; words are grouped into sentence segments; each segment's audio is embedded with NVIDIA NeMo TitaNet; the embeddings are clustered on a UMAP projection; and the transcript comes out labeled Speaker 1, Speaker 2, … Every stage is cached on content hash, so re-runs and config tweaks are cheap.

Quickstart

ffmpeg and ffprobe must be on PATH (brew install ffmpeg on macOS).

# No install needed:
uvx --from "dialogue-transcriber[all]" transcriber transcribe interview.mp3

# Or install the tool:
uv tool install "dialogue-transcriber[all]"

transcriber transcribe interview.mp3 --participants 2
transcriber transcribe "https://www.youtube.com/watch?v=..." --backend openai
transcriber serve interview.mp3        # review UI on http://127.0.0.1:8000

The default backend runs faster-whisper locally; --backend openai uses the OpenAI Whisper API instead (requires OPENAI_API_KEY, much faster on machines without a GPU). The key can be exported in the environment or kept in a .env file in your project — the CLI loads .env from the working directory (or nearest parent), and exported variables always take precedence over the file.

Where does data go?

  • Pipeline cache: ./.transcriber-cache/ in the directory you run from (override with --work-dir) — chunks, per-segment clips, embeddings, YouTube downloads, and the web UI's job state. Safe to delete; it will be rebuilt.
  • Transcripts: written next to the input audio (interview.txt), or wherever --output points; --output - prints to stdout.
  • Model weights (local backend): downloaded once into ~/.cache (Hugging Face / NeMo). The Whisper large-v3 download is ~3 GB, so the first local run takes a while.

Nothing leaves your machine with the default local backend; --backend openai sends audio to the OpenAI API.

Picking your extras

[all] is the easy button. For smaller installs:

uv pip install dialogue-transcriber              # core only
uv pip install "dialogue-transcriber[local]"     # + faster-whisper backend
uv pip install "dialogue-transcriber[openai]"    # + OpenAI Whisper API backend
uv pip install "dialogue-transcriber[cluster]"   # + scikit-learn / UMAP
uv pip install "dialogue-transcriber[embed]"     # + NeMo TitaNet speaker embedder
uv pip install "dialogue-transcriber[api]"       # + FastAPI backend (powers the web UI)
uv pip install "dialogue-transcriber[youtube]"   # + yt-dlp downloader
uv pip install "dialogue-transcriber[oip]"       # + MCP server for OIP consumers

CLI

# Full pipeline; writes a speaker-labeled transcript next to the audio
transcriber transcribe path/to/audio.mp3

# Speakers, language, format
transcriber transcribe interview.mp3 --participants 3 --language sv --format vtt

# Machine-readable output on stdout (see "For AI agents" below)
transcriber transcribe interview.mp3 --format json --output -

# Pull audio from YouTube
transcriber download "https://www.youtube.com/watch?v=..."

# Pipeline + web UI
transcriber serve interview.mp3 --participants 3

Formats: txt (merged speaker turns), vtt, srt, json. Pass --context "names, jargon" to prime Whisper with vocabulary it should expect. --output - streams the transcript to stdout and the summary to stderr, so the output pipes cleanly.

Web UI

transcriber serve runs a FastAPI backend and serves the bundled React frontend. You get:

  • a UMAP scatter where each dot is one segment, colored by cluster — lasso a cluster to bulk-rename it;
  • a continuous waveform with one region per segment — click or scrub to play anything;
  • a Gantt-style speaker timeline;
  • a virtualized transcript with full-text search;
  • inline-renameable speaker chips (renames persist server-side);
  • TXT / VTT / SRT export;
  • keyboard navigation (↑/↓ segments, Space play/pause, / search).

Multiple jobs can run side by side; add more via the sidebar.

serve picks its backend automatically: openai when an OPENAI_API_KEY is available (environment or .env), otherwise local. Pass --backend to choose explicitly. (A legacy single-job Dash UI is still available as transcriber ui.)

For AI agents

This project is built to be driven by agents as well as humans.

Claude Code skill — the repo doubles as a plugin marketplace. Install the skill and Claude Code will know how to transcribe and diarize audio on demand:

/plugin marketplace add Novia-RDI-Seafaring/transcriber
/plugin install dialogue-transcriber@dialogue-transcriber

Structured output--format json --output - emits a stable shape on stdout:

{
  "speakers": ["Speaker 1", "Speaker 2"],
  "n_segments": 42,
  "duration": 512.3,
  "segments": [
    {"speaker": "Speaker 1", "start": 0.0, "end": 4.2, "text": "..."}
  ]
}

MCP / OIP — the package is an Open Ingestion Protocol producer, so transcripts can be ingested by any OIP-aware consumer (e.g. Anchor) with no consumer-side changes:

transcriber oip install --data-dir ~/transcripts     # register the producer
transcriber oip ingest audio.mp3 --data-dir ~/transcripts
transcriber oip serve                                # MCP server (also: transcriber-mcp)

Tool namespace: transcribe. Region kind: transcript_segment. source_ref.kind: audio-timestamp.

Library use

from transcriber.config import ClusterConfig, PipelineConfig, TranscribeConfig
from transcriber.pipeline import run_pipeline
from transcriber.render import render_txt

cfg = PipelineConfig(
    transcribe=TranscribeConfig(backend="local", language="en"),
    cluster=ClusterConfig(participants=2),
)
result = run_pipeline("interview.mp3", config=cfg)
print(render_txt(result.segments))

PipelineResult.segments is a list of SpeakerSegment records with the sentence text, time range, the on-disk clip, and the assigned speaker. PipelineResult.cluster.projection is the 2-D UMAP for plotting.

Backends

Concern Default Override via
Transcribe faster-whisper large-v3 --backend openai
Embed nvidia/speakerverification_en_titanet_large pass embedder= to run_pipeline
Cluster UMAP(2) + KMeans + silhouette pass a ClusterConfig
YouTube yt-dlp replace YouTubeDownloader

All backends are Protocols — see transcriber/transcribe/base.py and transcriber/embed/base.py. Tests use in-memory fakes, so the heavy models are not required to run the suite.

Development

See CONTRIBUTING.md for guidelines and CHANGELOG.md for release history.

git clone https://github.com/Novia-RDI-Seafaring/transcriber
cd transcriber
uv venv
uv pip install -e ".[dev,cluster,api,openai,embed,youtube]"
(cd web && pnpm install && pnpm build)   # so `transcriber serve` can serve the UI

pytest                  # core + clustering + api tests
pytest -m "not slow"    # skip heavy/network tests
ruff check src tests

For frontend work: cd web && pnpm dev (http://127.0.0.1:5173, proxies /api to :8000) with transcriber serve … --port 8000 in another shell.

Releases: publishing a GitHub release triggers .github/workflows/release.yml, which builds the frontend, bundles it into the wheel, and publishes to PyPI via trusted publishing.

License

Apache-2.0 — see LICENSE.

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选