audio2score-mcp

audio2score-mcp

An MCP server that transcribes audio files into MIDI and MusicXML scores, converts MIDI to notation, and extracts playable note data for DAWs. It enables Claude to turn raw recordings into editable, shareable music scores via three simple tools.

Category
访问服务器

README

audio2score-mcp

Turn a recorded audio file into an editable music score: audio → MIDI → MusicXML.

MusicXML is the target because it's the "SVG of music notation" — an open, text-based format any notation app (MuseScore, Sibelius, Guitar Pro, Finale, Dorico) can open, edit, and re-export without owning the pipeline that produced it. This project only produces the .mid and .musicxml files; opening, editing, and exporting to anything else (PDF, audio, tab) happens in whichever notation app you choose, by hand — see "Formats this connects to" below for why this project deliberately doesn't wrap that part.

Two ways to run it: as plain CLI scripts, or as an MCP server exposing the same steps as tools Claude can call.

What's here

  • transcribe.py — audio file → MIDI, via Spotify's basic-pitch
  • to_score.py — MIDI file → MusicXML, via music21
  • score_to_notes.py — a score file (MIDI, MusicXML, or anything music21 can read) → JSON note array in daw-mcp's batch_set_notes format
  • mcp_server.py — MCP server wrapping all three as tools (transcribe_audio, midi_to_score, score_to_notes)

Each step is a separate, real artifact on disk — not a hidden intermediate. The pause between MIDI and notation is deliberate: automatic transcription is lossy, so the raw MIDI is worth a look (or a manual fix) before it becomes a score.

score_to_notes.py is intentionally format-agnostic, not MIDI-specific - music21.converter.parse() handles MIDI and MusicXML identically, so feeding it a .mid from transcribe.py or an .mxl from an external OMR tool (see "Formats this connects to") takes the same code path. There is no separate MusicXML→MIDI or MusicXML→PDF tool in this project - once a .musicxml exists, any notation app already opens and exports it, so building that here would just duplicate what's already installed.

Where files go

Recommended: drop input files in workspace/ (repo-local, gitignored - see .gitignore - nothing placed here ever gets committed, inputs included). There's no hard requirement though - any path works. Every output lands next to its input file, same base name, different extension:

workspace/song.mp3          <- you put this here (any format basic-pitch/librosa reads: mp3, wav, ogg, flac...)
workspace/song.mid          <- transcribe_audio writes this
workspace/song.musicxml     <- midi_to_score writes this (open in MuseScore/Guitar Pro/Sibelius/Finale/Dorico)
workspace/song.notes.json   <- score_to_notes writes this (feed into daw-mcp's batch_set_notes)

To view: open the .mid or .musicxml directly in whatever notation app you have - nothing here launches one for you. If MuseScore calls the .musicxml "corrupted," see the polyphony caveat below before assuming the file is broken.

Setup

Requires Python 3.11 specifically — basic-pitch pulls in TensorFlow 2.15, whose wheels stop at cp311; 3.12 and 3.13 will fail to resolve. The resulting venv is ~2GB (full TensorFlow, not a lighter backend).

uv venv --python 3.11 venv
uv pip install -r requirements.txt --python venv/Scripts/python.exe

Dependencies are pinned exactly (basic-pitch==0.4.0, music21==10.5.0, setuptools==65.5.0, mcp==2.0.0) — this project has no automated test suite, so a fresh environment matching exactly what was verified is the substitute. setuptools specifically is pinned because newer versions break a transitive resampy import that basic-pitch needs.

Usage: CLI

venv/Scripts/python.exe transcribe.py "C:\path\to\song.mp3"
# -> C:\path\to\song.mid

venv/Scripts/python.exe to_score.py "C:\path\to\song.mid"
# -> C:\path\to\song.musicxml

venv/Scripts/python.exe score_to_notes.py "C:\path\to\song.mid"
# -> C:\path\to\song.notes.json  (daw-mcp's batch_set_notes format - also
#    takes a .musicxml/.mxl directly, e.g. from OMR, no separate step needed)

Output always lands next to the input, same base filename, different extension. All three scripts refuse to overwrite an existing output file — delete or move it first if you want to re-run. Errors (missing input, a library failure) print a clear message to stderr and exit non-zero; nothing fails silently.

Run only the tool(s) your actual goal needs - don't chain all three by default. Each tool produces exactly one file; running more than you need just adds files nobody asked for.

Goal Run Files produced
View/edit a recording as notation transcribe_audio → midi_to_score .mid, .musicxml
Get a recording's notes into daw-mcp transcribe_audio → score_to_notes .mid, .notes.json (skip midi_to_score - not needed for this goal)
Get scanned/typeset sheet music into daw-mcp Audiveris (external, see "Formats this connects to") → score_to_notes on the .mxl .mxl, .notes.json (no MIDI step at all)
View/edit scanned sheet music as notation Audiveris only .mxl - already MusicXML, open it directly, no tool here needed

.mid in the first two rows isn't really "output" so much as an unavoidable checkpoint - basic-pitch can only emit MIDI, and it's worth a look before trusting what comes after it (see "Known issue" below on why).

Worked example

A real run, not a hypothetical one. Input: a synthetic mono WAV, a C major arpeggio (C4-E4-G4-C5, quarter notes with a short decaying envelope so onsets are clean) - "real audio" in the sense this project cares about (an actual waveform on disk, not hand-typed MIDI), just synthesized instead of recorded, so the transcript is reproducible without a copyrighted file lying around in a public repo.

$ venv/Scripts/python.exe transcribe.py c_major_arpeggio.wav
WARNING:root:Coremltools is not installed. ...
WARNING:root:tflite-runtime is not installed. ...
WARNING:root:onnxruntime is not installed. ...
Wrote c_major_arpeggio.mid

$ venv/Scripts/python.exe to_score.py c_major_arpeggio.mid
Wrote c_major_arpeggio.musicxml

$ venv/Scripts/python.exe score_to_notes.py c_major_arpeggio.mid
Wrote c_major_arpeggio.notes.json

The three WARNING:root lines are basic-pitch noting that optional backends (CoreML, TFLite, ONNX) aren't installed - harmless, TensorFlow is the backend actually used, and this is exactly what transcribe.py v1.1.1 now correctly hides without corrupting mcp_server.py's stdout stream (see CHANGELOG.md) - it only lands on the terminal, not the MCP protocol channel.

c_major_arpeggio.notes.json, the daw-mcp-ready output:

[[0.0, 60, 83, 1.0], [1.25, 64, 80, 1.0], [2.3333, 67, 80, 1.0], [3.5, 72, 78, 0.5], [4.0, 72, 78, 0.5]]

Four notes went in (C4, E4, G4, C5); basic-pitch correctly detected pitch and velocity for all four (60/64/67/72, matching the arpeggio exactly) but split the last note (C5) into two consecutive entries instead of one - the decaying envelope's tail apparently read as a second onset. This is the automatic-transcription lossiness the "What's here" section above warns about, caught in the wild on the very first note that had a naturalistic (non-flat) volume shape: check the .mid before trusting the .musicxml/.notes.json blindly, especially around sustained or decaying notes.

c_major_arpeggio.musicxml opens cleanly in any notation app (verified well-formed: correct MusicXML 4.0 DOCTYPE, <step>/<octave> pitches for C4/E4/G4/G4/C5/C5/C5 - the split C5 shows up as tied notes across a measure boundary, which is standard MusicXML for a note that doesn't fit in one measure, not a second bug).

Usage: MCP server

Registered in Claude Code's config as audio2score — restart Claude Code after a fresh install for it to appear (MCP servers load at startup).

Three tools, mirroring the three scripts exactly:

  • transcribe_audio(audio_path) → returns the .mid path
  • midi_to_score(midi_path) → returns the .musicxml path
  • score_to_notes(score_path) → returns the .notes.json path (daw-mcp's batch_set_notes note-array format; accepts MIDI or MusicXML)

Same behavior as the CLI underneath (same overwrite guard, same errors) — the MCP server is a thin wrapper, not a different implementation.

One thing to know if a call seems to hang: if a transcribe_audio call appears to time out or gets cancelled, the transcription may still be running in the background and will finish writing the .mid file regardless. A retry will then hit the overwrite guard ("already exists") even though the first call looked like it never succeeded. This isn't a bug — check whether the .mid already exists before retrying.

To register the server yourself elsewhere, add this to your MCP config (mcpServers), using absolute paths for both fields — the client launches stdio servers without a defined working directory, so relative paths won't resolve:

"audio2score": {
  "type": "stdio",
  "command": "<absolute path to>\\venv\\Scripts\\python.exe",
  "args": ["<absolute path to>\\mcp_server.py"],
  "env": {}
}

What this doesn't do

  • No score/MIDI → audio, PDF, or tab output, no MusicXML → MIDI conversion either — see "Formats this connects to" below for why and what to use instead
  • No stem separation or multi-instrument splitting
  • No automated test suite by design — verification is always a real run against real audio

Formats this connects to

Once a .musicxml exists, this project deliberately stops - every notation app already opens MusicXML natively and exports whatever's needed (PDF, audio, tab, MIDI) from its own menu. Building automated wrappers around those exports was tried and mostly reverted (see CHANGELOG.md v1.2.0 through v2.0.0 for the full back-and-forth) - the one direction still worth automating turned out to be none of them, once score_to_notes.py was confirmed to accept MusicXML directly with no separate conversion step.

Direction Use Notes
MusicXML → PDF, audio, tab, MIDI MuseScore Studio or Guitar Pro, opened normally Not wrapped here on purpose - see above. (MuseScore's CLI converter mode, -j job.json, genuinely can automate PDF export reliably if you want it for your own scripting - just isn't built into this project)
PDF (scanned/typeset sheet music) → MusicXML Audiveris (C:\Program Files\Audiveris\Audiveris.exe): Audiveris.exe -batch -export -output "<folder>" "<input>.pdf" -batch genuinely skips its GUI. Tested on 3 real PDFs: 2 clean one-page scores exported correctly (one with a minor time-signature warning); a 24-page guitar tab book hit real internal Audiveris crashes (NullPointerException/IndexOutOfBoundsException in its rhythm analysis) on several pages - OMR reliability drops fast on complex, multi-page, or tab-heavy input
PDF → daw-mcp's note format Audiveris (above) → this project's score_to_notes.py, directly on the .mxl Two steps, both real and tested end-to-end on actual sheet music - no MIDI conversion needed in between

Treat OMR output with at least as much suspicion as basic-pitch's audio transcription - check the intermediate .musicxml before trusting it, and don't expect Audiveris to succeed on every PDF (see the tab-book failure above).

Known issue: MuseScore can reject a transcribed .musicxml as "corrupted"

Heavily polyphonic transcriptions can produce a .musicxml that MuseScore Studio refuses to open, calling it "corrupted." Root cause: music21's own MusicXML writer omits the <voice> tag on some <note> elements when a piece needs many simultaneous voices (5+) - verified on a real 45-second recording that transcribed into dense, often-overlapping notes (a side effect of basic-pitch picking up harmonics/artifacts on real audio, not a clean single melodic line). Confirmed this is music21's writer, not this project's code: to_score.py is a two-line parse() + write() call with no note/voice logic of its own, and explicitly calling score.makeNotation() before writing doesn't fix it either. Re-parsing the same file with music21 itself only warns (Cannot put in an element with a missing voice tag) and recovers by defaulting those notes to voice 1 - MuseScore's importer is simply stricter and rejects outright instead of tolerating it. Workaround: click "Open anyway" - it loads fine, just with those specific notes in voice 1 instead of their originally-detected voice, a minor layout quirk, not lost data. Not seen on clean, low-polyphony input (a hand-authored melody MIDI transcribed and re-verified with zero voice-tag issues) - this is specific to messy, dense, real-audio-transcription output.

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选