audio2score-mcp
An MCP server that transcribes audio files into MIDI and MusicXML scores, converts MIDI to notation, and extracts playable note data for DAWs. It enables Claude to turn raw recordings into editable, shareable music scores via three simple tools.
README
audio2score-mcp
Turn a recorded audio file into an editable music score: audio → MIDI → MusicXML.
MusicXML is the target because it's the "SVG of music notation" — an open,
text-based format any notation app (MuseScore, Sibelius, Guitar Pro, Finale,
Dorico) can open, edit, and re-export without owning the pipeline that
produced it. This project only produces the .mid and .musicxml files;
opening, editing, and exporting to anything else (PDF, audio, tab) happens
in whichever notation app you choose, by hand — see "Formats this connects
to" below for why this project deliberately doesn't wrap that part.
Two ways to run it: as plain CLI scripts, or as an MCP server exposing the same steps as tools Claude can call.
What's here
transcribe.py— audio file → MIDI, via Spotify's basic-pitchto_score.py— MIDI file → MusicXML, via music21score_to_notes.py— a score file (MIDI, MusicXML, or anything music21 can read) → JSON note array in daw-mcp'sbatch_set_notesformatmcp_server.py— MCP server wrapping all three as tools (transcribe_audio,midi_to_score,score_to_notes)
Each step is a separate, real artifact on disk — not a hidden intermediate. The pause between MIDI and notation is deliberate: automatic transcription is lossy, so the raw MIDI is worth a look (or a manual fix) before it becomes a score.
score_to_notes.py is intentionally format-agnostic, not MIDI-specific -
music21.converter.parse() handles MIDI and MusicXML identically, so
feeding it a .mid from transcribe.py or an .mxl from an external OMR
tool (see "Formats this connects to") takes the same code path. There is
no separate MusicXML→MIDI or MusicXML→PDF tool in this project - once a
.musicxml exists, any notation app already opens and exports it, so
building that here would just duplicate what's already installed.
Where files go
Recommended: drop input files in workspace/ (repo-local, gitignored -
see .gitignore - nothing placed here ever gets committed, inputs
included). There's no hard requirement though - any path works. Every
output lands next to its input file, same base name, different
extension:
workspace/song.mp3 <- you put this here (any format basic-pitch/librosa reads: mp3, wav, ogg, flac...)
workspace/song.mid <- transcribe_audio writes this
workspace/song.musicxml <- midi_to_score writes this (open in MuseScore/Guitar Pro/Sibelius/Finale/Dorico)
workspace/song.notes.json <- score_to_notes writes this (feed into daw-mcp's batch_set_notes)
To view: open the .mid or .musicxml directly in whatever notation app
you have - nothing here launches one for you. If MuseScore calls the
.musicxml "corrupted," see the polyphony caveat below before assuming
the file is broken.
Setup
Requires Python 3.11 specifically — basic-pitch pulls in TensorFlow
2.15, whose wheels stop at cp311; 3.12 and 3.13 will fail to resolve. The
resulting venv is ~2GB (full TensorFlow, not a lighter backend).
uv venv --python 3.11 venv
uv pip install -r requirements.txt --python venv/Scripts/python.exe
Dependencies are pinned exactly (basic-pitch==0.4.0, music21==10.5.0,
setuptools==65.5.0, mcp==2.0.0) — this project has no automated test
suite, so a fresh environment matching exactly what was verified is the
substitute. setuptools specifically is pinned because newer versions
break a transitive resampy import that basic-pitch needs.
Usage: CLI
venv/Scripts/python.exe transcribe.py "C:\path\to\song.mp3"
# -> C:\path\to\song.mid
venv/Scripts/python.exe to_score.py "C:\path\to\song.mid"
# -> C:\path\to\song.musicxml
venv/Scripts/python.exe score_to_notes.py "C:\path\to\song.mid"
# -> C:\path\to\song.notes.json (daw-mcp's batch_set_notes format - also
# takes a .musicxml/.mxl directly, e.g. from OMR, no separate step needed)
Output always lands next to the input, same base filename, different extension. All three scripts refuse to overwrite an existing output file — delete or move it first if you want to re-run. Errors (missing input, a library failure) print a clear message to stderr and exit non-zero; nothing fails silently.
Run only the tool(s) your actual goal needs - don't chain all three by default. Each tool produces exactly one file; running more than you need just adds files nobody asked for.
| Goal | Run | Files produced |
|---|---|---|
| View/edit a recording as notation | transcribe_audio → midi_to_score |
.mid, .musicxml |
| Get a recording's notes into daw-mcp | transcribe_audio → score_to_notes |
.mid, .notes.json (skip midi_to_score - not needed for this goal) |
| Get scanned/typeset sheet music into daw-mcp | Audiveris (external, see "Formats this connects to") → score_to_notes on the .mxl |
.mxl, .notes.json (no MIDI step at all) |
| View/edit scanned sheet music as notation | Audiveris only | .mxl - already MusicXML, open it directly, no tool here needed |
.mid in the first two rows isn't really "output" so much as an
unavoidable checkpoint - basic-pitch can only emit MIDI, and it's worth a
look before trusting what comes after it (see "Known issue" below on why).
Worked example
A real run, not a hypothetical one. Input: a synthetic mono WAV, a C major arpeggio (C4-E4-G4-C5, quarter notes with a short decaying envelope so onsets are clean) - "real audio" in the sense this project cares about (an actual waveform on disk, not hand-typed MIDI), just synthesized instead of recorded, so the transcript is reproducible without a copyrighted file lying around in a public repo.
$ venv/Scripts/python.exe transcribe.py c_major_arpeggio.wav
WARNING:root:Coremltools is not installed. ...
WARNING:root:tflite-runtime is not installed. ...
WARNING:root:onnxruntime is not installed. ...
Wrote c_major_arpeggio.mid
$ venv/Scripts/python.exe to_score.py c_major_arpeggio.mid
Wrote c_major_arpeggio.musicxml
$ venv/Scripts/python.exe score_to_notes.py c_major_arpeggio.mid
Wrote c_major_arpeggio.notes.json
The three WARNING:root lines are basic-pitch noting that optional
backends (CoreML, TFLite, ONNX) aren't installed - harmless, TensorFlow is
the backend actually used, and this is exactly what transcribe.py v1.1.1
now correctly hides without corrupting mcp_server.py's stdout stream
(see CHANGELOG.md) - it only lands on the terminal, not the MCP protocol
channel.
c_major_arpeggio.notes.json, the daw-mcp-ready output:
[[0.0, 60, 83, 1.0], [1.25, 64, 80, 1.0], [2.3333, 67, 80, 1.0], [3.5, 72, 78, 0.5], [4.0, 72, 78, 0.5]]
Four notes went in (C4, E4, G4, C5); basic-pitch correctly detected pitch
and velocity for all four (60/64/67/72, matching the arpeggio exactly) but
split the last note (C5) into two consecutive entries instead of one -
the decaying envelope's tail apparently read as a second onset. This is
the automatic-transcription lossiness the "What's here" section above
warns about, caught in the wild on the very first note that had a
naturalistic (non-flat) volume shape: check the .mid before trusting the
.musicxml/.notes.json blindly, especially around sustained or decaying
notes.
c_major_arpeggio.musicxml opens cleanly in any notation app (verified
well-formed: correct MusicXML 4.0 DOCTYPE, <step>/<octave> pitches for
C4/E4/G4/G4/C5/C5/C5 - the split C5 shows up as tied notes across a
measure boundary, which is standard MusicXML for a note that doesn't fit
in one measure, not a second bug).
Usage: MCP server
Registered in Claude Code's config as audio2score — restart Claude Code
after a fresh install for it to appear (MCP servers load at startup).
Three tools, mirroring the three scripts exactly:
transcribe_audio(audio_path)→ returns the.midpathmidi_to_score(midi_path)→ returns the.musicxmlpathscore_to_notes(score_path)→ returns the.notes.jsonpath (daw-mcp'sbatch_set_notesnote-array format; accepts MIDI or MusicXML)
Same behavior as the CLI underneath (same overwrite guard, same errors) — the MCP server is a thin wrapper, not a different implementation.
One thing to know if a call seems to hang: if a transcribe_audio call
appears to time out or gets cancelled, the transcription may still be
running in the background and will finish writing the .mid file
regardless. A retry will then hit the overwrite guard ("already exists")
even though the first call looked like it never succeeded. This isn't a
bug — check whether the .mid already exists before retrying.
To register the server yourself elsewhere, add this to your MCP config
(mcpServers), using absolute paths for both fields — the client
launches stdio servers without a defined working directory, so relative
paths won't resolve:
"audio2score": {
"type": "stdio",
"command": "<absolute path to>\\venv\\Scripts\\python.exe",
"args": ["<absolute path to>\\mcp_server.py"],
"env": {}
}
What this doesn't do
- No score/MIDI → audio, PDF, or tab output, no MusicXML → MIDI conversion either — see "Formats this connects to" below for why and what to use instead
- No stem separation or multi-instrument splitting
- No automated test suite by design — verification is always a real run against real audio
Formats this connects to
Once a .musicxml exists, this project deliberately stops - every
notation app already opens MusicXML natively and exports whatever's
needed (PDF, audio, tab, MIDI) from its own menu. Building automated
wrappers around those exports was tried and mostly reverted (see
CHANGELOG.md v1.2.0 through v2.0.0 for the full back-and-forth) - the
one direction still worth automating turned out to be none of them, once
score_to_notes.py was confirmed to accept MusicXML directly with no
separate conversion step.
| Direction | Use | Notes |
|---|---|---|
| MusicXML → PDF, audio, tab, MIDI | MuseScore Studio or Guitar Pro, opened normally | Not wrapped here on purpose - see above. (MuseScore's CLI converter mode, -j job.json, genuinely can automate PDF export reliably if you want it for your own scripting - just isn't built into this project) |
| PDF (scanned/typeset sheet music) → MusicXML | Audiveris (C:\Program Files\Audiveris\Audiveris.exe): Audiveris.exe -batch -export -output "<folder>" "<input>.pdf" |
-batch genuinely skips its GUI. Tested on 3 real PDFs: 2 clean one-page scores exported correctly (one with a minor time-signature warning); a 24-page guitar tab book hit real internal Audiveris crashes (NullPointerException/IndexOutOfBoundsException in its rhythm analysis) on several pages - OMR reliability drops fast on complex, multi-page, or tab-heavy input |
| PDF → daw-mcp's note format | Audiveris (above) → this project's score_to_notes.py, directly on the .mxl |
Two steps, both real and tested end-to-end on actual sheet music - no MIDI conversion needed in between |
Treat OMR output with at least as much suspicion as basic-pitch's audio
transcription - check the intermediate .musicxml before trusting it,
and don't expect Audiveris to succeed on every PDF (see the tab-book
failure above).
Known issue: MuseScore can reject a transcribed .musicxml as "corrupted"
Heavily polyphonic transcriptions can produce a .musicxml that MuseScore
Studio refuses to open, calling it "corrupted." Root cause: music21's own
MusicXML writer omits the <voice> tag on some <note> elements when a
piece needs many simultaneous voices (5+) - verified on a real 45-second
recording that transcribed into dense, often-overlapping notes (a side
effect of basic-pitch picking up harmonics/artifacts on real audio, not a
clean single melodic line). Confirmed this is music21's writer, not this
project's code: to_score.py is a two-line parse() + write() call with
no note/voice logic of its own, and explicitly calling score.makeNotation()
before writing doesn't fix it either. Re-parsing the same file with music21
itself only warns (Cannot put in an element with a missing voice tag) and
recovers by defaulting those notes to voice 1 - MuseScore's importer is
simply stricter and rejects outright instead of tolerating it.
Workaround: click "Open anyway" - it loads fine, just with those specific
notes in voice 1 instead of their originally-detected voice, a minor layout
quirk, not lost data. Not seen on clean, low-polyphony input (a hand-authored
melody MIDI transcribed and re-verified with zero voice-tag issues) - this is
specific to messy, dense, real-audio-transcription output.
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。