speak
Lets Claude convert its replies into natural spoken scripts and play them with pause, resume, and rewind controls, plus voice and engine selection.
README
speak
Ask Claude to read its answer out loud.
Not a text-to-speech dump of the raw reply. You say "read that back to me simply" and Claude rewrites what it said as a short spoken script, then plays it. The script never appears in chat. It is a conversation between you and the model, with the text kept out of your way.
Works in Claude Code today. The playback half is a plain MCP server, so it also drops into Claude Desktop, Cursor, or any MCP client.
What you get
speakMCP server with toolsspeak,stop,pause,resume,back,skip,add_pronunciation,list_pronunciations,list_voices,set_default- Floating overlay on macOS while Claude speaks: a small pill above every window with a live waveform, back one sentence, skip one sentence, pause and resume from the exact point, stop
/speakskill for Claude Code, with styles:simple(default),brief,decisions,full,eli5- Auto-speak: a
Stophook that reads abriefscript after every answer, off by default speakCLI:echo "hello" | speak- Raycast script commands:
Speakreads the selected text as written,Speak Simplyrewrites it withclaude -pas a short spoken summary first,Speak Translatedtranslates it first (English by default; type another language as the command's argument, or setSPEAK_TRANSLATE_LANGUAGEto change the default). A non-English translation is read with a matching installed macOS voice, best available first (Premium, then Enhanced, then compact); English keeps your normal voice untouched. Download the Premium voice for a language in System Settings > Accessibility > Read & Speak > System Voice > Manage Voices to get the good one.Speak Stopstops. They copy the selection with a simulated Cmd+C, so Raycast needs Accessibility permission.
Engines
| Engine | Cost | Needs | Notes |
|---|---|---|---|
say |
free | macOS | Default. Download a Premium or Enhanced voice in System Settings > Accessibility > Spoken Content for good quality. Set a Siri voice as system default and bare say uses it. |
edge |
free | network | Microsoft neural voices via edge-tts-universal. Very good quality. Unofficial API. |
kokoro |
free | npm i kokoro-js, ~300MB model on first run |
Local neural TTS, runs on CPU. |
openai |
paid | OPENAI_API_KEY |
gpt-4o-mini-tts. |
elevenlabs |
free tier | ELEVENLABS_API_KEY |
eleven_flash_v2_5. |
Picking a voice on macOS
The default say engine uses your Mac's system voice, so set it up once in macOS and every reading uses it:
- Open System Settings > Accessibility > Read & Speak (Spoken Content on older versions).
- Under System Voice, pick a voice. Apple ships a large library per language; the Premium and Enhanced ones download on demand and sound far better than the compact defaults. Siri voices can be chosen here too.
- Set the speaking rate to taste.
Pronunciations
Tell Claude in chat:
- "pronounce fancyapp as fan-see-app from now on"
- "what pronunciations have you saved"
They live in ~/.config/speak/pronunciations.json (override the folder with SPEAK_CONFIG_DIR), matched whole-word and case-insensitive, and applied to every reading:
{
"fancyapp": "fan-see-app",
"OAuth": "oh auth"
}
Links and emails are read as spoken: https://example.com/login becomes "example dot com slash login", jane@example.de becomes "jane at example dot d e". Short endings like de, io, co.uk are spelled letter by letter, since "de" read as a word comes out as "duh".
Words with two readings ("live", "read", "lead") cannot be fixed by a word list. The skill tells Claude to respell them in the script so the voice cannot guess wrong: "lyive" for live as in alive, "red" for read in the past tense, "led" for the metal. The skill carries a table of the usual offenders and checks the script against it before speaking.
Leave SPEAK_VOICE unset to follow the system voice, or set it to a voice name from say -v ? to override for Claude only.
First available engine wins unless you set SPEAK_ENGINE. SPEAK_VOICE sets the voice. Both can be changed per session with the set_default tool ("switch to the edge engine").
Install
Set RAYCAST_SCRIPTS_DIR to a directory already added in Raycast (Settings > Extensions > Script Commands) and the commands are linked into it. Without it, add raycast/ from this repo there once by hand.
RAYCAST_SCRIPTS_DIR=~/raycast-scripts ./install.sh
Needs Node 20+, jq, and Claude Code.
git clone https://github.com/CareyScott/speak
cd speak
./install.sh
This builds the server and the macOS overlay helper (needs swiftc from the Xcode command line tools), registers the MCP server at user scope, links the skill into ~/.claude/skills, adds the Stop hook to ~/.claude/settings.json, and links speak and speak-auto into ~/.local/bin.
Restart Claude Code.
Use
Talk to Claude:
- "read that back to me"
- "say it simply"
- "what do you need from me, out loud"
- "/speak decisions"
- "stop"
Auto-speak after every answer:
speak-auto on # brief style
speak-auto on decisions # only the questions for you
speak-auto off
While it speaks, use the overlay or just say it:
- "pause", "resume", "go back a sentence", "skip that", "stop"
Pause holds the exact position. Resume continues from it. Back replays the previous sentence. Skip jumps to the next. Nothing is spoken about pausing or resuming; it just does it.
Pick an engine or voice:
export SPEAK_ENGINE=edge
export SPEAK_VOICE=en-IE-ConnorNeural
Or in chat: "list the say voices", "use Jamie from now on".
How it works
The skill tells Claude how to write for the ear: short sentences, lead with the point, describe code instead of reading it, end with the decision it needs from you. Claude calls the speak tool with that script.
The server strips any leftover markdown, splits the script into sentences, and synthesises the next sentence while the current one plays.
On macOS a small Swift helper (overlay/main.swift) owns playback with AVAudioPlayer and draws the overlay: an always-on-top, non-activating panel that never steals focus. Node sends it one sentence file at a time over JSON lines on stdin and it reports finished, back, or stop on stdout. Pause and resume happen inside the helper, so the position is exact. Back and skip tell the server which sentence to play next. Elsewhere the server falls back to afplay or ffplay with no overlay.
Auto-speak is a Claude Code Stop hook. When the flag file ~/.config/speak/auto exists, the hook blocks the stop once and asks Claude to speak a script in the style named in the file. The stop_hook_active guard stops it looping.
Development
npm run dev # run the server with tsx
npm test # vitest
npm run typecheck
MIT.
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。