AI Voice Assistant MCP Server

AI Voice Assistant MCP Server

Enables a voice-enabled AI assistant to call 7 built-in MCP tools including calculator, web search (DuckDuckGo), weather (wttr.in), date/time, and local file read/write/list operations, integrating with Gemini 2.0 Flash for tool-calling conversations.

Category
访问服务器

README

AI Voice Assistant with MCP Tool Calling

A fully free, end-to-end voice-enabled AI assistant built with Python.

Component Technology
🧠 AI Brain Google Gemini API (gemini-2.0-flash) — free tier
🎙️ Speech-to-Text Google Web Speech API via SpeechRecognition — free
🔊 Text-to-Speech pyttsx3 (offline) — free
🛠️ Tool Calling Model Context Protocol (MCP) — free
🔍 Web Search DuckDuckGo — free, no key needed
🌤️ Weather wttr.in REST API — free, no key needed

Features

  • 🎙️ Voice Input — speak naturally; the assistant understands you
  • 🔊 Voice Output — responses are read aloud via offline TTS
  • 🤖 Gemini AI — context-aware, multi-turn conversation
  • 🛠️ 7 MCP Tools available to the AI:
    1. Calculator — safe math expression evaluator
    2. Web Search — DuckDuckGo (no API key)
    3. Weather — real-time via wttr.in (no API key)
    4. Date/Time — current date and time
    5. Read File — read any local file
    6. Write File — write/append to a local file
    7. List Directory — browse local folders
  • 💬 Multi-turn memory — remembers conversation context
  • ⌨️ Text mode — works without a microphone (--text flag)

Prerequisites

  • Python 3.11+
  • A free Gemini API key — get one at https://aistudio.google.com/app/apikey
  • Internet connection (for speech recognition, Gemini API, and web tools)
  • Microphone (optional — text mode works without one)

Installation

1. Clone / download the project

# If using git:
git clone <your-repo-url>
cd "AI voice assistant"

# Or just open the folder in your terminal
cd "C:\Users\sasid\Downloads\AI voice assistant"

2. Create and activate a virtual environment (recommended)

python -m venv .venv

# Windows:
.venv\Scripts\activate

# macOS / Linux:
source .venv/bin/activate

3. Install PyAudio (Windows — required for microphone)

PyAudio on Windows needs a pre-built binary. The easiest way:

pip install pipwin
pipwin install pyaudio

Or download the correct .whl from https://www.lfd.uci.edu/~gohlke/pythonlibs/#pyaudio and install with:

pip install PyAudio‑0.2.14‑cpXX‑cpXX‑win_amd64.whl

4. Install remaining dependencies

pip install -r requirements.txt

5. Set your Gemini API key

Option A — .env file (recommended):

copy .env.example .env
# Then open .env and replace "your_gemini_api_key_here" with your actual key

Option B — edit config.py directly:

Open config.py and change:

GEMINI_API_KEY: str = os.getenv("GEMINI_API_KEY", "YOUR_GEMINI_API_KEY_HERE")

to:

GEMINI_API_KEY: str = "your_actual_api_key"

Running

Voice mode (default — microphone + TTS)

python main.py

Text-only mode (no microphone needed)

python main.py --text

List available TTS voices

python main.py --list-voices

Example Interactions

You say What happens
"What's the weather in Tokyo?" Calls get_weather MCP tool → speaks result
"Calculate 2 to the power of 32" Calls calculator tool → speaks 4294967296
"Search for the latest Python news" Calls web_search → summarises top results
"What day is today?" Calls get_datetime → speaks date & time
"Read the file notes.txt" Calls read_file → speaks file contents
"Write 'Hello World' to test.txt" Calls write_file → creates/updates file
"Reset conversation" Clears chat history
"Goodbye" / "Exit" Exits the assistant

Project Structure

AI voice assistant/
├── main.py          # Entry point — CLI, banner, main loop
├── assistant.py     # Gemini + MCP integration (agentic tool-call loop)
├── speech.py        # SpeechRecognition (STT) + pyttsx3 (TTS)
├── mcp_server.py    # MCP tool server with 7 built-in tools
├── config.py        # All settings and API key placeholder
├── requirements.txt # Python dependencies
├── .env.example     # API key template
└── README.md        # This file

Customisation

Change the AI's personality

Edit SYSTEM_PROMPT in config.py.

Adjust microphone sensitivity

Edit MIC_ENERGY_THRESHOLD in config.py (lower = more sensitive).

Change TTS voice or speed

Edit TTS_RATE and TTS_VOICE_PREFERENCE in config.py. Run python main.py --list-voices to see available voice names.

Add more MCP tools

Open mcp_server.py, add a new function, then register it in list_tools() and call_tool().


Gemini Free Tier Limits

Limit Value
Requests per minute 15
Tokens per day 1,000,000
Cost $0

Get your key at: https://aistudio.google.com/app/apikey


Troubleshooting

"No module named 'pyaudio'" → See PyAudio installation step above.

"Could not understand audio" → Speak clearly; adjust MIC_ENERGY_THRESHOLD lower in config.py.

"Speech recognition service error" → Check your internet connection (Google Web Speech API requires internet).

Gemini 429 / rate limit error → You've hit the free tier limit. Wait a minute and try again.

Assistant doesn't speak / TTS silent → Check system audio / volume. Try python main.py --list-voices to verify pyttsx3 works.


License

MIT — free to use, modify, and distribute.

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选