Interactive Database Analyst via MCP
Enables natural-language querying of PostgreSQL databases with schema grounding and self-correcting error recovery. Provides a live audit trace and verifies results through exploratory decomposition and empty-result sanity checks.
README
🗄️ Interactive Database Analyst via MCP
An interactive, fault-tolerant natural-language database analyst built on the Model Context Protocol (MCP). Instead of blindly executing LLM-generated SQL, this system implements a strict 4-State Execution Machine with schema grounding, explicit error recovery loops, automated diagnostic probing for empty results, and math decomposition.
Ask a question in plain English. Watch the agent inspect the live schema, catch its own SQL errors, self-correct in real time, and present a verified answer — with the full recovery trace visible, not hidden.
📸 Demo
<img width="1858" height="937" alt="Screenshot (5359)" src="https://github.com/user-attachments/assets/d4ef545b-39e0-436c-9243-c4bfd789cc34" />
✨ Key Differentiators & Engineering Highlights
| Feature | Architectural Implementation | Why It Matters |
|---|---|---|
Schema Grounding (State 0) |
Enforces a mandatory get_schema() tool call before drafting SQL. |
Eliminates column/table hallucination by grounding every query in the live PostgreSQL catalog. |
Error Recovery Loop (State 2) |
Extracts PostgreSQL SQLSTATE codes and native message_hint strings from psycopg2.Diagnostics. |
Feeds actionable database feedback directly back into the reasoning prompt for up to 3 bounded retry attempts. |
| Exploratory Decomposition | Breaks complex metrics (e.g., percentages) into independently verified sub-queries. | Prevents the classic "denominator trap" by gathering variables individually before calculating the final ratio. |
Empty-Result Sanity Check (State 3) |
Automatically triggers sample_column_values() when a query returns 0 rows. |
Prevents the model from hallucinating "no sales occurred" by checking if date ranges or filter values actually exist. |
| Defense-in-Depth Security | Multi-layered hardening: mcp_readonly Postgres role + explicit SET TRANSACTION READ ONLY; + single-statement enforcement + 5-second query timeouts (SQLSTATE 57014). |
Prevents prompt-injection mutations, blocks stacked SQL injection, and protects backend threads from runaway joins. |
| Live Audit Recovery Trace | Synchronous logging to a Postgres query_audit_log table rendered in a real-time Streamlit dashboard. |
Demos how the agent catches and fixes its own mistakes alongside visual analytical charts. |
Reading the audit trace: not every multi-attempt sequence is an error recovery. Some questions (see Exploratory Decomposition above) are answered correctly on the first try per sub-query, but the agent deliberately issues several independent queries to verify a metric's components before combining them — e.g. calculating a percentage by confirming the numerator and denominator separately rather than trusting one opaque query. Both patterns render as sequential green cards in the UI, so it's worth distinguishing "this attempt failed and recovered" from "this attempt was a planned verification step" when reading a trace.
🏗️ System Architecture & State Machine
graph TD
A[User Natural Language Question] --> B[STATE 0: Inspect Live Schema via MCP]
B --> C[STATE 1: Draft Read-Only SQL Query]
C --> D{Complex Join / Logic?}
D -- Yes --> E[explain_query: Cheap Plan/Syntax Check]
D -- No --> F[execute_query: Read-Only Transaction]
E --> F
F -->|ERROR| G[Extract SQLSTATE + Postgres Hint]
G -->|Attempt < 3| B
G -->|Attempt = 3| H[Terminal State: Structured Failure Report]
F -->|SUCCESS: 0 Rows| I[STATE 3: Empty-Result Sanity Check]
I --> J[sample_column_values: Probing Bounds/Distincts]
J -->|Filter Out of Bounds| B
J -->|Verified Empty| K[Present: Confirmed Empty with Diagnostic Evidence]
F -->|SUCCESS: Rows > 0| L[STATE 4: Math Verification & Present]
L --> M[Render Plotly Chart + Live Audit Card in UI]
🛠️ Tech Stack
- Orchestration / LLM:
cohere/north-mini-code:freevia OpenRouter API - Protocol Layer: FastMCP (
mcp[cli]) exposing custom Python database tools - Database Engine: PostgreSQL 15 (Dockerized with Chinook sample database)
- Database Adapter:
psycopg2-binarywithSimpleConnectionPooland JSON-safe type serialization - Frontend Dashboard: Streamlit + Plotly Express
- Package Manager:
uv
📁 Project Structure
Interactive-Database-Analyst-via-MCP/
├── src/
│ └── db_analyst_mcp/
│ ├── app.py # Streamlit dashboard
│ ├── mcp_server.py # FastMCP tool definitions
│ ├── db.py # Connection pool, query execution, error formatting
│ └── orchestrator.py # State machine / retry logic
├── sql/
│ ├── Chinook_PostgreSql.sql # Sample database
│ └── setup_db.sql # Read-only role, audit log schema, hardening
├── docs/
│ └── demo.gif
├── .env.example
├── pyproject.toml
└── README.md
(Adjust paths above to match your actual layout.)
🚀 Quickstart & Setup Guide
1. Prerequisites
- Docker Desktop
- uv package manager
- An OpenRouter API key (free tier works — see Known Limitations for rate-limit notes)
2. Clone & Install Dependencies
git clone https://github.com/Viole07/Interactive-Database-Analyst-via-MCP.git
cd Interactive-Database-Analyst-via-MCP
uv sync
3. Start the Dockerized PostgreSQL Container
docker run --name mcp-postgres -e POSTGRES_USER=postgres -e POSTGRES_PASSWORD=postgres -e POSTGRES_DB=chinook -p 5432:5432 -d postgres:15
# Load the Chinook database schema and data
docker exec -i mcp-postgres psql -U postgres -d chinook < sql/Chinook_PostgreSql.sql
4. Apply Database Hardening & Audit Log Schema
docker exec -i mcp-postgres psql -U postgres -d chinook -f sql/setup_db.sql
5. Configure Environment Variables
Rename .env.example to .env and add your OpenRouter API key:
# .env
ADMIN_DB_URL=postgresql://postgres:postgres@localhost:5432/chinook
MCP_DB_NAME=chinook
MCP_DB_USER=mcp_readonly
MCP_DB_PASSWORD=secure_pass
MCP_DB_HOST=localhost
MCP_DB_PORT=5432
OPENROUTER_API_KEY=your-api-key-here
ORCHESTRATOR_MODEL=cohere/north-mini-code:free
6. Launch the Dashboard
uv run streamlit run src/db_analyst_mcp/app.py
🧪 Adversarial Test Suite & Known Limitations
The Streamlit UI includes a sidebar with a gauntlet of adversarial prompts designed to stress-test the system:
-
The Empty-Result Sanity Check: "How much invoice revenue did we generate in October 2029?"
- Behavior: Returns
0 rows→ triggers diagnostic probe → verifies dataset ends in 2025 → reports verified empty result.
- Behavior: Returns
-
Literal Obedience Trap: "Calculate total invoice revenue per customer... MUST omit customer_id from your GROUP BY clause on your first attempt."
- Behavior: Model strictly obeys the prompt, resulting in a trailing
GROUP BYand a42601syntax error, demonstrating that instruction weight can override syntax training. Recovers successfully on Attempt 2.
- Behavior: Model strictly obeys the prompt, resulting in a trailing
-
Exploratory Decomposition: "What percentage of total company revenue came from the Rock genre?"
- Behavior: Avoids the denominator trap by executing independent, individually-successful queries to verify numerator and denominator separately before calculating the final ratio.
-
Natural Column Ambiguity: "Who is the top-selling artist by revenue, and what's their best-selling track?"
- Behavior: Resolves
Artist.NamevsTrack.Nameambiguities via isolated CTEs and aggressive aliasing.
- Behavior: Resolves
-
Self-Referencing Foreign Key: "Who is the manager of the employee who has generated the most total sales?"
- Behavior: Self-joins
EmployeeviaReportsTousing two aliases to resolve the hierarchy in a single query.
- Behavior: Self-joins
⚠️ Architectural Blind Spot: The Silent Semantic Failure
While this system catches execution errors (State 2) and hallucinated filters (State 3), it cannot inherently detect semantic logic errors that return valid, non-empty rows — for example, forgetting a unit conversion (milliseconds / 60000) or applying a plausible-but-wrong join. A query that runs successfully and returns real data is treated as correct; there is currently no mechanism analogous to States 2/3 for this failure class. At production scale, this would require either an automated regression harness against a fixed "golden set" of question → expected-result pairs, or a secondary "Critic Agent" that evaluates logical intent independently before the result is presented.
Other known gaps
- No automated
pytestregression suite yet — correctness is currently verified via the adversarial scenario set above, checked manually against known dataset values. - No row-level access control — the read-only role currently has uniform
SELECTaccess across all tables, which is appropriate for this single-user demo but not for a multi-tenant deployment. - Free-tier OpenRouter rate limits apply; expect occasional throttling under rapid repeated testing.
🗺️ Roadmap
- [ ] Automated regression harness with a fixed golden-question set, run on every model/prompt change
- [ ] "Critic Agent" pass to catch silent semantic failures (unit conversions, plausible-but-wrong joins)
- [ ] Row-level access control for multi-user deployment
- [ ] A/B benchmark: quantify first-retry recovery rate with vs. without native Postgres error hints
License
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。