Agentic RAG MCP
A multi-agent Retrieval-Augmented Generation system exposed as an MCP server. Ask a question and a LangGraph pipeline plans the retrieval, pulls evidence from a pgvector knowledge base, optionally augments it with live web research, drafts a cited answer, and then self-critiques it for grounding — revising until the answer is supported by the sources.
README
Agentic RAG MCP
<p align="center"> <img src="docs/img/01-architecture.png" alt="Agentic RAG — Multi-Agent Graph" width="880"> </p>
A multi-agent Retrieval-Augmented Generation system exposed as an MCP server. Ask a question and a LangGraph pipeline plans the retrieval, pulls evidence from a pgvector knowledge base, optionally augments it with live web research, drafts a cited answer, and then self-critiques it for grounding — revising until the answer is supported by the sources.
It plugs into any MCP client (Claude Code/Desktop, Cursor, Windsurf, …) as three tools:
ingest, ask, and search.
Why this design? A bare RAG endpoint is easy to copy; a multi-agent system that verifies its own answers and ships as an MCP server is not. The architecture is the moat — "easy to buy, hard to replicate."
Architecture
flowchart LR
Q([Question]) --> P[🧭 Planner<br/>plan + search queries]
P --> R[📚 Retriever<br/>pgvector top-k]
R --> W[🌐 Web Researcher<br/>Firecrawl • optional]
W --> S[✍️ Synthesizer<br/>cited answer]
S --> C{🔎 Critic<br/>grounded?}
C -- needs revision --> S
C -- grounded --> A([Answer + citations])
subgraph Stores
DB[(Supabase<br/>pgvector)]
end
R <-->|cosine search| DB
classDef agent fill:#1e293b,stroke:#7C3AED,color:#e2e8f0;
class P,R,W,S,C agent;
| Agent | Model / tool | Responsibility |
|---|---|---|
| Planner | Claude (claude-opus-4-8, adaptive thinking) |
Decompose the question into focused search queries |
| Retriever | Voyage embeddings + pgvector | Cosine top-k over the knowledge base |
| Web Researcher | Firecrawl (optional) | Augment with live web results when a key is set |
| Synthesizer | Claude | Draft an answer grounded in context, with [n] citations |
| Critic | Claude | Verify grounding; loop back for revision if unsupported |
MCP tools
| Tool | Arguments | Returns |
|---|---|---|
ingest |
url: str |
Scrapes the URL, chunks + embeds it, stores it. { url, chunks_added } |
ask |
question: str |
Runs the full pipeline. { answer, citations, plan, grounded } |
search |
query: str, k: int = 5 |
Retrieval only — top-k chunks with similarity scores |
Quickstart
# 1. Install (Python 3.10+)
uv venv && uv pip install -e ".[dev]" # or: pip install -e ".[dev]"
# 2. Configure
cp .env.example .env # fill in ANTHROPIC_API_KEY, VOYAGE_API_KEY, DATABASE_URL
# 3. Create the vector table (Supabase SQL editor or psql)
psql "$DATABASE_URL" -f sql/schema.sql
# 4. Run the MCP server (stdio by default)
agentic-rag-mcp
Connect it to Claude Code
claude mcp add agentic-rag -s user \
--env ANTHROPIC_API_KEY=sk-ant-... \
--env VOYAGE_API_KEY=pa-... \
--env DATABASE_URL=postgresql://... \
-- agentic-rag-mcp
Then, from the client: "ingest https://example.com/docs" → "ask: how do I configure X?".
How it works
- Plan — Claude turns the question into a short plan + 1–5 search queries.
- Retrieve — each query is embedded (Voyage
voyage-3.5) and matched against pgvector by cosine distance; results are de-duplicated and ranked. - Research — if
FIRECRAWL_API_KEYis set, live web results are added to the context. - Synthesize — Claude writes an answer grounded only in the numbered context, citing
each claim as
[n]. - Critique — a strict fact-checker pass decides whether the answer is fully supported. If not (and revisions remain), it loops back to the synthesizer with feedback.
Configurable via env: RAG_MODEL, RAG_TOP_K, RAG_MAX_REVISIONS, RAG_EMBED_MODEL.
<p align="center"> <img src="docs/img/05-live-retrieval.png" alt="Live retrieval over pgvector" width="820"> </p> <p align="center"><sub>Live retrieval over pgvector — Voyage embeddings, real cosine similarity (illustrative demo corpus).</sub></p>
Evaluation
Answer quality is tracked with promptfoo — faithfulness, citation presence, and latency — so quality is measured, not asserted:
cd evals && promptfoo eval -c promptfooconfig.yaml
See evals/ for the rubric and test cases.
Deploy
Containerised and ready for Railway (HTTP transport):
railway up # uses Dockerfile + railway.json; set RAG_TRANSPORT=http
Expose RAG_HTTP_PORT and connect over --transport http. A cloudflared tunnel works for
local demos.
Project layout
src/agentic_rag_mcp/
config.py # env-driven settings
llm.py # Anthropic (Claude) helper — adaptive thinking, JSON parsing
embeddings.py # Voyage embeddings
store.py # pgvector store (psycopg)
web.py # Firecrawl web research (optional)
ingest.py # chunking + ingestion
state.py # LangGraph state
nodes.py # planner / retriever / researcher / synthesizer / critic
graph.py # graph assembly
server.py # FastMCP server (ingest / ask / search)
sql/schema.sql # pgvector schema
evals/ # promptfoo eval suite
License
MIT — see LICENSE.
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。