DocMind
An MCP-native document research agent that provides search, summarization, and citation tools for iterative multi-step research, ending in validated structured JSON reports.
README
DocMind — MCP-Native Document Research Agent
An iterative, multi-step research agent built on LangGraph state graphs and LangChain retrieval chains, with search, summarization, and citation exposed as a real MCP (Model Context Protocol) server — so those tools are consumable by any MCP-compatible client, not locked to this agent. Retrieval combines hybrid search (vector embeddings + keyword) and every research session ends in a validated structured JSON report, not free text.
Runs on a BYOK (Bring Your Own Key) model — you supply your own
Anthropic or OpenAI API key via a local .env file. DocMind never ships,
stores, or transmits your key anywhere except directly to your chosen LLM
provider.
Why this is a research agent, not just RAG
A single retrieve-then-answer pass often isn't enough for open-ended questions: the first search may surface part of the answer while missing another part entirely. DocMind's graph is a loop, not a pipeline:
┌──────┐ ┌────────┐ ┌──────────┐
│ plan │ ───▶ │ search │ ───▶ │ evaluate │
└──────┘ └───┬────┘ └────┬─────┘
▲ │
│ "search" (gap │ "synthesize" (evidence
│ found, budget │ sufficient, or budget
│ remains) │ exhausted)
└─────────────────┤
▼
┌─────────────┐
│ synthesize │ ──▶ END
└─────────────┘
planturns the raw question into a focused first search query.searchcalls the MCPsearch_documentstool and accumulates evidence across iterations (deduplicated by document + chunk).evaluate— an LLM call — judges whether the evidence gathered so far actually answers the question. If not, it proposes a refined query and the graph loops back tosearch, up toMAX_RESEARCH_ITERATIONS.synthesizeproduces the final answer as a validatedResearchReportobject via LangChain'swith_structured_output— never a raw string a downstream service has to parse.
MCP server: search, summarize, cite
app/mcp_server/server.py exposes three tools over the corpus:
| Tool | Purpose |
|---|---|
search_documents(query, top_k) |
Hybrid vector+keyword search across the whole corpus |
summarize_document(source, focus) |
LLM summary of one named document, optionally focused on an aspect |
cite_passage(claim, source) |
Finds the best-matching passage within one document that supports a claim — the citation/grounding primitive |
Because these are MCP tools rather than plain Python functions, any
MCP-compatible client can connect to this same server and use them —
including, for example, Claude Desktop, once pointed at
python -m app.mcp_server.server. The LangGraph agent shipped in this
repo deliberately consumes its own server the same way an external client
would (app/mcp_client/tools.py), rather than importing the retrieval
module directly — that's what makes it "MCP-native" end to end.
Hybrid retrieval
app/retrieval/hybrid_search.py fuses Postgres full-text keyword search
with pgvector cosine-similarity search (HNSW index), normalizing and
weighting both signals (HYBRID_ALPHA in .env). The fusion logic is a
pure function, unit-tested without needing a live database.
BYOK — Bring Your Own Key
- Set
LLM_PROVIDER=anthropicoropenaiin.env. - Set the matching
ANTHROPIC_API_KEYorOPENAI_API_KEY. app/config.pyraises an explicit, readable error if the selected provider's key is missing.- Local embeddings (
sentence-transformers) run entirely on your machine and need no API key — only the plan/evaluate/synthesize LLM calls and thesummarize_documenttool use your BYOK key.
Project structure
DocMind/
├── app/
│ ├── config.py # BYOK settings, loaded from .env
│ ├── llm.py # LLM factory (Anthropic / OpenAI)
│ ├── schemas.py # ResearchReport structured-output schema
│ ├── main.py # CLI entry point
│ ├── api.py # optional FastAPI /research endpoint
│ ├── graph/
│ │ ├── state.py # ResearchState (question, evidence, iteration...)
│ │ ├── nodes.py # plan / search / evaluate / synthesize
│ │ └── build_graph.py # wires the iterative loop
│ ├── retrieval/
│ │ ├── embeddings.py # local sentence-transformers embeddings
│ │ ├── vector_store.py # documents + corpus_chunks (pgvector)
│ │ └── hybrid_search.py # score fusion (unit-testable, pure fn)
│ ├── mcp_server/
│ │ └── server.py # FastMCP: search / summarize / cite tools
│ └── mcp_client/
│ └── tools.py # discovers + calls MCP tools at runtime
├── data/corpus/ # sample research docs (RAG, vector DBs, agents)
├── scripts/
│ ├── init_db.sql # pgvector schema (documents + corpus_chunks)
│ └── seed_corpus.py # paragraph-aware chunking + embedding + insert
├── tests/ # pure-logic unit tests (no live DB/LLM needed)
├── examples/example_research_session.md
├── docker-compose.yml # local Postgres + pgvector (port 5433)
├── requirements.txt
└── .env.example
Setup
1. Clone and install dependencies
git clone <your-fork-url>
cd DocMind
python -m venv venv && source venv/bin/activate # Windows: venv\Scripts\activate
pip install -r requirements.txt
2. Configure your BYOK key
cp .env.example .env
# edit .env — set LLM_PROVIDER and the matching API key
3. Start Postgres + pgvector
docker-compose up -d
4. Seed the corpus
python -m scripts.seed_corpus
5. Run DocMind
python -m app.main
Or as an API:
uvicorn app.api:app --reload
# POST http://localhost:8000/research {"question": "How does HNSW work?"}
Inspecting the MCP server directly (useful for demoing that these tools are consumable by any MCP client, not just this agent):
mcp dev app/mcp_server/server.py
See examples/example_research_session.md
for a full sample session showing the loop refine itself across two
iterations.
Running tests
pytest
Covers the pure-logic pieces — score fusion and loop routing — that don't require a live database or LLM call, so they run in any environment, including CI.
Tech stack
| Layer | Technology |
|---|---|
| Agent orchestration | LangGraph (iterative state graph, conditional looping, checkpointer) |
| Retrieval chains | LangChain |
| Tool exposure | MCP (Model Context Protocol) — mcp Python SDK |
| Structured output | LangChain with_structured_output + Pydantic |
| Vector store | PostgreSQL + pgvector (HNSW index) |
| Embeddings | sentence-transformers, local, CPU-only |
| LLM | Anthropic Claude or OpenAI (BYOK, user-selected) |
| API | FastAPI (optional) |
Known limitations
- The MCP client opens a fresh stdio session per tool call rather than a persistent connection — simple and reliable for a research workload with a handful of tool calls per session, not tuned for high frequency.
evaluate_node's sufficiency judgment is itself an LLM call and can be wrong in either direction (stopping early or looping unnecessarily);MAX_RESEARCH_ITERATIONSbounds the cost of the latter.- The sample corpus is small and topic-specific (RAG, vector databases, agent architectures) — meant to demonstrate the pipeline end to end, not as a production knowledge base.
License
MIT — see LICENSE.
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。
mcp-server-qdrant
这个仓库展示了如何为向量搜索引擎 Qdrant 创建一个 MCP (Managed Control Plane) 服务器的示例。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。