AI Research Paper Assistant
Enables asking questions about foundational AI/ML research papers, with retrieval-augmented generation and an agentic layer for complex queries, via MCP-compatible clients.
README
AI Research Paper Assistant
A Retrieval-Augmented Generation (RAG) system that answers questions about foundational AI/ML research papers, with an agentic reasoning layer, automated evaluation, and an MCP (Model Context Protocol) server so any MCP-compatible LLM client can use it as a tool.
What it does
Ask questions like:
- "What is the attention mechanism in Transformers?"
- "How does RAG combine retrieval with generation?"
- "Compare how the Transformer paper and ReAct paper approach reasoning"
The system retrieves relevant passages from three indexed papers — Attention Is All You Need, Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, and ReAct: Synergizing Reasoning and Acting in Language Models — and generates a grounded answer, citing the source paper and page.
For complex, multi-part questions, an agent layer breaks the question into sub-questions, retrieves context for each, and synthesizes a combined answer — rather than relying on a single retrieval pass.
Architecture
┌─────────────┐
│ papers/ │ Raw PDFs (downloaded from arXiv)
└──────┬──────┘
│ ingest.py: load → chunk → embed
▼
┌─────────────┐
│ chroma_db/ │ Vector store (291 chunks, all-MiniLM-L6-v2 embeddings)
└──────┬──────┘
│
├─────────────────┬─────────────────────┐
▼ ▼ ▼
┌─────────────┐ ┌──────────────┐ ┌────────────────────┐
│ rag_chain.py│ │ agent.py │ │ evaluation.py │
│ Simple LCEL │ │ LangGraph: │ │ RAGAS metrics: │
│ retrieve→ │ │ plan→ │ │ faithfulness, │
│ generate │ │ retrieve→ │ │ answer relevancy, │
│ │ │ synthesize │ │ context precision │
└──────┬──────┘ └──────┬───────┘ └─────────────────────┘
│ │
└───────┬────────┘
▼
┌─────────────────┐
│ mcp_server.py │ Exposes both as MCP tools
│ + app.py (CLI) │ for external LLM clients
└─────────────────┘
Why this shape: all shared setup (embeddings, vector store connection, LLM client) lives in one place, src/config.py. Every other module imports from it instead of redefining it — so changing the embedding model or LLM provider means editing one file, not four.
Tech stack and why each piece was chosen
| Component | Choice | Why |
|---|---|---|
| Orchestration | LangChain (LCEL) | Industry-standard way to compose retrieval + prompt + LLM into a pipeline |
| Agent framework | LangGraph | Lets the system plan multi-step retrieval instead of one fixed pass — needed for comparison-style questions |
| Vector store | Chroma | Local, free, simple to set up and inspect for a single-developer project |
| Embeddings | HuggingFace all-MiniLM-L6-v2 |
Runs locally on CPU, free, no API key or vendor lock-in — deliberate alternative to OpenAI embeddings |
| LLM | Groq (Llama 3.3 70B) | Free tier, very fast inference, good enough quality for this scope |
| Evaluation | RAGAS | Automated, LLM-judged scoring instead of just eyeballing answers |
| Tool exposure | MCP (Model Context Protocol) | Turns the pipeline into a callable tool for any MCP client (e.g. Claude Desktop), not just a terminal script |
Evaluation results
Measured with RAGAS on a 3-question test set against known ground-truth answers:
| Metric | Score | What it measures |
|---|---|---|
| Faithfulness | 0.87 | Does the answer stick to what's actually in the retrieved context, or hallucinate? |
| Answer Relevancy | 0.93 | Does the answer actually address the question asked? |
| Context Precision | 0.94 | Were the most relevant chunks ranked near the top of retrieval? |
Project structure
ai-research-paper-assistant/
├── src/
│ ├── config.py # Single source of truth: embeddings, vectorstore, LLM setup
│ ├── ingest.py # PDF loading, chunking, embedding into Chroma
│ ├── rag_chain.py # Simple retrieve-then-generate LCEL chain
│ ├── agent.py # LangGraph agent: plan → retrieve → synthesize
│ └── evaluation.py # RAGAS evaluation harness
├── mcp_server.py # MCP tool wrapper around rag_chain and agent
├── app.py # CLI entry point (ingest / ask / agent / evaluate)
├── download_papers.py # One-time script to pull papers from arXiv
├── requirements.txt
└── eval_results.txt # Saved RAGAS output
Setup
# 1. Clone and enter the project
git clone <your-repo-url>
cd ai-research-paper-assistant
# 2. Create and activate a virtual environment
python -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
# 3. Install dependencies
pip install -r requirements.txt
# 4. Add your Groq API key (free at console.groq.com)
echo "GROQ_API_KEY=your_key_here" > .env
# 5. Download the papers and build the vector store
python download_papers.py
python app.py ingest
Known dependency issue: ragas==0.3.9 ships with a broken import (langchain_community.chat_models.vertexai, a module path that's since moved). If you hit ModuleNotFoundError on that import when running evaluate, it's a known upstream bug, not a problem with this code — open an issue or check for a patched ragas release. A workaround is to wrap that one import line in the installed package in a try/except ImportError, since this project doesn't use VertexAI at all.
Usage
python app.py ask # Simple RAG: ask direct questions, one retrieval pass
python app.py agent # Agentic RAG: handles multi-part/comparison questions
python app.py evaluate # Re-run RAGAS scoring, saves to eval_results.txt
python mcp_server.py # Start as an MCP server for external clients (e.g. Claude Desktop)
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。