Document Assistant
Enables natural language search and analysis of uploaded PDF, CSV, and Excel documents using retrieval-augmented generation and MCP tools, providing contextual answers to user queries.
README
🤖 AI Document Scanner Using RAG & MCP Tools with Python
An intelligent AI Document Scanner and Assistant built with Python that allows users to upload documents, search their contents, and ask questions using Retrieval-Augmented Generation (RAG) and Model Context Protocol (MCP) tools.
The application combines document processing, text chunking, vector embeddings, FAISS similarity search, MCP tools, and Large Language Models (LLMs) to provide contextual answers from uploaded documents.
📌 Project Overview
Traditional document search often depends on exact keyword matching. This project uses semantic search through vector embeddings, allowing users to ask questions naturally.
The system:
- Accepts documents from the user.
- Extracts text from the documents.
- Splits the text into smaller chunks.
- Converts chunks into vector embeddings.
- Stores embeddings in a FAISS vector database.
- Retrieves relevant chunks when the user asks a question.
- Uses an LLM to generate an answer based on the retrieved context.
- Uses MCP tools to expose document search and analysis capabilities to an AI agent.
✨ Features
- 📄 PDF document processing
- 📊 CSV and Excel data analysis
- 🔍 Semantic document search
- 🧠 Retrieval-Augmented Generation (RAG)
- 🗂️ FAISS vector database
- 🤖 LLM-powered question answering
- 🔌 Model Context Protocol (MCP) tool integration
- 💬 Interactive document chat
- 🌐 Streamlit web interface
- 📑 Document chunking and embeddings
- 🔎 Context-aware information retrieval
- 🧩 Modular project architecture
- 🔐 Environment-variable based API key configuration
🏗️ Architecture
┌─────────────────────┐
│ User │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ Streamlit UI │
│ app.py │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ Document Loader │
│ PDF / CSV / Excel │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ Text Splitting & │
│ Preprocessing │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ Embedding Model │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ FAISS Vector Store │
└──────────┬──────────┘
│
User Question
│
▼
┌─────────────────────┐
│ Semantic Retrieval │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ MCP Tools │
│ PDF / CSV / Excel │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ RAG Pipeline │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ LLM │
│ Gemini / Other LLM │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ AI Response │
└─────────────────────┘
🔄 RAG Workflow
1. Document Loading
Documents are loaded using appropriate Python libraries.
- PDF →
PyPDF - Excel →
Pandas/OpenPyXL - CSV →
Pandas
2. Text Splitting
Large documents are divided into smaller chunks using a text splitter. This improves retrieval accuracy and keeps prompts manageable.
3. Embeddings
Each text chunk is converted into a numerical vector representing its semantic meaning.
Document Text
|
v
Embedding Model
|
v
Numerical Vector
4. Vector Storage
The generated vectors are stored in FAISS, enabling efficient similarity-based retrieval.
5. Semantic Search
When the user asks a question, the question is converted into an embedding and compared against stored document vectors. The most relevant chunks are retrieved.
6. Response Generation
The retrieved context is passed to the LLM along with the user's question. The model generates an answer based on the retrieved information.
🔌 MCP Integration
The project uses Model Context Protocol (MCP) to expose document-related functionality as tools that an AI agent can call.
Example MCP Tools
search_pdf(question)
find_employee(name)
analyze_csv(question)
This allows an AI system to retrieve information from documents or analyze structured data when required.
Example MCP Workflow
User
|
v
"Find the details of Manoj Sarkar."
|
v
AI Agent
|
v
MCP Tool
find_employee("Manoj Sarkar")
|
v
Vector / Document Search
|
v
Relevant Information
|
v
LLM
|
v
Final Answer
🛠️ Technologies Used
| Technology | Purpose |
|---|---|
| Python | Core programming language |
| Streamlit | Web application interface |
| LangChain | RAG and document processing |
| FAISS | Vector similarity search |
| MCP | AI tool integration |
| Google Gemini | Large Language Model |
| PyPDF | PDF text extraction |
| Pandas | Data processing |
| OpenPyXL | Excel file processing |
| Vector Embeddings | Semantic representation |
| python-dotenv | Environment variable management |
📁 Project Structure
AI-Document-Scanner/
│
├── app.py
├── server.py
├── client.py
├── requirements.txt
├── README.md
│
├── assistants/
│ ├── pdf_assistant.py
│ ├── excel_assistant.py
│ ├── csv_assistant.py
│ └── docx_assistant.py
│
├── utils/
│ ├── loaders.py
│ ├── splitter.py
│ ├── vectorstore.py
│ ├── rag_pipeline.py
│ ├── llm_provider.py
│ └── prompts.py
│
├── mcp_tools/
│ ├── pdf_tool.py
│ ├── rag_tool.py
│ ├── excel_tool.py
│ └── csv_tool.py
│
├── data/
│ └── sample_documents/
│
└── .env
The exact files and folders may vary depending on the current implementation.
⚙️ Installation
1. Clone the Repository
git clone https://github.com/your-username/AI-Document-Scanner.git
cd AI-Document-Scanner
2. Create a Virtual Environment
Windows
python -m venv venv
venv\Scripts\activate
Linux/macOS
python3 -m venv venv
source venv/bin/activate
3. Install Dependencies
pip install -r requirements.txt
🔑 Environment Variables
Create a .env file in the project root.
Google Gemini
GOOGLE_API_KEY=your_google_api_key
If your implementation also supports OpenAI:
OPENAI_API_KEY=your_openai_api_key
Important: Never commit your
.envfile or API keys to GitHub.
Add the following to .gitignore:
.env
venv/
__pycache__/
*.pyc
.faiss/
▶️ Running the Application
Start the Streamlit application:
streamlit run app.py
Then open the application in your browser:
http://localhost:8501
💡 Example Questions
After uploading a document, users can ask questions such as:
What is this document about?
Summarize the document.
Who is Manoj Sarkar?
Find the employee with the highest sales.
What is the total revenue?
What are the main points discussed in the document?
Find information related to a specific topic.
🧪 Example RAG Pipeline
documents = load_documents(file_path)
chunks = split_documents(documents)
vectorstore = create_vectorstore(chunks)
results = vectorstore.similarity_search(query, k=4)
context = "\n".join(
document.page_content
for document in results
)
response = llm.invoke(
f"""
Answer the question using the following context:
{context}
Question:
{query}
"""
)
🔌 Example MCP Tool
A simplified MCP tool can look like:
from mcp.server.fastmcp import FastMCP
mcp = FastMCP("Document Assistant")
@mcp.tool()
def search_pdf(question: str) -> str:
"""Search the uploaded PDF and return relevant information."""
# Vector search implementation
return "Relevant document information"
The MCP server exposes this functionality so that an AI client or agent can use it when required.
🎯 Use Cases
This project can be useful for:
- 📚 Research document assistants
- 🏢 Company knowledge bases
- 📄 Legal document search
- 🎓 Educational document analysis
- 👨💼 HR document assistants
- 📊 Business report analysis
- 🧾 Invoice and report processing
- 📑 Policy and documentation search
- 🤖 AI-powered knowledge management systems
🚀 Future Enhancements
- [ ] Multi-document conversational memory
- [ ] DOCX support
- [ ] Image document scanning
- [ ] OCR integration
- [ ] Voice input
- [ ] Source/page citations
- [ ] Chat history
- [ ] User authentication
- [ ] Multi-user support
- [ ] Cloud deployment
- [ ] Advanced agentic workflows
- [ ] Additional MCP tools
- [ ] Database integration
- [ ] Document summarization
- [ ] Hybrid keyword + semantic search
- [ ] Reranking for improved retrieval accuracy
🔐 Security
For security:
- Store API keys in
.env. - Never upload API keys to GitHub.
- Add
.envto.gitignore. - Avoid storing sensitive documents in public repositories.
- Validate uploaded files before processing.
🧠 Key Concepts Demonstrated
This project demonstrates practical knowledge of:
- Python
- Generative AI
- Large Language Models
- Retrieval-Augmented Generation
- Vector Databases
- Semantic Search
- Embeddings
- LangChain
- FAISS
- Model Context Protocol
- AI Agents
- Streamlit
- Document Processing
- API Integration
👨💻 Author
Manoj Sarkar
B.Tech in Computer Science & Engineering
Interested in Python, Generative AI, Agentic AI, RAG, MCP, and AI-powered applications.
⭐ Support
If you find this project useful, consider giving the repository a ⭐ on GitHub.
📄 License
This project is intended for educational and development purposes. Add an appropriate license file if you plan to distribute or reuse the project publicly.
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。