Google Scholar MCP Server
Enables academic research through Google Scholar by searching for papers, finding author publications, discovering recent research, and identifying highly cited works through web scraping with natural language queries.
README
🔬 Google Scholar MCP Server
A Model Context Protocol (MCP) server that provides access to Google Scholar for academic research through web scraping. This server enables you to search for papers, find author publications, discover recent research, and identify highly cited works.
✨ Features
- 🔍 Paper Search: Search Google Scholar for academic papers with flexible filtering
- 👨🔬 Author Research: Find papers by specific authors
- 📅 Recent Papers: Discover recent publications in any field
- 🏆 Highly Cited Papers: Find influential papers with citation filtering
- ⏱️ Rate Limiting: Respectful scraping with built-in delays
- 🛡️ Error Handling: Robust error handling and logging
- 🌐 Local Web Interface: Optional Flask web interface for testing
- 🧠 Smart Query Processing: Natural language query processing with AI integration
🚀 Quick Start
Prerequisites
- Python 3.8 or higher
- pip (Python package manager)
Installation
-
Clone the repository
git clone https://github.com/yourusername/google-scholar-mcp.git cd google-scholar-mcp -
Install dependencies
pip install -r requirements.txt -
Optional: Set up environment variables
cp env.example .env # Edit .env with your preferred settings
Running the MCP Server
Run the MCP server for use with MCP clients:
python main.py
Testing with Local Web Interface
For testing and development, you can run the local web interface:
python local_server.py
Then open your browser to http://localhost:5000
🔧 Configuration
The server can be configured through environment variables. Copy env.example to .env and modify as needed:
# Request delay between Google Scholar requests (seconds)
REQUEST_DELAY=5
# Maximum results per request
MAX_RESULTS_PER_REQUEST=20
# HTTP timeout (seconds)
TIMEOUT=15
Available Tools
1. search_papers
Search for academic papers on Google Scholar.
Parameters:
query(required): Search query for papersnum_results(optional): Number of results to return (1-20, default: 10)start_year(optional): Earliest publication year to includeend_year(optional): Latest publication year to include
Example:
{
"query": "machine learning neural networks",
"num_results": 15,
"start_year": 2020,
"end_year": 2024
}
2. get_author_papers
Search for papers by a specific author.
Parameters:
author_name(required): Name of the author to search fornum_results(optional): Number of results to return (default: 10)
Example:
{
"author_name": "Geoffrey Hinton",
"num_results": 20
}
3. search_recent_papers
Search for recent papers in a specific field.
Parameters:
field(required): Research field or topicyears_back(optional): How many years back to search (1-10, default: 2)num_results(optional): Number of results to return (default: 10)
Example:
{
"field": "quantum computing",
"years_back": 3,
"num_results": 15
}
4. get_highly_cited_papers
Search for highly cited papers in a topic.
Parameters:
topic(required): Research topic or fieldmin_citations(optional): Minimum number of citations (default: 100)num_results(optional): Number of results to return (default: 10)
Example:
{
"topic": "transformer neural networks",
"min_citations": 500,
"num_results": 10
}
Response Format
Each tool returns a JSON response with paper information including:
title: Paper titleauthors: Author namesurl: Link to the paperyear: Publication yearsnippet: Paper abstract/description snippetcited_by: Number of citations (when available)pdf_url: Direct PDF link (when available)publication_info: Journal/conference information
Rate Limiting and Ethics
This server implements respectful scraping practices:
- 2-second delays between requests
- Proper User-Agent headers
- Error handling for rate limits
- Designed for research and educational purposes
🔍 MCP Client Integration
Claude Desktop
Add this to your Claude Desktop configuration (~/Library/Application Support/Claude/claude_desktop_config.json on macOS):
{
"mcpServers": {
"google-scholar": {
"command": "python",
"args": ["/path/to/google-scholar-mcp/main.py"],
"cwd": "/path/to/google-scholar-mcp"
}
}
}
Other MCP Clients
The server follows the standard MCP protocol and should work with any MCP-compatible client.
🧠 Smart Query Processing
The server includes intelligent query processing that can understand natural language requests:
# Example natural language queries:
"Find recent computer vision papers from CVPR 2023"
"Show me highly cited papers by Geoffrey Hinton"
"What are the latest developments in quantum computing?"
📊 Response Format
All tools return structured JSON with paper information:
{
"title": "Paper Title",
"authors": "Author Names",
"url": "Link to paper",
"year": 2023,
"snippet": "Abstract excerpt...",
"cited_by": 150,
"pdf_url": "Direct PDF link",
"publication_info": "Journal/Conference"
}
⚖️ Legal and Ethical Considerations
- 🎓 Educational Use: This tool is intended for research and educational purposes
- 📜 Terms of Service: Respect Google Scholar's terms of service
- 🤝 Responsible Use: Use responsibly and avoid excessive requests
- 🔌 Official APIs: Consider using official APIs when available
- 📚 Copyright: Be mindful of copyright and fair use policies
🔧 Troubleshooting
Common Issues
- Rate Limiting: If you get blocked, wait and reduce request frequency
- Network Errors: Check your internet connection
- Parsing Errors: Google Scholar may change their HTML structure
- Import Errors: Make sure all dependencies are installed
Debug Mode
Enable debug logging by setting DEBUG=true in your .env file.
Logging
The server includes detailed logging. Check the console output for error messages and debugging information.
📦 Dependencies
mcp: Model Context Protocol libraryrequests: HTTP library for web scrapingbeautifulsoup4: HTML parsinglxml: XML/HTML parserurllib3: HTTP clientflask: Web interface (optional)
🤝 Contributing
Contributions are welcome! Please ensure:
- Respectful scraping practices
- Error handling for edge cases
- Clear documentation
- Testing with various queries
- Follow the existing code style
Development Setup
# Clone the repository
git clone https://github.com/yourusername/google-scholar-mcp.git
cd google-scholar-mcp
# Install dependencies
pip install -r requirements.txt
# Run tests
python test_server.py
python test_query_processor.py
# Run local development server
python local_server.py
📄 License
This project is licensed under the MIT License - see the LICENSE file for details.
🙏 Acknowledgments
- Built on the Model Context Protocol by Anthropic
- Inspired by the need for accessible academic research tools
- Thanks to the open-source community for the excellent libraries used
⚠️ Disclaimer
This tool is for educational and research purposes. Please respect Google Scholar's terms of service and use responsibly. The authors are not responsible for any misuse of this tool.
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。