Research Guard
An MCP server that detects duplicated research directions and probes blind spots before committing to a project, helping researchers avoid wasted effort.
README
Research Guard
Detect duplicated research directions before you waste months on them.
An MCP server that makes your AI agent check "has this been done before, where are the pitfalls, what are my blind spots" before cheering you on.
Why?
The most painful cost for researchers is not money - it's time. "Halfway through, I found out someone already did it" is a nightmare almost every researcher has lived through.
Existing research AI tools (Elicit, Consensus, Semantic Scholar) help you find known information faster. Research Guard helps you discover what you don't know you don't know.
How it differs from similar projects
| Project | Type | Focus | Limitation |
|---|---|---|---|
| paper-search-mcp / semantic-scholar-mcp | MCP server | Multi-source paper search & download | Search only, no novelty assessment |
| idea_novelty_checker (AI2/MSR) | Research code | Novelty assessment benchmark (ACL 2025) | Paper reproduction code, not an installable tool |
| SciSpace Paper Novelty Check | Commercial SaaS | Pre-submission novelty check | Closed-source, paid, not agent-integrable |
| Research Guard | MCP server | Pre-project novelty check + blind spot probing | See Scope & limitations |
Methodology
The pipeline follows the literature-validated retrieve-then-rerank design (Shahid et al., Literature-Grounded Novelty Assessment of Scientific Ideas, SDP@ACL 2025, arXiv:2506.22026). Their ablations show that keyword-only retrieval flags non-novel ideas with just 5.2% accuracy, while facet-based re-ranking reaches 89.7% - which is exactly why Research Guard does not do naive keyword matching.
idea → ① contribution triplet decomposition → ② 5-dimension concept expansion
→ ③ multi-path retrieval (3 sources, concurrent) → ④ citation network expansion
→ ⑤ facet re-ranking (purpose/mechanism/evaluation/application)
→ ⑥ facet-grounded novelty assessment → ⑦ blind spot probing
→ RED / YELLOW / GREEN / UNKNOWN
Novelty definition (facet-based, from the paper above): an idea is novel if and only if it differs from all retrieved papers in at least one core facet (purpose, mechanism, or evaluation), or uniquely combines known facets / applies them to a new application domain.
Core capabilities
| Capability | Description |
|---|---|
| Novelty check | Decomposes the idea into a "problem + method + innovation" triplet and matches contribution structure, not text similarity |
| Facet re-ranking | LLM re-ranks candidates by facet priority, focusing on the most overlap-prone prior work |
| Concept expansion | 5-dimension terminology expansion: synonyms, paradigm upshifts, alternative methods, cross-domain transfer, known critiques |
| Citation network expansion | Forward citations + backward references to find related work that keyword search misses |
| Blind spot probing | Forced counter-evidence retrieval: surfaces approaches and negative results you haven't considered |
Data sources
Three built-in sources, all enabled by default and searched concurrently; results are deduplicated by paper id (papers indexed under different ids across sources may still appear more than once):
| Source | API key | Citation graph | Notes |
|---|---|---|---|
| OpenAlex | Not required | ✅ | 250M+ works; set OPENALEX_MAILTO to join the polite pool for faster responses |
| arXiv | Not required | ❌ | Preprint coverage |
| Semantic Scholar | Optional | ✅ | Rate-limited without a key; set SEMANTIC_SCHOLAR_API_KEY to lift it |
Select sources via the GUARD_SOURCES env var (comma-separated), e.g. GUARD_SOURCES=openalex,arxiv runs fully key-free.
Quick start
Install
Not yet published to PyPI - install from source:
git clone https://github.com/Fengrru/research-guard.git
cd research-guard
pip install -e packages/guard-core
pip install -e packages/guard-mcp
Configure environment variables
# Required: LLM API
export LLM_PROVIDER="openai"
export OPENAI_API_KEY="your_key"
export LLM_MODEL="gpt-4o-mini" # optional, defaults to gpt-4o-mini
# Optional: data source configuration
export GUARD_SOURCES="semantic_scholar,openalex,arxiv" # default: all
export SEMANTIC_SCHOLAR_API_KEY="your_key" # optional, rate-limited without it
export OPENALEX_MAILTO="you@example.com" # optional, speeds up OpenAlex
Use with Claude Desktop
Edit ~/Library/Application Support/Claude/claude_desktop_config.json (on Windows: %APPDATA%\Claude\claude_desktop_config.json):
{
"mcpServers": {
"research-guard": {
"command": "research-guard",
"env": {
"OPENAI_API_KEY": "your_key",
"OPENALEX_MAILTO": "you@example.com"
}
}
}
}
Example
Just tell Claude:
I want to use Transformers for time-series anomaly detection - check if anyone has done this
Claude calls the check_novelty tool and returns:
🔴 RED - High overlap
Found 3 highly related works:
1. "Anomaly Transformer" (2022) - same core method, differs in attention design
2. "AnomalyGPT" (2023) - combines anomaly detection with LLMs
3. ...
Suggested differentiation: time-series anomaly detection + LLM integration
remains under-explored
MCP tools
| Tool | Purpose | When to use |
|---|---|---|
check_novelty |
Novelty check (core) | Before committing to a direction |
expand_concepts |
Concept expansion | To understand the terminology space |
smart_search |
Smart literature search | For deep literature surveys |
explore_citation_network |
Citation network expansion | To dig deeper after finding seed papers |
find_blind_spots |
Blind spot probing | To fill gaps after initial retrieval |
Project layout
research-guard/
├── packages/
│ ├── guard-core/ # Core engine (usable standalone)
│ │ └── guard_core/
│ │ ├── sources/ # Built-in sources: OpenAlex / arXiv / Semantic Scholar
│ │ ├── reranker.py # Facet re-ranking (retrieve-then-rerank)
│ │ └── ...
│ └── guard-mcp/ # MCP server
├── plugins/ # Data source plugins (implement the DataSource protocol)
└── tests/
Scope & limitations
Important: assessments are based on indexed public literature. The following may be missed:
- Unpublished work (someone is working on it but hasn't published yet)
- Cross-language literature (Chinese, Japanese, German, and other non-English papers)
- Gray literature (patents, theses, technical reports)
- Work that is semantically equivalent but superficially very different
When retrieval or LLM assessment fails, the system returns an UNKNOWN rating with the reason listed in warnings - it never disguises a failure as GREEN (novel). If facet re-ranking fails, it degrades to retrieval order and says so in warnings.
The system presents facts, it does not make value judgments. The final decision always belongs to the researcher.
Development
# Clone
git clone https://github.com/Fengrru/research-guard.git
cd research-guard
# Install with dev dependencies
pip install -e "packages/guard-core[dev]"
pip install -e "packages/guard-mcp[dev]"
# Lint + tests
ruff check .
pytest
Contributing
Contributions welcome! See CONTRIBUTING.md.
Main directions:
- New data sources (PubMed, IEEE, CNKI, ... - just implement the
DataSourceprotocol) - Domain concept graphs (biomedicine, materials science, ...)
- Better novelty assessment (embedding pre-filtering, expert few-shot examples, ...)
- Share your "saved me months" cases
Citation
Methodology reference:
@inproceedings{shahid2025novelty,
title = {Literature-Grounded Novelty Assessment of Scientific Ideas},
author = {Shahid, Simra and Radensky, Marissa and Fok, Raymond and
Siangliulue, Pao and Weld, Daniel S. and Hope, Tom},
booktitle = {Proceedings of the Workshop on Scholarly Document Processing (SDP@ACL)},
year = {2025}
}
License
Apache-2.0
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。