Polymath MCP

Polymath MCP

Aggregates 17 free, keyless research sources into a single MCP server, enabling unified search for papers, code, models, trends, standards, and electronic components with deduplicated results.

Category
访问服务器

README

Polymath MCP

An MCP (Model Context Protocol) server that aggregates free, keyless research sources for computer science, computer engineering, electronics and IT.

One server, one normalized output format, seventeen data sources. Every source works without an API key and without payment. Results from different sources are merged and deduplicated by DOI and arXiv id, so the same paper found on arXiv, Crossref and OpenAlex comes back as a single record that combines the best fields of each.

Why

Research data is scattered across many APIs, each with its own format, quirks and rate limits. Existing MCP servers each cover one slice. Polymath gives an agent a single interface where a paper, its open-access PDF, its citation count, its implementation on GitHub, the model trained from it on Hugging Face and the practical discussion around it on Hacker News or Stack Exchange are all one tool call apart.

Tools

Tool What it does Sources
search_papers Search academic papers, deduplicated across sources arXiv, Crossref, DBLP, Europe PMC, OpenAlex, CORE, Zenodo, Semantic Scholar, Hugging Face Papers
get_paper Resolve one paper by DOI or arXiv id, merging every source that knows it, including open-access PDF link and citation counts the above, plus Unpaywall and OpenCitations
search_code Find repositories, e.g. the implementation of a paper GitHub
search_hub Search models or datasets, with links back to their papers Hugging Face Hub
search_trends Practitioner discussions and relevance signals Hacker News, Stack Overflow, Electronics Stack Exchange
search_standards RFCs and Internet-Drafts IETF Datatracker
search_components Electronic component documents: scanned datasheets and KiCad symbols Internet Archive, KiCad libraries
list_sources List every source and its category -

Every search response is an envelope with results, sources_ok and sources_failed. Partial failure is the normal case for an aggregator: when a source is down or rate-limited, the others still answer and the envelope says exactly which sources contributed.

Installation

Requires Python 3.11+.

pip install git+https://github.com/Moreti2002/polymath-mcp

Or from a clone:

git clone https://github.com/Moreti2002/polymath-mcp
cd polymath-mcp
pip install .

Claude Code

claude mcp add polymath -- polymath-mcp

Claude Desktop / other MCP clients

{
  "mcpServers": {
    "polymath": {
      "command": "polymath-mcp"
    }
  }
}

Configuration

No API keys are needed. One optional environment variable:

  • POLYMATH_EMAIL: a real contact e-mail. Unlocks the Unpaywall source (which requires it) and puts Crossref requests in its polite pool with better rate limits. Without it everything else still works.
{
  "mcpServers": {
    "polymath": {
      "command": "polymath-mcp",
      "env": { "POLYMATH_EMAIL": "you@example.com" }
    }
  }
}

Sources and their limits

All limits below are for keyless access, as observed in practice.

Source Role Keyless limit
arXiv CS/EE preprints, primary 1 request / 3 s (self-imposed politeness)
Crossref DOI metadata, primary 3 req/s in the polite pool
DBLP Canonical CS bibliography (venues) ~1 req/s, aggressive blocking on bursts
Europe PMC Full-text biomedical + applied ML ~10 req/s
OpenAlex Enrichment by DOI (free); text search is credit-rationed (~100 searches/day) 1000 credits/day/IP
CORE Institutional repository full-text ~10 req/min
Zenodo Software and datasets with DOIs 30 req/min
Semantic Scholar Citation graph, best effort only shared anonymous pool, frequent 429
Unpaywall Open-access PDF resolution by DOI 100k/day, requires POLYMATH_EMAIL
OpenCitations Citation counts and references unthrottled, high latency
GitHub Repository search 10 req/min
Hugging Face Models, datasets, papers generous, undocumented
Hacker News (Algolia) Discussion search ~10k req/h
Stack Exchange Stack Overflow + Electronics SE 300 req/day/IP
IETF Datatracker RFCs and drafts unthrottled
Internet Archive Scanned datasheets and databooks unthrottled
KiCad libraries (GitLab) Component symbol existence unthrottled

The server caches responses in memory and throttles each source to stay within these limits. Sources with hard daily budgets (Stack Exchange, OpenAlex search) are cached the longest.

Known coverage gaps

  • There is no legitimate keyless API for commercial electronic component catalogs (Octopart/Nexar, Digi-Key and Mouser all require credentials). search_components covers what keyless sources can: scanned datasheets on the Internet Archive and KiCad symbol libraries. For part selection discussion, search_trends includes the Electronics Stack Exchange.
  • Papers with Code shut down in 2025; its role is covered by the Hugging Face Papers source.
  • Semantic Scholar without a key shares a global anonymous pool and fails often. It is wired as best-effort and never blocks a search.

Architecture

src/polymath/
  server.py        MCP server and tool definitions
  models.py        normalized result models (Paper, CodeRepo, HubItem, ...)
  aggregate.py     concurrent fan-out with per-source failure isolation
  dedupe.py        cross-source paper merging (DOI, arXiv id, title)
  http.py          shared HTTP client, retries, per-source throttling
  cache.py         in-memory TTL cache
  providers/       one module per source, registered in a plug-in registry

Adding a source is one file: subclass Provider (or PaperLookupProvider), map the API response to the normalized models, decorate the class with @register. The aggregator, dedupe and MCP tools pick it up automatically.

Development

python -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"

pytest -m "not live"   # offline tests (mocked HTTP)
pytest -m live         # integration tests against the real APIs

Offline tests must always pass. Live tests depend on third-party services and may fail when a source is down or rate-limited; CI runs them as non-blocking.

License

MIT. See LICENSE.

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选