paper-mcp

paper-mcp

Remotely-callable MCP server for academic paper search, full-text retrieval and image to LaTeX conversion across arXiv, Semantic Scholar, and OpenAlex.

Category
访问服务器

README

paper-mcp

<!-- mcp-name: io.github.MCPServings/paper-mcp -->

Remotely-callable MCP server for academic paper search, full-text retrieval & image→LaTeX, served at https://latex-tools.online/mcp.

Three corpora behind one normalized interface:

  • arxiv (default) — search, metadata, and full-text (HTML / markdown / LaTeX source)
  • semanticscholar (alias s2) — the full S2 API surface: citation graph, authors, recommendations, full-text snippets, bulk datasets
  • openalex (alias oa) — 316M all-field works: citation graph, authors with h-index, institutions, topics, influence metrics

Plus a unified search_all that fuses all three corpora, image→LaTeX OCR, and LaTeX lint + PDF→text tooling.


Tools (41)

Generic / source-agnostic (8)

Tool Purpose
search_all(query, max_results=10, sources='arxiv,semanticscholar,openalex') Unified search. Fans out to all three corpora concurrently, de-duplicates the same work (by DOI/title) and re-ranks with Reciprocal Rank Fusion. Each hit carries sources (who found it) + an ids map for follow-up calls. Prefer this for broad lookups.
search_papers(query, source='arxiv', max_results=10, sort_by='relevance') Single-corpus search. arXiv query accepts plain text or field syntax (ti: au: cat:cs.CL abs: + AND/OR).
get_paper(paper_id, source='arxiv') One paper's full record. S2 id accepts S2 id / DOI: / ARXIV: / CorpusId:.
search_by_author(author, source='arxiv') Papers by author, newest first.
list_recent(category, source='arxiv') Latest in a category (arXiv code or S2 field of study).
list_categories(source='arxiv') Common category codes.
read_paper(paper_id, format='markdown') FULL text (arXiv). markdown = body with formulas as $LaTeX$; html = raw LaTeXML page; latex = original manuscript .tex source.
list_paper_sources() Available corpora.

read_paper fetch chain: arxiv.org/html/{id}ar5iv fallback (markdown/html), or arxiv.org/e-print/{id} tarball main .tex (latex). Formulas are recovered from the LaTeXML alttext invariant.

Medical / evidence-graded (1)

Tool Purpose
search_medical(query, study_types='rct,meta-analysis,systematic-review', year_from=0, max_results=10, fetch_fulltext=True) Clinical literature search. Queries PubMed, filters by research type via Publication-Type tags and re-ranks by the evidence pyramid (meta-analysis / systematic review > RCT > cohort > ...), so real trials surface above high-cited reviews/guidelines that pure-citation ranking floats up. Open-access full text is attached from Europe PMC by PMID. If the type filter yields nothing it auto-relaxes (flagged filter_relaxed). query is English keyword/boolean text — do NL/multilingual query understanding upstream. Backed by NCBI E-utilities + Europe PMC (both free, no key required).

Image → LaTeX (3)

Turn a formula or table image back into LaTeX (e.g. a figure cropped from a paper) without needing your own vision model. Backed by the co-located recognize service (PaddleOCR-VL / DeepSeek-OCR / texify).

Tool Purpose
recognize_formula(image_url=... or image_base64=..., model='deepseek-ocr') Formula image → LaTeX. image_url is downloaded server-side (with SSRF guards). Returns {latex, model, elapsed_ms}.
recognize_table(image_url=... or image_base64=..., model='deepseek-ocr') Table image → LaTeX tabular.
list_ocr_models() Available OCR models (deepseek-ocr, paddleocr-vl, texify).

LaTeX tooling (3)

Companions to the LaTeX/PDF web tools at latex-tools.online — same backends, exposed over MCP.

Tool Purpose
lint_latex(code) Check a LaTeX snippet for errors and return an auto-fixed version. Returns {errors, fixed_code, summary_en, summary_zh, elapsed_ms}.
extract_pdf(pdf_url=... or pdf_base64=..., formula=True, table=True) PDF → clean Markdown/LaTeX text via MinerU (useful for papers with no open-access full text). pdf_url is downloaded server-side (SSRF-guarded). Content-addressed + cached: a recently-seen or small PDF returns content in one call; a fresh PDF (MinerU is GPU-heavy, minutes) returns status='running' + a task_id.
extract_pdf_result(task_id) Fetch an extract_pdf job by task_id. Returns content once status='done'; while 'running', content is null — call again shortly.

OpenAlex (8)

  • Works: get_openalex_work · get_openalex_citations · get_openalex_references · search_openalex_works (filters: year range, open-access, min-citations, institution)
  • Authors/Institutions: search_openalex_authors · search_openalex_institutions
  • Analytics: get_openalex_trends · list_openalex_topics

Semantic Scholar (18)

  • Graph: get_paper_citations · get_paper_references · get_paper_authors
  • Lookup: match_paper_title · autocomplete_papers
  • Bulk: search_papers_bulk (≤1000, sortable, token paging) · get_papers_batch
  • Authors: search_authors · get_author · get_author_papers · get_authors_batch
  • Full-text: search_snippets (search inside paper body)
  • Recommend: recommend_papers_for_paper · recommend_papers_from_examples
  • Datasets: list_dataset_releases · get_dataset_release · get_dataset_download_links · get_dataset_diffs

Layout

paper_mcp/
  server.py            FastMCP server (tool registrations + instructions)
  models.py            normalized Paper model
  aggregate.py         cross-source fusion (dedup + Reciprocal Rank Fusion)
  sources/
    base.py            source registry (get_source / list_sources)
    arxiv.py           arXiv Atom API + read_paper (HTML/markdown/latex)
    semanticscholar.py Semantic Scholar full API surface
    openalex.py        OpenAlex REST API (works/authors/institutions/topics)
    recognize.py       image→LaTeX client over the co-located recognize service
    latextools.py      lint + PDF-extract clients over the latex-tools services
pyproject.toml

Run locally

cd paper-mcp
python -m venv .venv && . .venv/bin/activate
pip install -e .
PAPER_MCP_PORT=9400 python -m paper_mcp.server
# MCP endpoint at http://127.0.0.1:9400/mcp (JSON-RPC; a plain GET returns 406)

Env

Var Default Notes
PAPER_MCP_HOST 127.0.0.1
PAPER_MCP_PORT 9400
PAPER_MCP_PATH /mcp
SEMANTIC_SCHOLAR_API_KEY optional; raises S2 rate limit. Set via /etc/paper-mcp.env in prod.
MCP_MAX_PER_HOUR 300 Direct-client JSON-RPC POST budget per IP.
MCP_WORKER_MAX_PER_HOUR 300 Trusted reverse-proxy Worker budget per HMAC-derived connection key. Raw keys are not retained.
MCP_WORKER_SHARED_MAX_PER_HOUR 2400 Shared ceiling across all trusted Worker connections.
MCP_RATE_COOLDOWN_SEC 300 Minimum fast-rejection cooldown after a bucket reaches its limit.

Deployment (latex-tools.online)

  • Runs as paper-mcp.service on tencent-us (43.130.32.180), WorkingDirectory /opt/paper-mcp, loopback port 9400.
  • nginx reverse-proxies https://latex-tools.online/mcp127.0.0.1:9400/mcp.
  • Worker-aware buckets activate only when a trusted reverse proxy overwrites X-MCP-Worker after validating the upstream platform. Never pass through a client-supplied value.
  • uvicorn access logging is disabled because legacy MCP clients may put connection keys and profiles in the endpoint URL. The reverse proxy must also log $uri, not $request, for the MCP route.
  • Secrets in /etc/paper-mcp.env (SEMANTIC_SCHOLAR_API_KEY).
  • Runtime systemd/nginx/env files are managed by the tencent-us operations backup, not by this source repository; never commit /etc/paper-mcp.env.

Update flow

This repo is the source of truth. The server runs an independent copy under /opt/paper-mcp (not auto-synced):

# edit here → push → deploy the complete canonical Python package
rsync -a --delete paper_mcp/ tencent-us:/opt/paper-mcp/paper_mcp/
ssh tencent-us 'systemctl restart paper-mcp'
ssh tencent-us 'curl -s -o /dev/null -w "%{http_code}\n" http://127.0.0.1:9400/mcp'  # 406 = healthy (needs JSON-RPC handshake)

Production parity verified on 2026-07-23: main@07f6bbe8622aa063f56ee222a40d19c5d4264048 matches all 12 deployed Python source files byte-for-byte. The older copy embedded in latex-tools-deploy/paper-mcp/ is not a deployment source.

Notes

  • arXiv calls are politely rate-limited + retried (_USER_AGENT, backoff).
  • read_paper covers ~80%+ of papers via official HTML; older scan-only papers may have no full text.
  • Moved here from the docs repo on 2026-06-07; that copy is gone.

License

MIT © MCPServings. See LICENSE.

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选