reddit-research
MCP server that answers questions using Reddit discussions as evidence, searching threads, fetching comments, and ranking useful context, with optional cited answer synthesis.
README
reddit-research
A personal-use, retrieval-augmented research tool that answers questions using Reddit discussions as evidence. It searches for relevant threads via a search API, fetches thread/comment data through the official Reddit Data API, ranks the most useful evidence, and returns structured context (optionally synthesized into a cited answer).
It is intended for interactive question answering and short-lived local research. It does not train models on Reddit data, bulk-archive Reddit, or redistribute collected datasets. See the spec for the full design and non-goals.
Install
uv sync # core CLI
uv sync --extra mcp # + MCP server
uv sync --extra dev # + test tooling
Configure
Configuration splits cleanly in two:
-
Secrets → environment /
.env(never the TOML). Copy.env.example:# Reddit creds — only for the "praw" backend (see Fetch backends below) REDDIT_CLIENT_ID=... REDDIT_CLIENT_SECRET=... BRAVE_SEARCH_API_KEY=... # or TAVILY_API_KEY (search always needs a key) ANTHROPIC_API_KEY=... # only if you enable synthesis -
Non-secret behavior →
reddit-research.toml(seereddit-research.example.toml):backend,user_agent, search/synthprovider, comment limits, cache TTLs, etc.
A .env in the current directory (or ~/.config/reddit-research/.env) is loaded
automatically. Precedence is real env var > .env file > TOML > default, so any
TOML value can be overridden by its env var when needed — but each setting has
one canonical home to avoid duplication.
Fetch backends
Reddit gated self-serve API app creation in 2026 (the "Responsible Builder
Policy"), so an approved OAuth app is no longer guaranteed. The fetcher therefore
supports several backends, selectable per-command with --backend or via
[reddit] backend / REDDIT_RESEARCH_BACKEND:
| Backend | Auth | Notes |
|---|---|---|
auto (default) |
— | praw if Reddit creds are set, else arctic |
arctic (recommended keyless) |
none | Arctic Shift archive with PullPush fallback. Works from any IP/host. Serves periodically-updated historical data, so very recent threads may lag; retains deleted/removed content. |
praw |
Reddit OAuth app | Official API; live data + full comment-tree expansion. Needs an approved app. |
json |
none | Reddit's public .json endpoints. Largely unusable in 2026: Reddit fingerprint-blocks the .json path (403) for non-browser clients — even from a residential IP, and even with a matching browser User-Agent (it checks the TLS/HTTP2 fingerprint of a current browser). Kept for the rare environment where it still works. |
reddit-research evidence "..." # auto -> arctic (keyless)
reddit-research answer "..." --backend praw # live data, needs an approved app
Compliance note: arctic/json let you run without an approved app, but using
them to sidestep API approval sits in tension with the spec's "don't circumvent
access controls" non-goal. Intended for genuine personal, low-volume use.
CLI
reddit-research search "best backup strategy for homelab"
reddit-research fetch "https://www.reddit.com/r/selfhosted/comments/..."
reddit-research evidence "what do selfhosted users recommend for backups?" -r selfhosted -r datahoarder
reddit-research answer "what do Reddit users recommend for homelab backups?"
reddit-research cache stats
reddit-research cache purge --older-than-days 30
Useful flags: --subreddit/-r (repeatable), --limit-threads, --max-comments,
--sort, --since-days, --format json|markdown, --no-llm, --show-queries,
--verbose.
MCP server
Exposes search_reddit_threads, fetch_reddit_thread, rank_reddit_evidence,
and the high-level answer_from_reddit. Two transports (REDDIT_RESEARCH_MCP_TRANSPORT):
stdio (default) — local use; the client spawns the process:
{
"mcpServers": {
"reddit-research": {
"command": "reddit-research-mcp",
"env": { "BRAVE_SEARCH_API_KEY": "..." }
}
}
}
http / streamable-http — a long-lived network server for container/NAS
deployment (see below). A bearer token is required; the server refuses to
start over HTTP without REDDIT_RESEARCH_MCP_TOKEN, and every request except
GET /healthz must send Authorization: Bearer <token>.
Deploy on a NAS (Docker)
The included Dockerfile + docker-compose.yml run the MCP server over HTTP.
-
Configure — create
.envnext to the compose file:BRAVE_SEARCH_API_KEY=... # search needs a key REDDIT_RESEARCH_MCP_TOKEN=$(openssl rand -hex 32) # required bearer token # ANTHROPIC_API_KEY=... # only if you enable synthesis -
Run — build from source:
docker compose up -d --build curl http://127.0.0.1:8000/healthz # -> okTo deploy a prebuilt image from your own registry instead, set
REDDIT_RESEARCH_IMAGEin.envand pull:export REGISTRY=your-registry.example.com echo "REDDIT_RESEARCH_IMAGE=$REGISTRY/reddit-research-mcp:latest" >> .env docker login "$REGISTRY" # once docker compose pull && docker compose up -dTo publish a new image after code changes:
docker build --provenance=false \ -t "$REGISTRY/reddit-research-mcp:latest" . docker push "$REGISTRY/reddit-research-mcp:latest"The SQLite cache persists in
./data. The container binds to127.0.0.1:8000by default, so it's reachable only through the NAS's reverse proxy (change theports:mapping to8000:8000to expose it on the LAN instead). -
Reverse proxy (Synology) — Control Panel → Login Portal → Advanced → Reverse Proxy → Create:
- Source:
https://reddit-mcp.<your-domain>(port 443, HTTPS — enables TLS) - Destination:
http://localhost:8000 - Enable HSTS as desired; the streamable-HTTP transport streams responses, so leave response buffering off (the default reverse-proxy behavior is fine).
- Source:
-
Connect from Claude Code:
claude mcp add --transport http reddit-research \ https://reddit-mcp.<your-domain>/mcp \ --header "Authorization: Bearer <your REDDIT_RESEARCH_MCP_TOKEN>"
Security notes: the token gates outbound calls that spend your API keys, so keep
it secret and prefer a random 32-byte value. The compose file sets the keyless
arctic backend (the public .json path is fingerprint-blocked by Reddit for
non-browser clients — see Fetch backends). Switch to praw only if you have an
approved OAuth app and set its credentials in .env.
Architecture
question -> query planner -> search provider -> reddit URL extractor
-> Reddit API fetcher -> cache -> evidence ranker
-> structured evidence -> optional answer synthesis
Package layout under src/reddit_research/:
| Module | Responsibility |
|---|---|
core/config.py |
Env + TOML config, secret handling |
core/models.py |
Pydantic data contracts |
core/search.py |
Query planner + Brave/Tavily providers |
core/reddit.py |
PRAW fetcher + normalization |
core/public_json.py |
Keyless .json fetcher |
core/archive.py |
Arctic Shift / PullPush fetcher |
core/fetchers.py |
Backend selection (auto/praw/json/arctic) |
core/urls.py |
Reddit URL/ID parsing |
core/cache.py |
SQLite cache (TTL + purge) |
core/ranking.py |
Explainable lexical ranking |
core/synth.py |
Optional, provider-agnostic synthesis |
core/orchestrator.py |
Pipeline + run metadata/warnings |
cli.py / mcp_server.py |
Interfaces |
Develop
uv run pytest
Status
MVP implemented: search/fetch/evidence/answer CLI, Brave + Tavily
providers, three fetch backends (praw/json/arctic), SQLite cache, lexical
ranking, JSON/Markdown output, and the MCP server. Later enhancements (async
fetching, semantic reranking, branch summarization) are tracked in the spec.
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。