recall-mcp

recall-mcp

A shared, local-first memory layer for AI CLIs, providing persistent, layered memory across Claude Code, Gemini CLI, and other MCP-aware clients.

Category
访问服务器

README

recall-mcp

One shared, layered, local-first brain for every AI CLI you use. Claude Code, Gemini CLI, Cursor, Continue, Zed — they all forget. recall-mcp is the memory they share.

License: MIT Python 3.10+ MCP Local-first

<p align="center"> <img src="docs/demo.gif" alt="Gemini CLI calling recall-mcp's memory_recall tool to answer 'what did we ship today and why isn't it called brain-mcp?' — surfacing the rename decision and shipped-today architecture from the layered memory store" width="820"> <br> <em>Gemini CLI recalling today's decisions from a brain it shares with Claude Code and Hermes.</em> </p>

Built on the layered memory engine from Hermes Agent by Nous Research (MIT). recall-mcp packages that engine as a standalone MCP server so any AI client — not just Hermes — can plug into the same brain. Original architecture: theirs. Packaging, MCP surface, cross-CLI integration: this project. See Credits.

Quick start

# Install
pipx install recall-mcp

# Wire it into Claude Code (one-time)
echo '{"mcpServers":{"recall-mcp":{"type":"stdio","command":"recall-mcp"}}}' >> ~/.claude.json

# Restart Claude Code. Done.

That's it. Every conversation now writes to and reads from the same persistent brain — and so do Gemini CLI, Cursor, and any other MCP-aware client you wire up the same way.

What it does

flowchart TD
    A[Claude Code] -- MCP --> M[recall-mcp]
    B[Gemini CLI] -- MCP --> M
    C[Cursor / Continue / Zed] -- MCP --> M
    M --> S[(SQLite<br/>facts)]
    M --> V[(ChromaDB<br/>vectors)]
    M --> E[(Entity<br/>graph)]
    M --> T[(Temporal<br/>lineage)]
    M --> F[(FTS5<br/>keyword)]
    classDef client fill:#1f6feb,stroke:#1f6feb,color:#fff,stroke-width:0
    classDef brain fill:#a371f7,stroke:#a371f7,color:#fff,stroke-width:0
    classDef store fill:#0d1117,stroke:#30363d,color:#7d8590
    class A,B,C client
    class M brain
    class S,V,E,T,F store

Every AI CLI has the same blind spot: each new session starts with amnesia. Native save_memory tools store flat lists that bloat the system prompt over time. Cloud memory services need accounts, paid tiers, and trust your data to a vendor.

recall-mcp gives you one brain shared by every MCP-aware AI client:

  • 🧠 7 memory layers — vector similarity, BM25 keyword, entity graph, temporal lineage, importance scoring, forgetting engine, hybrid retrieval
  • 🔌 Drop-in via MCP — works with Claude Code, Gemini CLI, Cursor, Continue, Zed, any client speaking Model Context Protocol
  • 🏠 Local-first — SQLite + ChromaDB on your machine. No accounts, no Docker, no cloud lock-in
  • 🔄 Brain-swappable — switch between Claude, Gemini, MiniMax, Qwen — they all share the same memory
  • 🛡️ Graceful degradation — when embeddings hit rate limits, BM25 + entity + temporal carry the load. Never poisons the index

Install

pipx install recall-mcp

Or with uv:

uv tool install recall-mcp

Or from source:

git clone https://github.com/Dhari-Q/recall-mcp
cd recall-mcp
pip install -e .

Configure your AI client

Claude Code

Add to ~/.claude.json under your project's mcpServers:

{
  "mcpServers": {
    "recall-mcp": {
      "type": "stdio",
      "command": "recall-mcp"
    }
  }
}

Gemini CLI

Add to ~/.gemini/settings.json:

{
  "mcpServers": {
    "recall-mcp": {
      "command": "recall-mcp",
      "trust": true
    }
  }
}

Cursor

Add to ~/.cursor/mcp.json:

{
  "mcpServers": {
    "recall-mcp": {
      "command": "recall-mcp"
    }
  }
}

Restart your client. Done.

Five tools you'll use

Tool Purpose
memory_recall(query, top_k) Hybrid search across all layers — vector + BM25 + entity + temporal
memory_remember(content, type, confidence, tags) Store a fact, decision, preference, or gotcha
memory_recent_sessions(limit) List recent session summaries with decisions and bug fixes
memory_search_entity(name, limit) Find memories tied to a specific file, project, person, or tool
memory_stats() Sanity-check counts across every layer

Optional: real semantic search

By default, recall-mcp ships with BM25 keyword + entity graph + temporal retrieval — those work without any API key.

To enable vector / semantic search (queries like "how do I swap the AI" finding "switchable via /model" without shared keywords), point recall-mcp at an embeddings provider:

Create ~/.recall-mcp/.env (or export in your shell):

# MiniMax (global) — fastest path
MINIMAX_API_KEY=sk-...

# Or OpenAI
OPENAI_API_KEY=sk-...

# Or OpenRouter
OPENROUTER_API_KEY=sk-...

Vector layer activates automatically on next start.

Optional: auto-prefetch hook for Claude Code

The MCP tools above are deliberate — the model has to call them. For silent automatic recall on every prompt (like Claude Code's native memory but layered), add a UserPromptSubmit hook. See examples/claude_code_hook.md for the recipe.

Memory types

When you ask the model to remember something, it picks one of:

Type Decay Examples
architecture Permanent "We use ChromaDB for vectors"
decision Permanent "We chose MIT over GPL"
convention Permanent "All API calls go through retry_utils"
pattern Permanent "Use with statements for sqlite connections"
gotcha Permanent "MiniMax embeddings are NOT OpenAI-compatible"
preference Permanent "User prefers terse responses"
progress 7 days "Finished MCP wiring on 2026-04-28"
context 30 days Misc. background facts

Storage location

All data lives in $RECALL_MCP_HOME (defaults to ~/.recall-mcp/):

~/.recall-mcp/
├── memory/          # SQLite — facts + entity graph + temporal lineage
├── episodic/        # SQLite — session summaries
└── chroma/          # ChromaDB — vector embeddings

Set RECALL_MCP_HOME to point multiple machines at a synced folder (e.g., Syncthing) and your AI's memory follows you.

Architecture

recall-mcp exposes seven memory layers (originally designed in Hermes Agent), each backed by a focused storage engine:

  1. Episodic (per-turn / per-session events) — SQLite
  2. Semantic (extracted facts, decisions) — SQLite + ChromaDB
  3. Entity graph (who/what/why, dependencies) — SQLite
  4. Temporal lineage (millisecond timestamps, before/after queries) — SQLite
  5. Importance scoring (not all memories equal) — derived
  6. Forgetting engine (decay + Jaccard dedup) — derived
  7. Hybrid retrieval (BM25 + vector + entity + temporal, fused with optional LLM re-rank) — runtime

When you call memory_recall, all four retrieval paths run in parallel, results are deduplicated, scored by source quality + importance, and returned ranked.

Credits

Memory architecture derived from Hermes by Nous Research (MIT). recall-mcp generalizes the layered memory + retrieval engine into a standalone MCP server that any AI client can plug into.

License

MIT — see LICENSE.

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选