mnema
Provides a file-first personal memory layer for AI agents, enabling them to store and retrieve memories as markdown files with an SQLite index. The MCP server offers read-only search by default, with optional write tools for manual memory addition and conflict resolution.
README
mnema
(Greek μνῆμα — memory, memorial. Same root as Mnemosyne.)
File-first personal memory layer for AI agents. Markdown is the source of truth; SQLite is a disposable index. Conflicting memories are recorded, never auto-deleted — resolution is a human decision, made when it is cheap to ask.
Born from a source-level audit of existing memory systems (memmy-agent, Honcho, mem0). They converge on hybrid retrieval and LLM extraction — and diverge on exactly the things mnema bets on: human-readable storage, mandatory provenance, and honest conflict handling.
Why another memory system
| Problem in existing systems | mnema's answer |
|---|---|
| Memory locked in a DB you can't grep, diff, or edit | Markdown files are canonical; the SQLite index rebuilds from them |
| Contradictions silently accumulate, or an LLM silently deletes the "old" fact | Conflicts become records with provenance; a human resolves; losers are downweighted, never deleted |
| Memories can't be traced back to their source | Every derived memory carries the session ref, message ids, and a redacted excerpt inline — auditable even after transcripts are purged |
| Everything gets stored, quality control deferred to retrieval | Extraction is an explicit, human-confirmed distill step with a write-time filter |
Status: working, pre-release
Everything under Implemented is covered by the test suite (50 tests, offline and deterministic) and has been exercised live against real Claude Code sessions and the real Anthropic API.
Quick start
npm install -g @raytien/mnema # or: pnpm install && pnpm build (from source)
mnema --help
# manual memory
node dist/cli.js add --body "Always use pnpm, never npm" --stable
# hybrid search (first run downloads a ~118MB local embedding model;
# set MNEMA_NO_EMBED=1 for keyword-only, zero download)
node dist/cli.js search "package manager"
# extract memories from a Claude Code session (needs ANTHROPIC_API_KEY)
node dist/cli.js distill ~/.claude/projects/<proj>/<session>.jsonl --output run.json
# review run.json, then:
node dist/cli.js distill --apply run.json
MCP (read-only search from any Claude Code session):
claude mcp add mnema -- mnema-mcp
Env: MNEMA_ROOT (default ~/.mnema), MNEMA_NO_EMBED=1,
MNEMA_ENABLE_WRITE=1 (MCP write tools), MNEMA_MODEL, ANTHROPIC_API_KEY.
Recommended: cd ~/.mnema && git init — your memory history is just files.
Implemented
Storage — files first, crash-safe
- Canonical markdown memories with Zod-validated frontmatter (ULID ids, versioned schema, discriminated source union)
- Write protocol: intent journal → cross-process lock → manifest generation → atomic file write (fsync + rename) → DB transaction. Fault-injection tests cover every crash point; recovery replays the journal without re-calling any LLM
- Repair-before-read across processes (durable dirty marker, not in-memory state); in-place index rebuild that never unlinks an open DB
- Manual edits detected by content hash: revision bumps, source wraps as
revised, embeddings recompute — automatically, on the next index op_key/op_hashidempotency: same key replays, same key with a different payload errors (never silently dropped)
Retrieval — hybrid, multilingual
- FTS5 (contentless-delete,
remove_diacritics 2) + local vector search (sqlite-vec, pinned multilingual MiniLM, q8) fused with RRF - Shared tokenizer for index and query:
Intl.Segmenter+ Han bigrams — Chinese two-character terms actually hit (raw unicode61 scores zero); English/Spanish/Portuguese/French work as-is;cafefindscafé - Query hardening: user input never reaches FTS MATCH raw (
C++,alpha -beta, emoji-only queries are all safe); input caps on every untrusted surface - Time decay after fusion (stable memories exempt); superseded memories downweighted, derived from the resolution graph
- Cross-lingual retrieval via multilingual embeddings (verified live: English queries matching Chinese memories)
Distill — explicit, audited capture
- Claude Code JSONL parsed defensively (no public schema; bad lines counted, never a crash)
- Versioned redaction runs before anything reaches the LLM; secrets never survive into stored excerpts
- One extraction call per preview; write-time filter (preferences, decisions, constraints only — empty sessions honestly yield zero)
run.jsonis immutable (tamper-detected by hash): accept/reject only; apply is LLM-free and idempotent- Fabricated citations are dropped — provenance must be real
Conflicts — record, never auto-delete
- Batched LLM judging of semantically-near pairs; verdicts stored as deterministic relation files with the input hashes they were judged on
- Relations go
stalewhen a member is edited,orphanedwhen deleted; stale verdicts stop affecting ranking resolve keep:<id> | keep_bothre-verifies hashes under lock and rejects supersession cycles (A>B>C>A)- Judge failures are recorded (
error) and retryable — a network blip never becomes a permanently missed conflict - Unresolved conflicts surface alongside search results
Interfaces
- CLI:
add / search / distill / resolve / check-conflicts / index / eval - MCP over stdio: read-only
searchby default;addand two-phaseresolve(preview token required to commit) only behind an explicit flag - Eval harness with a draft gold set; deterministic FTS-only baseline pinned in CI (recall@3 0.70, gate at 0.62)
Security posture
0700/0600permissions, atomic temp-file writes, stdio-only MCP- Honest residual risk: a poisoned conversation distilled into memory is persistent prompt injection; the mitigation is the human confirm step in distill, not a technical control
Roadmap
Near-term (blocked on a confirmed gold set):
- [ ] Calibrate RRF k, decay half-life, and the conflict-candidate distance threshold (current values are literature defaults)
- [ ] Held-out test set and scheduled (non-CI) LLM quality evals: distill precision/recall, redaction precision, contradiction F1 with negative pairs
Planned:
- [ ]
mnema review— batch conflict triage in the terminal - [ ] Distill sources beyond Claude Code (Cursor, Codex session formats)
- [ ] Freshness re-verification (
source_status) against still-existing transcripts - [ ] Custom FTS tokenizer preserving symbol terms (
C++vsC— currently a documented limitation; recall via the vector path only) - [ ] npm publish + prebuilt binary matrix (macOS arm64, Linux x64)
Explicitly out of scope (v1 promises, not omissions):
- No agent runtime, no desktop app, no background daemon — one CLI, one MCP server, LLM calls only in explicit steps
- No auto-deletion of memories, ever
- No multi-user / workspace / auth — a personal, local tool
- Japanese/Korean text: detected and warned, not usefully indexed
Design history
The full design history (comparison matrix, reviewed plan, implementation specs, milestone tracking) is maintained privately.
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。