ragvault

ragvault

MCP server for local-first RAG over Obsidian vaults, enabling AI agents to search and ask questions about notes with grounded citations.

Category
访问服务器

README

rag-vault

CI Python License

Local-first RAG over my Obsidian vault — ask my second brain questions in the terminal, get grounded answers with citations that deep-link back into Obsidian. The retrieval internals are hand-rolled (BM25 + embeddings + reciprocal rank fusion), and the whole engine doubles as an MCP server so AI agents on my machine can search my notes as a tool.

demo

Why hand-rolled?

At vault scale (dozens–hundreds of notes), a vector database is overkill — brute-force cosine over a numpy matrix answers in under a millisecond. So this repo implements the interesting parts itself, in plain Python I can defend line by line:

  • Vector store → SQLite + a numpy matrix (unit-normalized float32; cosine = dot product)
  • BM25 → ~40 lines of term-frequency math (k1=1.5, b=0.75)
  • Hybrid fusion → Reciprocal Rank Fusion: score = Σ 1/(60 + rank) — no weights to tune
  • Chunking → markdown-aware: splits on headings, keeps a Note > Heading breadcrumb, extracts [[wikilinks]]

Frameworks (LangChain, LlamaIndex) would hide exactly the parts this project exists to understand.

Local-first by design. My vault contains journal entries and career notes. Embeddings (nomic-embed-text) and default answer generation (qwen2.5:3b) run on my Mac via Ollama. Nothing leaves the machine unless I explicitly pass --provider claude — and even then, only the question plus the retrieved excerpts are sent, never the vault.

How it works

flowchart LR
    V[Obsidian vault\n*.md] --> C[chunker\nheading-aware]
    C --> E[embeddings\nnomic-embed-text]
    E --> S[(SQLite\nincremental sync)]
    Q[question] --> H{hybrid search}
    S --> H
    H -->|cosine| F[RRF fusion]
    H -->|BM25| F
    F --> A[grounded answer\nqwen local / claude opt-in]
    A --> T[cited answer\nobsidian:// links]

Retrieval quality, measured

vault eval scores retrieval against 12 real questions with known source notes (hit@k: was the right note in the top k; MRR: mean reciprocal rank of the first hit):

mode hit@1 hit@3 hit@6 MRR
vector 0.67 1.00 1.00 0.82
bm25 0.58 0.75 0.83 0.67
hybrid 0.67 0.83 0.92 0.77

(numbers from my vault — rerun with python -m ragvault eval)

On this vault, plain vector search actually beats hybrid on hit@3/hit@6/MRR — my questions phrase concepts close to how the notes word them, so dense embeddings alone do well, and RRF's rank-based blend lets BM25's misses drag hybrid down a bit. Hybrid still beats BM25 alone across the board, and I'd expect it to pull ahead on a vault with more exact-keyword lookups (IDs, code, jargon).

Quickstart

Works on any Obsidian vault (or any folder of markdown):

git clone https://github.com/maxrotemberg04-spec/rag-vault && cd rag-vault
python3 -m venv .venv && .venv/bin/pip install -r requirements.txt
ollama pull nomic-embed-text   # embeddings (local)
ollama pull qwen2.5:3b         # answers (local)

export RAGVAULT_VAULT=~/path/to/your/vault
.venv/bin/python -m ragvault index
.venv/bin/python -m ragvault ask "what did I decide about X?"
.venv/bin/python -m ragvault search "keyword hunt" --mode bm25   # no LLM needed

--provider claude uses the Anthropic API for answers if ANTHROPIC_API_KEY is set (retrieved excerpts only — the vault itself never uploads).

Agents can use it too (MCP)

The same engine runs as an MCP server:

claude mcp add ragvault -e RAGVAULT_VAULT=$HOME/Documents/FOCUS -- \
    $PWD/.venv/bin/python $PWD/mcp_server.py

Claude Code sessions then get two tools — search_vault (hybrid retrieval; the agent synthesizes) and ask_vault (fully local answer). My "Educator" Claude session uses this instead of grepping the vault.

Design decisions

  • No vector DB — at this scale the honest engineering answer is a numpy dot product. At ~100k documents I'd reach for HNSW indexes (or pgvector) and this section would change.
  • RRF over weighted score fusion — rank-based fusion is scale-free, so BM25 and cosine scores never need calibrating against each other.
  • Two texts per chunk — verbatim display_text for humans, cleaned embed_text (wikilinks resolved, callout markers stripped, breadcrumb prepended) for the models.
  • Grounding contract — the answer prompt allows only the retrieved excerpts, requires [n] citations, and must say "That's not in the vault" rather than guess.
  • Evals are part of the product — same philosophy as my eval-harness: if you can't measure retrieval, you can't improve it.

Limitations & roadmap

Single-user, single-vault by design. No re-ranker; no stemming (BM25 is exact-token). Roadmap: vault ui (local web page), link-graph ranking boost using the [[wikilink]] graph, --watch auto-reindex, semantic query cache.

Full design doc: docs/design.md

MIT © Max Rotemberg

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选