gurbani-mcp

gurbani-mcp

A local self-hosted MCP server for searching and verifying Gurmukhi quotes from Sri Guru Granth Sahib, offering keyword search, authenticity verification, and text guarding via an MCP or HTTP API.

Category
访问服务器

README

gurbani-mcp

A local, self-hosted tool for looking up and verifying that a quote is authentically from Sri Guru Granth Sahib (SGGS). This project covers SGGS only (not other Banis).

It runs entirely on your own machine. No data leaves your computer unless you choose to expose it to an AI client (Claude, ChatGPT, etc.), and even then, only the query text you send is transmitted — never the underlying database.

  • Search — keyword search with Gurbani concept expansion (e.g. "sewa" also matches "service", "selfless service")
  • Verify — check a Gurmukhi quote against SGGS; get back the verbatim text + citation (Ang, Shabad, author), or a clear "not found"
  • Guard — scan a block of text for every Gurmukhi quote in it and verify each one individually
  • Two ways to connect an AI client: an MCP server (Claude Desktop, Claude Code, Cursor) and a plain HTTP API (ChatGPT via Custom GPT Actions, or any REST client)

Gurbani text is only ever returned from the source database — never paraphrased, summarized, or generated. See Quote Verification below.


Quick start

Prerequisites:

Tool Why Install
Docker (or Colima on macOS) one-time database build brew install colima docker && colima start --cpu 4 --memory 8
uv run the Python servers curl -LsSf https://astral.sh/uv/install.sh | sh
git clone <this-repo-url> gurbani-mcp
cd gurbani-mcp
bash scripts/setup.sh

setup.sh builds database/dist/banidb.sqlite from the official Khalis Foundation BaniDB Docker image (the dataset behind SikhiToTheMax), then installs Python dependencies. It takes several minutes the first time (downloading + seeding a ~640MB dataset); nothing about your database is uploaded anywhere.

Verify it worked:

bash scripts/test_search.sh "benefits of sewa"
uv run --extra dev pytest

Connect to an AI client

Claude Desktop / Claude Code (MCP)

Add this to your Claude Desktop config (~/Library/Application Support/Claude/claude_desktop_config.json on macOS, %APPDATA%\Claude\claude_desktop_config.json on Windows):

{
  "mcpServers": {
    "gurbani": {
      "command": "uv",
      "args": ["run", "--directory", "/absolute/path/to/gurbani-mcp", "python", "-m", "mcp_server.server"]
    }
  }
}

Restart Claude Desktop. You should see a 🔨 tools icon indicating the gurbani server is connected, exposing search_gurbani, verify_quote, guard_text, get_shabad_by_ang, get_line, and get_shabad_by_line.

For Claude Code, add the same server with:

claude mcp add gurbani -- uv run --directory /absolute/path/to/gurbani-mcp python -m mcp_server.server

ChatGPT (Custom GPT Actions, via the HTTP API)

  1. Start the HTTP API:

    uv run uvicorn api_server.app:app --port 8421
    
  2. ChatGPT Actions need a public HTTPS URL — they can't reach localhost. The fastest way to get one without an account is a Cloudflare quick tunnel:

    cloudflared tunnel --url http://localhost:8421
    

    This prints a temporary https://<random>.trycloudflare.com URL. Treat it as sensitive while it's live — anyone with the URL can query your local API. It's meant for short sessions; for anything longer-lived, put an API key or auth layer in front of it first (not included here — see Follow-ups).

  3. In ChatGPT: Explore GPTs → Create → Configure → Actions → Import from URL, and paste https://<random>.trycloudflare.com/openapi.json. ChatGPT will pick up all the endpoints (/api/search, /api/verify, /api/guard, etc.) automatically.

  4. Give the GPT instructions like: "When asked to verify a Gurbani quote, always call the verify or guard action and quote its source_text back verbatim — never answer from your own memory."


Example: verifying every quote in a document

curl -s "http://localhost:8421/api/guard" --get \
  --data-urlencode "q=$(cat my_document.txt)" | jq

Returns every Gurmukhi span found in the text, each marked verified, verified_fuzzy (found, but with minor punctuation/spelling differences), or not_found — with the exact source citation (Ang, author, full line) for anything that verified.


Quote verification — how authenticity is guaranteed

Four layers, all sharing one matching core in gurbani_rag/verify.py:

  1. Build gate (scripts/validate_db.py) — structural checks (row counts, Ang coverage, no gaps/duplicates) run automatically during setup.sh, so an incomplete or corrupted database can never reach runtime.
  2. Retrieval — every search result comes straight from the source database with its citation attached. Authentic by construction.
  3. Verify (verify_quote) — exact match first (punctuation-agnostic), then fuzzy match via FTS5 + rapidfuzz with a 0.90 confidence floor. Below that: not_found. The text returned is always the source's own — never your input echoed back.
  4. Guard (guard_text) — scans arbitrary text for every Gurmukhi span and verifies each one independently. This is the tool for auditing any drafted or existing content.

See CLAUDE.md for the full architecture and schema.


Optional: semantic search

The default search is keyword + concept expansion (works well, no extra setup). An optional ChromaDB-based semantic index can also be built:

uv run python scripts/build_index.py

Running tests

uv run --extra dev pytest

tests/test_golden.py checks known-authentic quotes verify at their correct Ang, and known fakes are correctly rejected.

Data attribution

Scripture text, translations, and transliterations: Khalis Foundation BaniDB (SikhiToTheMax dataset).

BaniDB's compiled/proprietary form is not redistributed by this repo — database/dist/banidb.sqlite is always built locally from the official BaniDB Docker image (see scripts/setup.sh).

Follow-ups (not built yet)

  • The HTTP API has no authentication — fine for a short-lived tunnel session, not for leaving it exposed long-term.
  • No CI workflow yet.

License

MIT — see LICENSE.

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选