ragmcp

ragmcp

Retrieval augmented generation server for policy documents, exposing search, grounded Q&A with verified citations, and source management as MCP tools. Runs with local Ollama models for free or managed AWS backends, with an evaluation harness for retrieval and grounding metrics.

Category
访问服务器

README

ragmcp

A retrieval augmented generation service for policy and procedure documents, exposed to any MCP host as a set of tools, and measured by an evaluation harness rather than by spot checks.

Built to run three ways from the same code path:

  • Local, recommended: Ollama for embeddings and generation against an in memory hybrid index. Real learned embeddings, free, private, no account.
  • Managed: Amazon Bedrock for embeddings and generation, Amazon OpenSearch Service for kNN and BM25 hybrid retrieval.
  • Hermetic: a deterministic hashed embedder and an extractive answerer, so the test suite and a fresh clone run with nothing installed and nothing running.

Hermetic is the default. Every other backend is opt in through an environment variable, never autodetected from the presence of AWS credentials or a running Ollama server.

What it costs to run

Component Service Cost
Embeddings Ollama, nomic-embed-text Free, runs locally
Generation Ollama, llama3.1:8b Free, runs locally
Retrieval index In memory hybrid index Free
Retrieval index, managed Amazon OpenSearch Service, one t3.small.search node, 10 GB EBS Free tier, 750 hours per month for 12 months
Document storage Amazon S3 Free tier, 5 GB for 12 months
API surface AWS Lambda Free tier, 1M requests per month, always free
Embeddings, managed Amazon Bedrock, Titan Text Embeddings V2 Metered. No free tier
Generation, managed Amazon Bedrock, Converse API Metered. No free tier

The whole project runs at zero cost on Ollama. Bedrock has no free tier at any usage level, which is why it is opt in and why the default path never calls it.

Two traps worth knowing on the AWS side: OpenSearch Serverless is not free tier eligible, because it bills a minimum OCU allocation whether or not anything is indexed, so use a managed domain on t3.small.search. And the free OpenSearch tier is 12 months at 750 hours per month, not always free.

Inference stays local rather than going to a hosted free tier for a second reason beyond cost: most free inference tiers reserve the right to train on submitted prompts. That is fine for the synthetic corpus in data/, and it is exactly the question a security review asks about a real document corpus.

Why it is built this way

The interesting problems in a RAG deployment are not the embedding call. They are: what happens when the corpus does not contain the answer, what happens when the model cites a passage that was never retrieved, and how anyone knows whether a retrieval change made things better or worse. Each of those is a component here with a test behind it.

Architecture

document ──> chunking ──> embeddings ──> index ──> retrieval ──> gate ──> generation ──> citation
             heading      Ollama,        Amazon    BM25 +        abstain  Ollama,         verification
             scoped,      Bedrock or     OpenSearch vectors,     when the Bedrock or      drop unmapped
             overlapping  hashing        or local  fused         corpus   extractive      markers
                                                                 is thin  fallback
Module Responsibility
ragmcp/chunking.py Heading scoped, token budgeted chunks with sentence level overlap
ragmcp/embeddings.py Ollama and Bedrock Titan embeddings, plus a deterministic hashed backend
ragmcp/index.py BM25 and dense vector scoring fused with Reciprocal Rank Fusion
ragmcp/opensearch.py Amazon OpenSearch backend, HNSW cosine kNN field and a normalising fusion pipeline
ragmcp/llm.py Ollama and Bedrock generation, plus an extractive fallback for timeouts
ragmcp/ollama_client.py Dependency free HTTP client with actionable errors for a missing model or a stopped server
ragmcp/pipeline.py Ingest, retrieve, abstention gate, citation verification
ragmcp/server.py MCP server over stdio, JSON-RPC 2.0, no SDK dependency
eval/ Labelled question set, retrieval metrics, groundedness metrics

Fusion happens in a different place depending on the backend, which is worth knowing before reading the code. Locally, the two ranked lists are fused client side with Reciprocal Rank Fusion. On OpenSearch, both queries go up as one hybrid query and a search pipeline normalises the two score distributions before combining them, which is necessary because BM25 scores are unbounded while cosine scores sit in [0, 1].

MCP tools

Tool Purpose
search_documents Retrieve passages by hybrid, keyword or vector search
answer_question Grounded answer with verified citations, or an explicit refusal
list_sources What is indexed and how many chunks each document produced
ingest_document Chunk, embed and index a document at runtime

Every call is validated against its input schema before it reaches a handler. Bad input comes back as an MCP tool error with isError: true; a handler exception is caught, logged and returned the same way, so a malformed call from a host never takes the server down.

Register it with any MCP host:

{
  "mcpServers": {
    "ragmcp": {
      "command": "python",
      "args": ["-m", "ragmcp.server", "data"],
      "cwd": "/path/to/ragmcp"
    }
  }
}

Quick start

Hermetic, needs nothing installed:

python -m unittest discover -s tests    # 33 tests, no network, no credentials
python eval/run_eval.py --k 3           # evaluation report as JSON
python -m ragmcp.server data            # MCP server on stdio

With real local models, which is the interesting configuration:

ollama pull nomic-embed-text
ollama pull llama3.1:8b

export RAGMCP_EMBEDDINGS=ollama
export RAGMCP_GENERATOR=ollama
python eval/run_eval.py --k 3 --json eval/results.json

To use the managed backends, install requirements-aws.txt, set the variables in .env.example, and call OpenSearchIndex.create() once to provision the index mapping and the fusion pipeline.

Evaluation

48 questions over an 8 document, 43 chunk corpus. 40 are answerable and 8 are deliberately out of scope. Retrieval is judged at chunk level, not document level: with a corpus this small, document level judgements score every retriever at recall 1.0 and hide every regression worth catching. Gold chunks are weak labelled, meaning a chunk counts as relevant when it sits in a labelled document and contains the answer bearing string.

Retrieval, at k=3, on the hermetic backend (hashed embedder). These are not the numbers you get with Ollama embeddings, see the note below the table:

Mode recall@3 MRR nDCG@3 p50 latency
keyword 0.833 0.838 0.813 0.02 ms
vector 0.783 0.746 0.732 1.40 ms
hybrid 0.808 0.796 0.775 1.50 ms

End to end answers:

Metric Value
Citation support rate, in scope 0.875
Grounded rate, in scope 0.900
False abstention rate, in scope 0.100
Abstention rate, out of scope 1.000

What the numbers actually say

The published table is the hashed embedder, and it is the weakest configuration on purpose. It is the one that runs anywhere with nothing installed, so it is the one that can be reproduced from a clean clone. Rerun with RAGMCP_EMBEDDINGS=ollama to measure the configuration that is actually recommended, and replace the table above with what you get.

Hybrid does not beat keyword search in that table, and that is expected. RRF fuses two ranked lists and lands between them when one retriever is much weaker. The hashed embedder has no learned semantics, so it contributes little and drags the fused ranking down. Weighted RRF and a truncated SVD embedder were both tried; neither moved MRR above keyword alone. The fusion is kept because it is the piece that pays off once real embeddings sit behind it, and the hypothesis is that nomic-embed-text reverses this result. That hypothesis is untested here and is written down rather than assumed, because the whole point of the harness is that claims like it get measured before they get repeated.

The abstention gate was rebuilt after the first version failed. The first gate measured how many query terms appeared in the retrieved passages. It scored some out of scope questions above some in scope paraphrases, because words like "customers" and "review" appear throughout a policy corpus, and it falsely refused 35 percent of answerable questions. The BM25 top score separates the two populations far better through term saturation and length normalisation: 100 percent abstention on out of scope questions at 10 percent false abstention.

The remaining failures are vocabulary mismatch. The questions still missed are paraphrases that share almost no terms with the source text, such as asking what stops someone redirecting funds by fake email when the policy says callback verification against the vendor master record. This is precisely the gap a real embedding model closes, which is the argument for keeping the vector half of the index.

Known limitations

  • The BM25 gate threshold scales with corpus size and needs retuning per index. It is a constructor argument, not a constant, and the tests pin their own.
  • Weak labelling ties gold chunks to a literal string. It is cheap and it survives a chunking change, but it cannot label a question whose answer is spread across two chunks.
  • The extractive fallback returns leading sentences from the top passages. It keeps the tool honest and cited when the model is unreachable, but it is a degradation path, not an answer quality strategy.
  • The Bedrock and OpenSearch backends are exercised by their interface contract, not against live services, so those code paths carry no integration test here.
  • The published metrics come from the hashed embedder. Ollama or Bedrock embeddings change the vector and hybrid rows, so the table must be regenerated before those numbers are quoted for any other configuration.
  • Ollama generation quality is not measured here. The citation support and groundedness numbers come from the extractive answerer, which cannot hallucinate a citation by construction, so they set a floor rather than describing what a generative model does on this corpus.

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选