ollama-mcp

ollama-mcp

Enables AI agents to offload mechanical, high-token work to local Ollama models through MCP, with role-based model discovery, batch processing, and file-aware inputs.

Category
访问服务器

README

ollama-mcp

An MCP server that lets an agent — Claude Code, or anything else speaking MCP — hand work to your local Ollama models.

The point is not to wrap the Ollama API. It is to give a frontier-model agent a cheap local tier it can delegate to: summarizing a 40KB log, extracting fields from twenty dumped files, reformatting JSON — mechanical, high-token, low-judgment work that burns expensive context for no benefit.

No model name appears anywhere in src/. Models are discovered live from the daemon and addressed by capability or by role. Pull a new model and it becomes usable within a minute, with no code change, no config edit, and no restart. That property is enforced in CI, not by convention.


Install

git clone https://github.com/Clickt-Digital-Marketing-Inc/ollama-mcp.git
cd ollama-mcp
npm install
npm run build

Register with Claude Code (user scope — available in every project):

claude mcp add --scope user ollama -- node /absolute/path/to/ollama-mcp/dist/src/index.js

Or add it to any MCP client's config:

{
  "mcpServers": {
    "ollama": {
      "command": "node",
      "args": ["/absolute/path/to/ollama-mcp/dist/src/index.js"]
    }
  }
}

Requires Node ≥ 20 and a running Ollama daemon. No configuration is needed to start — roles resolve by capability against whatever you already have installed.

Raise your client's tool timeout

This is the first thing most people hit, and it is not something this server can fix for you.

Local generation is slow by frontier-API standards: a large model can take ~12s just to load, and a long generation or a batch runs for minutes. MCP clients apply their own request timeout, and it is usually 60 seconds. When it fires, your client kills the call while the server is still working correctly — so the timeout_ms argument on the tools is not sufficient on its own. It bounds the server's wait on Ollama; it cannot extend your client's patience.

Codex (~/.codex/config.toml):

[mcp_servers.ollama]
command = "node"
args = ["/absolute/path/to/ollama-mcp/dist/src/index.js"]
startup_timeout_sec = 60
tool_timeout_sec = 900

Writing your own client with the TypeScript SDK — the third argument to callTool:

await client.callTool({ name: 'ollama_dispatch', arguments: {...} }, undefined,
  { timeout: 900_000, maxTotalTimeout: 900_000 });

If dispatches die at almost exactly 60 seconds, this is why — not the daemon, and not timeout_ms.


Tools

ollama_dispatch

One local generation. The main tool.

{ "prompt": "Summarize this changelog in 3 bullets.", "model": "role:summarize" }

Supports system, multi-turn messages, structured output via format (either "json" or a full JSON Schema), tool-calling passthrough, the usual sampling controls, and files/file_globs (below).

Every response ends with a metrics line:

[ollama model=<name> via=role:summarize→role:fast→caps:completion
 tok=3120→412 dur=6.4s load=0.0s rate=64tok/s ctx=131072 done=stop think=off]

model and via are always present. A dispatcher that silently routes to the wrong model is the most expensive failure mode there is, so the resolution trail is never hidden.

ollama_dispatch_batch

Fan-out over many prompts. Items are grouped by resolved model and groups run sequentially, so a cold load is paid at most once per model instead of thrashing VRAM. Results come back in input order regardless of execution order, and one failing item never voids the run.

ollama_models

Discovery: capabilities, context window, size, residency. Two things worth knowing:

  • refresh: true re-reads the daemon after you pull something.
  • explain_selector: "role:coder" dry-runs the resolver and prints the whole fallback chain without spending a generation. When routing surprises you, start here.

ollama_lifecycle

status / warm / unload. A large model can take ~12s to load and occupy tens of GB of VRAM, so warming before a batch and unloading afterwards are both real operations you'll want.


Choosing a model

Three grammars for the model field:

Form Example Meaning
literal qwen3:32b that exact model (bare names resolve to :latest, or to a single installed tag)
role role:summarize an ordered fallback chain
capability caps:vision+tools any installed model with all those capabilities
(omitted) the configured default role

Ambiguity is refused rather than guessed: if foo matches three installed tags, you get an error listing them. A wrong-model run is invisible in the output, so it is not something to coin-flip.

Roles

A role is an ordered chain. Each link is a literal name, another role, or a capability predicate — and the first link that resolves wins:

{
  "roles": {
    "coder": { "chain": ["some-coding-model", "some-fallback-model", "caps:completion"] }
  }
}

If the preferred model isn't installed, the chain falls through. This is the future-proofing story: the chain documents your intent even when the model isn't there yet, and starts routing to it the moment you pull it.

Built-in roles — all defined purely as capability predicates, so they work against any install: general, fast, big, reasoner, coder, vision, tools, embed, summarize, extract.

Adding a model

Three ways, in increasing order of commitment:

  1. Just pull it. ollama pull <model>. Within ~60s it joins the candidate pool for every role and capability it qualifies for, and is addressable by name. Nothing else to do.
  2. Pin a preference. Add it to the front of a role's chain in ollama-mcp.config.json.
  3. Use an env var, no file at all. OLLAMA_MCP_ROLE_CODER="model-a,model-b,caps:completion". OLLAMA_MCP_ROLE_<NAME> is parsed generically, so this also creates roles — OLLAMA_MCP_ROLE_TRANSLATOR=... gives you role:translator with no code change.

Config is discovered at $OLLAMA_MCP_CONFIG, then ./ollama-mcp.config.json, then ~/.config/ollama-mcp/config.json; first hit wins. A malformed config is non-fatal — the server logs, falls back to defaults, and warns on the first response, because a typo should never take the server down.


File-aware inputs

files and file_globs make the server read files and feed them to the local model:

{ "prompt": "Extract every TODO with its file and line.",
  "file_globs": ["src/**/*.ts"],
  "model": "role:extract" }

File contents never enter the calling agent's context — only the model's distilled answer comes back. For large inputs this is the whole reason the server is worth having.

That is a statement about your agent's context, not about the network. File contents are sent to whichever Ollama host serves the request. By default that host is loopback, so they stay on this machine — but host is settable per call, so the server refuses file-bearing calls aimed anywhere else (see below).

Safety boundary

This is an LLM directing a server to read a disk, so the boundary is explicit:

  • Root allowlist. Only paths under OLLAMA_MCP_FILE_ROOTS (separated by the platform path delimiter — : on macOS/Linux, ; on Windows; defaults to the working directory) are readable. Paths are realpath-resolved before the check, so ../ traversal and symlink escapes both fail closed.
  • Sensitive-file deny-list, on by default: .env*, *.pem, *.key, id_rsa*, .ssh/**, .aws/**, .git/config, and anything named like a credential or secret. These can only be read by naming the file explicitly and passing allow_sensitive: true. A glob can never pull one in, whatever the flag says.
  • Loopback-only transfer. host is caller-controlled per call, so a file-bearing call to a non-loopback host would be an exfiltration path: read local files, POST them anywhere. Such a call is refused with REMOTE_FILE_TRANSFER_BLOCKED before the files are opened — nothing is read and nothing is sent. Loopback means localhost (and *.localhost), 127.0.0.0/8, and ::1; a LAN address or a hostname is not loopback even if it resolves back here. OLLAMA_MCP_ALLOW_REMOTE_FILES=1 is the explicit opt-in if you genuinely trust a remote host with these contents. Prompt-only calls to a remote host are unaffected.
  • Caps: 1MB per file, 4MB total, 50 files. Exceeding a cap is a hard error naming the file — never a silent drop.
  • Every file actually read is listed in the response, so an unexpected read is visible rather than silent.

Provenance, not immunity. File contents are untrusted input flowing into a model whose output comes back to your agent. The server wraps each file in explicit delimiters marking it as data rather than instructions. That makes the provenance legible; it does not make the output safe to act on blindly. Treat a dispatch result as untrusted text.


The thinking-token trap

Worth understanding, because it will bite you with any reasoning-capable model.

Reasoning tokens and answer tokens are drawn from the same num_predict budget. Set the cap too low with thinking enabled and the model spends the entire budget reasoning, then returns content: "" with done_reason: "length" — an HTTP 200, success-shaped, completely empty result. An agent will happily treat that as "the summary is empty" and carry on.

Three defences:

  1. think defaults to off. This server is for mechanical work where reasoning is cost without benefit. It's also capability-gated, so models that don't support thinking never receive the field.
  2. An unsafely low num_predict is raised to a workable floor (with a visible warning) when thinking is on. num_predict is a cap, not a target — raising it can't make a good run worse, but leaving it converts a guaranteed-empty result into a wasted model load.
  3. The exhausted case is detected and returned as an error, quantified, with ranked fixes — never as an empty success.

Configuration

Variable Default Purpose
OLLAMA_HOST http://localhost:11434 Daemon address
OLLAMA_MCP_CONFIG — Explicit config path
OLLAMA_MCP_DEFAULT_ROLE general Role used when model is omitted
OLLAMA_MCP_ROLE_<NAME> — Comma-separated chain; defines new roles
OLLAMA_MCP_ALIAS_<NAME> — Shorthand → real model name
OLLAMA_MCP_RANKING resident-then-smallest Tie-break policy among capable models
OLLAMA_MCP_TIMEOUT_MS 600000 Total request timeout
OLLAMA_MCP_CONNECT_TIMEOUT_MS 3000 Separate and short, so a down daemon fails fast
OLLAMA_MCP_REGISTRY_TTL_MS 60000 Model-list cache TTL
OLLAMA_MCP_MAX_OUTPUT_CHARS 100000 Output cap before truncation
OLLAMA_MCP_DEFAULT_THINK false See the trap above
OLLAMA_MCP_DEFAULT_TEMPERATURE 0 Determinism by default
OLLAMA_MCP_KEEP_ALIVE 10m Longer than Ollama's default; batch-friendly
OLLAMA_MCP_BATCH_CONCURRENCY 1 Within-group concurrency; 1 is VRAM-safe
OLLAMA_MCP_FILE_ROOTS cwd Readable roots, separated by the platform path delimiter (: POSIX, ; Windows)
OLLAMA_MCP_ALLOW_REMOTE_FILES false Permit file inputs when the target host is not loopback. Off by default — see the safety boundary above
OLLAMA_MCP_DETAIL concise Default response verbosity
DEBUG_OLLAMA_MCP — 1 for verbose stderr

Precedence everywhere: per-call argument > env var > config file > built-in default.

Ranking defaults to residency-first because an already-loaded model answers in seconds while a cold one can take ~12s to load — for high-volume mechanical work, "already in VRAM" beats every other signal.


Development

npm run build
npm run test:unit               # no daemon required
npm run check:no-model-literals # CI gate: no model names in src/
npm run smoke                   # end-to-end, needs a live daemon

The codebase is deliberately split: src/ is pure except for four files (index.ts, ollama/client.ts, registry/fetch.ts, config/load.ts). Model resolution, request shaping, response classification, batching and path validation are all total functions over plain data, tested against response fixtures captured from a real daemon. That's why the test suite needs no Ollama and CI is green on a clean runner.


Credits

Design inspiration for the file-aware tooling — reading files server-side so their contents never traverse the agent's context — came from Jadael/OllamaClaude. No code was copied; that project is AGPL-3.0 and this one is independently implemented under MIT.

License

MIT © Clickt Digital Marketing Inc.

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选