ollama-mcp
Enables AI agents to offload mechanical, high-token work to local Ollama models through MCP, with role-based model discovery, batch processing, and file-aware inputs.
README
ollama-mcp
An MCP server that lets an agent — Claude Code, or anything else speaking MCP — hand work to your local Ollama models.
The point is not to wrap the Ollama API. It is to give a frontier-model agent a cheap local tier it can delegate to: summarizing a 40KB log, extracting fields from twenty dumped files, reformatting JSON — mechanical, high-token, low-judgment work that burns expensive context for no benefit.
No model name appears anywhere in src/. Models are discovered live from the daemon and addressed by capability or by role. Pull a new model and it becomes usable within a minute, with no code change, no config edit, and no restart. That property is enforced in CI, not by convention.
Install
git clone https://github.com/Clickt-Digital-Marketing-Inc/ollama-mcp.git
cd ollama-mcp
npm install
npm run build
Register with Claude Code (user scope — available in every project):
claude mcp add --scope user ollama -- node /absolute/path/to/ollama-mcp/dist/src/index.js
Or add it to any MCP client's config:
{
"mcpServers": {
"ollama": {
"command": "node",
"args": ["/absolute/path/to/ollama-mcp/dist/src/index.js"]
}
}
}
Requires Node ≥ 20 and a running Ollama daemon. No configuration is needed to start — roles resolve by capability against whatever you already have installed.
Raise your client's tool timeout
This is the first thing most people hit, and it is not something this server can fix for you.
Local generation is slow by frontier-API standards: a large model can take ~12s just to load, and a long generation or a batch runs for minutes. MCP clients apply their own request timeout, and it is usually 60 seconds. When it fires, your client kills the call while the server is still working correctly — so the timeout_ms argument on the tools is not sufficient on its own. It bounds the server's wait on Ollama; it cannot extend your client's patience.
Codex (~/.codex/config.toml):
[mcp_servers.ollama]
command = "node"
args = ["/absolute/path/to/ollama-mcp/dist/src/index.js"]
startup_timeout_sec = 60
tool_timeout_sec = 900
Writing your own client with the TypeScript SDK — the third argument to callTool:
await client.callTool({ name: 'ollama_dispatch', arguments: {...} }, undefined,
{ timeout: 900_000, maxTotalTimeout: 900_000 });
If dispatches die at almost exactly 60 seconds, this is why — not the daemon, and not timeout_ms.
Tools
ollama_dispatch
One local generation. The main tool.
{ "prompt": "Summarize this changelog in 3 bullets.", "model": "role:summarize" }
Supports system, multi-turn messages, structured output via format (either "json" or a full JSON Schema), tool-calling passthrough, the usual sampling controls, and files/file_globs (below).
Every response ends with a metrics line:
[ollama model=<name> via=role:summarize→role:fast→caps:completion
tok=3120→412 dur=6.4s load=0.0s rate=64tok/s ctx=131072 done=stop think=off]
model and via are always present. A dispatcher that silently routes to the wrong model is the most expensive failure mode there is, so the resolution trail is never hidden.
ollama_dispatch_batch
Fan-out over many prompts. Items are grouped by resolved model and groups run sequentially, so a cold load is paid at most once per model instead of thrashing VRAM. Results come back in input order regardless of execution order, and one failing item never voids the run.
ollama_models
Discovery: capabilities, context window, size, residency. Two things worth knowing:
refresh: truere-reads the daemon after you pull something.explain_selector: "role:coder"dry-runs the resolver and prints the whole fallback chain without spending a generation. When routing surprises you, start here.
ollama_lifecycle
status / warm / unload. A large model can take ~12s to load and occupy tens of GB of VRAM, so warming before a batch and unloading afterwards are both real operations you'll want.
Choosing a model
Three grammars for the model field:
| Form | Example | Meaning |
|---|---|---|
| literal | qwen3:32b |
that exact model (bare names resolve to :latest, or to a single installed tag) |
| role | role:summarize |
an ordered fallback chain |
| capability | caps:vision+tools |
any installed model with all those capabilities |
| (omitted) | the configured default role |
Ambiguity is refused rather than guessed: if foo matches three installed tags, you get an error listing them. A wrong-model run is invisible in the output, so it is not something to coin-flip.
Roles
A role is an ordered chain. Each link is a literal name, another role, or a capability predicate — and the first link that resolves wins:
{
"roles": {
"coder": { "chain": ["some-coding-model", "some-fallback-model", "caps:completion"] }
}
}
If the preferred model isn't installed, the chain falls through. This is the future-proofing story: the chain documents your intent even when the model isn't there yet, and starts routing to it the moment you pull it.
Built-in roles — all defined purely as capability predicates, so they work against any install: general, fast, big, reasoner, coder, vision, tools, embed, summarize, extract.
Adding a model
Three ways, in increasing order of commitment:
- Just pull it.
ollama pull <model>. Within ~60s it joins the candidate pool for every role and capability it qualifies for, and is addressable by name. Nothing else to do. - Pin a preference. Add it to the front of a role's
chaininollama-mcp.config.json. - Use an env var, no file at all.
OLLAMA_MCP_ROLE_CODER="model-a,model-b,caps:completion".OLLAMA_MCP_ROLE_<NAME>is parsed generically, so this also creates roles —OLLAMA_MCP_ROLE_TRANSLATOR=...gives yourole:translatorwith no code change.
Config is discovered at $OLLAMA_MCP_CONFIG, then ./ollama-mcp.config.json, then ~/.config/ollama-mcp/config.json; first hit wins. A malformed config is non-fatal — the server logs, falls back to defaults, and warns on the first response, because a typo should never take the server down.
File-aware inputs
files and file_globs make the server read files and feed them to the local model:
{ "prompt": "Extract every TODO with its file and line.",
"file_globs": ["src/**/*.ts"],
"model": "role:extract" }
File contents never enter the calling agent's context — only the model's distilled answer comes back. For large inputs this is the whole reason the server is worth having.
That is a statement about your agent's context, not about the network. File contents are sent to whichever Ollama host serves the request. By default that host is loopback, so they stay on this machine — but host is settable per call, so the server refuses file-bearing calls aimed anywhere else (see below).
Safety boundary
This is an LLM directing a server to read a disk, so the boundary is explicit:
- Root allowlist. Only paths under
OLLAMA_MCP_FILE_ROOTS(separated by the platform path delimiter —:on macOS/Linux,;on Windows; defaults to the working directory) are readable. Paths arerealpath-resolved before the check, so../traversal and symlink escapes both fail closed. - Sensitive-file deny-list, on by default:
.env*,*.pem,*.key,id_rsa*,.ssh/**,.aws/**,.git/config, and anything named like a credential or secret. These can only be read by naming the file explicitly and passingallow_sensitive: true. A glob can never pull one in, whatever the flag says. - Loopback-only transfer.
hostis caller-controlled per call, so a file-bearing call to a non-loopback host would be an exfiltration path: read local files, POST them anywhere. Such a call is refused withREMOTE_FILE_TRANSFER_BLOCKEDbefore the files are opened — nothing is read and nothing is sent. Loopback meanslocalhost(and*.localhost),127.0.0.0/8, and::1; a LAN address or a hostname is not loopback even if it resolves back here.OLLAMA_MCP_ALLOW_REMOTE_FILES=1is the explicit opt-in if you genuinely trust a remote host with these contents. Prompt-only calls to a remote host are unaffected. - Caps: 1MB per file, 4MB total, 50 files. Exceeding a cap is a hard error naming the file — never a silent drop.
- Every file actually read is listed in the response, so an unexpected read is visible rather than silent.
Provenance, not immunity. File contents are untrusted input flowing into a model whose output comes back to your agent. The server wraps each file in explicit delimiters marking it as data rather than instructions. That makes the provenance legible; it does not make the output safe to act on blindly. Treat a dispatch result as untrusted text.
The thinking-token trap
Worth understanding, because it will bite you with any reasoning-capable model.
Reasoning tokens and answer tokens are drawn from the same num_predict budget. Set the cap too low with thinking enabled and the model spends the entire budget reasoning, then returns content: "" with done_reason: "length" — an HTTP 200, success-shaped, completely empty result. An agent will happily treat that as "the summary is empty" and carry on.
Three defences:
thinkdefaults to off. This server is for mechanical work where reasoning is cost without benefit. It's also capability-gated, so models that don't support thinking never receive the field.- An unsafely low
num_predictis raised to a workable floor (with a visible warning) when thinking is on.num_predictis a cap, not a target — raising it can't make a good run worse, but leaving it converts a guaranteed-empty result into a wasted model load. - The exhausted case is detected and returned as an error, quantified, with ranked fixes — never as an empty success.
Configuration
| Variable | Default | Purpose |
|---|---|---|
OLLAMA_HOST |
http://localhost:11434 |
Daemon address |
OLLAMA_MCP_CONFIG |
— | Explicit config path |
OLLAMA_MCP_DEFAULT_ROLE |
general |
Role used when model is omitted |
OLLAMA_MCP_ROLE_<NAME> |
— | Comma-separated chain; defines new roles |
OLLAMA_MCP_ALIAS_<NAME> |
— | Shorthand → real model name |
OLLAMA_MCP_RANKING |
resident-then-smallest |
Tie-break policy among capable models |
OLLAMA_MCP_TIMEOUT_MS |
600000 |
Total request timeout |
OLLAMA_MCP_CONNECT_TIMEOUT_MS |
3000 |
Separate and short, so a down daemon fails fast |
OLLAMA_MCP_REGISTRY_TTL_MS |
60000 |
Model-list cache TTL |
OLLAMA_MCP_MAX_OUTPUT_CHARS |
100000 |
Output cap before truncation |
OLLAMA_MCP_DEFAULT_THINK |
false |
See the trap above |
OLLAMA_MCP_DEFAULT_TEMPERATURE |
0 |
Determinism by default |
OLLAMA_MCP_KEEP_ALIVE |
10m |
Longer than Ollama's default; batch-friendly |
OLLAMA_MCP_BATCH_CONCURRENCY |
1 |
Within-group concurrency; 1 is VRAM-safe |
OLLAMA_MCP_FILE_ROOTS |
cwd | Readable roots, separated by the platform path delimiter (: POSIX, ; Windows) |
OLLAMA_MCP_ALLOW_REMOTE_FILES |
false |
Permit file inputs when the target host is not loopback. Off by default — see the safety boundary above |
OLLAMA_MCP_DETAIL |
concise |
Default response verbosity |
DEBUG_OLLAMA_MCP |
— | 1 for verbose stderr |
Precedence everywhere: per-call argument > env var > config file > built-in default.
Ranking defaults to residency-first because an already-loaded model answers in seconds while a cold one can take ~12s to load — for high-volume mechanical work, "already in VRAM" beats every other signal.
Development
npm run build
npm run test:unit # no daemon required
npm run check:no-model-literals # CI gate: no model names in src/
npm run smoke # end-to-end, needs a live daemon
The codebase is deliberately split: src/ is pure except for four files (index.ts, ollama/client.ts, registry/fetch.ts, config/load.ts). Model resolution, request shaping, response classification, batching and path validation are all total functions over plain data, tested against response fixtures captured from a real daemon. That's why the test suite needs no Ollama and CI is green on a clean runner.
Credits
Design inspiration for the file-aware tooling — reading files server-side so their contents never traverse the agent's context — came from Jadael/OllamaClaude. No code was copied; that project is AGPL-3.0 and this one is independently implemented under MIT.
License
MIT © Clickt Digital Marketing Inc.
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。