orchestrator-mcp

orchestrator-mcp

Enables sending each request to the model configured for that kind of work, routing through a configurable set of capabilities and model deployments with a consistent response envelope.

Category
访问服务器

README

orchestrator-mcp

Capability-routed MCP server: send each request to the model configured for that kind of work

PyPI Python License Tests

Quick Start · How It Works · Tools · Configuration · Guardrails · Troubleshooting · Contributing

Research to one model, coding to another, cheap extraction to a third — and every answer comes back through the same validated envelope. Pointing a capability at your own deployment is a YAML edit. There is no code to change.

Published to PyPI as orchestrator-mcp-server — the shorter name is an empty registered project owned by someone else. The import package is orchestrator_mcp.

Quick Start

Write a config.yaml — start from config.example.yaml — and check that it loads:

ORCHESTRATOR_CONFIG=config.yaml uvx --from orchestrator-mcp-server python -c "from orchestrator_mcp.server import build_server; build_server(); print('config ok')"

A bad config fails here rather than at request time: every deployment must route to a declared capability, every capability must have a deployment behind it, and every fallback must name a real capability.

Claude Code

claude mcp add orchestrator --env ORCHESTRATOR_CONFIG=$PWD/config.yaml -- uvx orchestrator-mcp-server

Codex

In ~/.codex/config.toml:

[mcp_servers.orchestrator]
command = "uvx"
args = ["orchestrator-mcp-server"]
env = { ORCHESTRATOR_CONFIG = "/absolute/path/to/config.yaml" }

Both speak stdio, and every tool result is returned as structured content and as JSON text, so a client that reads only one of the two still gets the whole envelope.

Note Provider keys are read from the environment the server process gets, which is the client's environment — not your shell's. If a capability returns auth_failed while the same config works from a terminal, add the key to the client's env block.

Homebrew

brew tap crAK1644/tap
brew install orchestrator-mcp-server

That puts orchestrator-mcp-server on your PATH, so a client can call it directly instead of going through uvx:

claude mcp add orchestrator --env ORCHESTRATOR_CONFIG=$PWD/config.yaml -- orchestrator-mcp-server

Note Apple Silicon pours a prebuilt bottle. Everything else builds every Python dependency from source, including several Rust crates, which takes around fifteen minutes — uvx orchestrator-mcp-server is the same program from a prebuilt wheel in about a second.

From a checkout

uv sync && uv run pytest -q
claude mcp add orchestrator --env ORCHESTRATOR_CONFIG=$PWD/config.yaml -- uv run --directory $PWD orchestrator-mcp-server

Documentation

Everything lives in this file and in the annotated config. The links below are the fast path to a specific answer.

Getting started

Using it

Understanding it

Development

How It Works

A capability is a LiteLLM model_name alias group. Several deployments can share one name, and litellm.Router already load-balances, retries, cools down, and falls back across them. So the routing engine is the config file:

model_list:
  - model_name: coding                  # capability, not a model
    litellm_params:
      model: anthropic/claude-sonnet-4-5
      api_key: os.environ/ANTHROPIC_API_KEY

  - model_name: coding                  # same capability, your own box
    litellm_params:
      model: openai/qwen-coder
      api_base: http://vllm.internal:8000/v1
      api_key: os.environ/LOCAL_VLLM_KEY

The pieces:

  • The config file is the router. Capabilities, deployments, fallbacks, and retry policy are all LiteLLM's own schema, so litellm --config config.yaml runs on it unchanged.
  • The caller states its capability. There is no intent classifier — the caller is already a language model and knows whether it is asking a coding question. Paying a second model to guess what the first one already knows buys a cost increase and a new failure mode.
  • One envelope for every outcome. Success, refusal, truncation, and timeout all return the same shape, so callers branch on a field instead of parsing prose.

Tools

ask

Argument Notes
capability Enum, built from your config. Bad values are rejected by the protocol layer.
prompt Required. Capped by limits.max_prompt_chars.
context Source material. When set, the model is told to answer only from it and to abstain otherwise.
system Extra instructions. Applied before the server's own directives, so it cannot disable them. Capped by limits.max_system_chars.
response_schema JSON Schema ("type": "object"), as an object or a JSON string. Switches on structured mode. Capped by limits.max_schema_chars — it is inlined into the prompt verbatim.
temperature Pinned to 0 whenever response_schema is set.
max_output_tokens Capped by limits.max_output_tokens.

list_capabilities

What each capability is for, the deployments behind it, and where it falls back.

The response envelope

{
  "ok": true,
  "content": "…",
  "data": null,
  "insufficient_context": false,
  "capability_requested": "coding",
  "model_used": "anthropic/claude-sonnet-4-5",
  "fallback_used": false,
  "finish_reason": "stop",
  "usage": { "prompt_tokens": 10, "completion_tokens": 20, "total_tokens": 30, "cost_usd": 0.0002 },
  "latency_ms": 412,
  "error": null
}

content holds the prose answer and is null in structured mode and on failure; data holds the validated object and is set only in structured mode; error is { code, message } whenever ok is false.

finish_reason rides along on failures too, when a provider replied at all — it is diagnosis rather than an answer, and length tells you to raise max_output_tokens instead of going looking for a bug. It is null when nothing was received.

Error codes

A closed set, so callers branch on a value instead of matching substrings.

Code Means
invalid_request Rejected at the boundary, before any provider was called.
no_deployment Every deployment for the capability is cooled down or absent.
upstream_error The provider failed, or returned no usable completion.
rate_limited The provider rate-limited the call.
context_exceeded The request did not fit the model's context window.
schema_validation_failed The reply never matched your schema, repairs included.
timeout limits.request_timeout_s elapsed for the whole call.
content_filtered The provider filtered the completion.
auth_failed Bad or missing credentials.
output_truncated The model hit the output limit mid-answer.

error.message is bounded at 500 characters and never quotes the rejected output back at you.

System Requirements

  • Python 3.11, 3.12, or 3.13
  • uv — or any installer, if you would rather use pip
  • An MCP client that speaks stdio (Claude Code, Codex, or your own)
  • At least one provider you hold credentials for, or a local endpoint
  • Network access to whatever your config.yaml points at — nothing else phones home

A local Ollama endpoint works and costs nothing, which makes it a reasonable way to try the server before wiring up paid providers:

model_list:
  - model_name: fast
    litellm_params:
      model: ollama_chat/qwen2.5:7b
      api_base: http://localhost:11434

Configuration

ORCHESTRATOR_CONFIG points at the file; it defaults to config.yaml in the working directory. Keep it out of version control — it holds your endpoints.

capabilities:
  coding: "Writing, refactoring, reviewing, and debugging code."
  fast: "Cheap, low-latency answers. Classification, extraction, short replies."

model_list:
  - model_name: coding
    litellm_params:
      model: anthropic/claude-sonnet-4-5
      api_key: os.environ/ANTHROPIC_API_KEY

  - model_name: fast
    litellm_params:
      model: anthropic/claude-haiku-4-5-20251001
      api_key: os.environ/ANTHROPIC_API_KEY

router_settings:
  num_retries: 2
  cooldown_time: 60
  fallbacks:
    - coding: [fast]

Keys are referenced as os.environ/NAME and read at request time. Never inline the value.

Limits

Every one of these is a boundary the caller cannot cross, checked before a provider is called. A nonsensical value fails at startup rather than mid-request.

Key Default What it bounds
max_prompt_chars 100000 The prompt argument
max_context_chars 400000 The context argument
max_system_chars 10000 Caller instructions, which reach the prompt verbatim
max_schema_chars 20000 response_schema, which is inlined verbatim
max_output_tokens 4096 Completion length
request_timeout_s 120 The whole call — retries, fallback, and repairs share it
schema_repair_attempts 1 Retries given to a schema-violating reply

Guardrails, and their limits

This server sees a prompt and a completion. It has no ground truth, so it cannot verify factual claims, and nothing here should be read as a hallucination detector. What it does enforce:

  • Shape is validated, not assumed. Structured replies are checked against your schema locally with jsonschema, regardless of whether the provider claims to enforce response_format. A violation is a failure, not a payload.
  • Bounded repair. An invalid structured reply gets limits.schema_repair_attempts retries carrying the validator's complaint, then fails as schema_validation_failed. Never a best-effort half-parsed object.
  • An unfinished answer is a failure, not a short answer. A completion cut off by the token limit comes back as output_truncated with content: null, and one the provider filtered as content_filtered. Neither is returned as prose, because a half answer reads exactly like a whole one.
  • The error tells you what broke, not what the model wrote. error.message gives the failing path and constraint (schema violation at answer/city: failed the 'maxLength' constraint) and is capped at 500 characters. The rejected value itself goes only back to the model that produced it, in the repair turn.
  • request_timeout_s bounds the call. Retries, cross-capability fallback, and repair turns all spend from one budget, so 120 cannot become 360.
  • Abstention is typed. With context set, the model is given an explicit way to say the material does not support an answer; it arrives as insufficient_context, not as prose you have to pattern-match.
  • The server never ghostwrites. When ok is false, content and data are both null. It will not put a "Sorry, I couldn't…" string where a model's answer goes, because callers cannot tell those apart. Enforced by an assertion on every response and covered by tests.
  • Degradation is visible. fallback_used and model_used always ride along, so an answer served by the backup after the primary died never passes as the intended one.
  • The caller cannot smuggle a model. There is no free-form model parameter, only the capability enum. Routing stays operator-controlled.
  • Boundaries reject early. Unknown capability, oversized prompt or system, empty prompt, and a malformed or oversized response_schema all fail before a provider is called.

Two known gaps. The MCP SDK drops unknown arguments before the handler sees them, so an unrecognized key is ignored at the protocol layer rather than rejected — direct calls into Orchestrator.ask do reject it. And a response_schema containing a pathological pattern can burn CPU on the event loop during validation: the schema is size-capped but not analyzed, so treat schema authorship as a trusted operation.

Tests

uv run pytest -q

76 tests, no network — deployments are stubbed with LiteLLM's mock_response, and the shapes it cannot express (no choices, null content, a truncated or filtered reply) are stubbed as raw ModelResponse objects. Includes the rate-limit-then-fallback path and the cooled-down-group path.

Because all of that is stubbed, it proves the orchestrator's logic and nothing about your providers. For that:

uv run python smoke_live.py

Real calls against your config.yaml, roughly four short ones per capability, so it costs a little money — run it deliberately, not in CI. It checks the things only a live endpoint can answer: whether response_format survives the round trip, whether the model honours the abstention path instead of inventing, and what the provider actually sends as finish_reason when it runs out of room. Name capabilities as arguments to check only some (uv run python smoke_live.py fast).

Troubleshooting

Symptom Cause
config not found: config.yaml No config where the server is looking. Set ORCHESTRATOR_CONFIG to an absolute path — a client launches the server from its own working directory, not yours.
auth_failed from the client, but the same config works in a terminal The key is in your shell, not the client's. Add it to the client's env block.
Startup fails naming a capability A capability has no deployment, a deployment names an undeclared capability, or a fallback points at one that does not exist.
no_deployment Every deployment in the group is cooling down after failures. router_settings.cooldown_time controls how long.
output_truncated The answer did not fit. Raise max_output_tokens, in the request or in limits.
schema_validation_failed after repairs The model cannot produce your schema. Simplify it, or route the capability at a model that supports response_format natively.
timeout on a local model request_timeout_s covers the entire call including model load. A 7B model on a laptop wants more than the default.

Bug Reports

Open an issue with the envelope you got back — it carries error.code, model_used, fallback_used, finish_reason, and latency_ms, which is most of a diagnosis already. Add your config.yaml with the keys removed.

For anything that looks like a routing or retry problem, LiteLLM's own trace is the useful attachment:

LITELLM_LOG=DEBUG uv run python smoke_live.py 2>debug.log

It writes to stderr only, so stdout stays clean and the MCP stream is unaffected.

Contributing

Issues and pull requests are welcome. The bar for a change is a test that fails without it — the suite runs offline, so there is no key to obtain and no cost to pay. Keep config.yaml out of your commits.

  1. Fork and branch.
  2. Make the change.
  3. Add the test that fails without it.
  4. uv run pytest -q.
  5. Open the pull request.

If you are adding a capability to your own setup, you do not need a pull request: it is a model_list entry.

Releasing

Bump version in pyproject.toml, then publish a GitHub Release tagged vX.Y.Z. release.yml runs the suite, checks the tag against the packaged version, and uploads to PyPI.

There is no API token in this repo. PyPI is configured as a Trusted Publisher for this workflow, so it mints a short-lived credential from the job's OIDC identity — nothing to store, nothing to rotate, nothing to leak.

Not included

Semantic/embedding routing and RouteLLM-style predictive routing (the caller states its capability); Redis-backed distributed cooldown state (single process — LiteLLM enables it via config when you need a second node); streaming (MCP tool results return whole); sampling/createMessage loops; PII redaction and telemetry callbacks (available as LiteLLM callbacks when a requirement names one).

License

MIT — see LICENSE. Use it, fork it, ship it in something commercial. The only thing it asks is that the copyright notice travels with the code.

Support


Built on LiteLLM, Pydantic, and the Python MCP SDK.

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选