Payment Orchestrator MCP Server
Enables agents to interrogate payment routing decisions through six tools: route transactions, explain decisions, simulate scenarios, inspect segment evidence, normalize decline codes, and review backtest summaries. It provides read-only access to the routing engine, allowing natural-language queries without modifying any decisions.
README
Payment routing orchestrator
Evidence in the hot path, AI at the edges.
Given a card authorization about to be sent, it decides which payment service provider should receive it — on empirical approval evidence, at a fee tolerance the operator states in real units. If the attempt declines, a state machine keyed on the decline's error class decides what happens next: same PSP later, failover now, a different channel, or stop. The routing decision itself is deterministic and auditable; no language model runs inside it.
Built by Bryan Rodríguez Abarca — a payments product manager who decides where a model earns its place, and where it doesn't. More at https://vryahn.com/work/routing
Built and shipped inside Claude Code: the engine, the AI edge layer, the eval harness, the web UI, the API and this README were produced in agentic sessions — one orchestrating session delegating to subagents (backend, UI, publish, case study) — and are gated by the tests and evals you can run yourself. The engineering discipline is the same one described in nutri.: conventions the agent must load, structural guardrails, and a machine — not a promise — as the definition of done. The guardrails caught what would otherwise have shipped: a rewrite that dropped every API path (found by post-deploy verification), a UI mislabeling valid issuers as unseen, and two wrong assumptions of the author's (a page-count rule, a DNS setup) that subagents refused to act on.
Live demo https://orchestrator.vryahn.com · Case study
https://vryahn.com/work/routing · API api/README.md ·
MCP MCP.md
Out-of-sample replay on 84,011 TEST transactions (days 22–31, tables trained on
days 1–21): expected approval 72.02% at cost_bias=0 against 66.13%
actually observed — +5.89 pp, directional, not an A/B result. See
Limitations.
Run it
Python 3.11. requirements.txt is the runtime — fastapi plus the standard
library, and the only thing Vercel installs. requirements-dev.txt adds the
offline stack (duckdb, pandas, numpy, pyarrow, mcp, uvicorn, httpx) needed to
regenerate data, run the backtest, serve MCP, or test locally.
python3.11 -m venv .venv && .venv/bin/pip install -r requirements-dev.txt
.venv/bin/python cli.py --txn-file demo_transactions.json # 8 decision-boundary cases
.venv/bin/uvicorn api.index:app --reload --port 8000 # API on /api/*, UI from public/
The tables the engine reads (routing_tables.json, routing_meta.json) are
committed, so a fresh clone routes — and deploys — without the offline
pipeline. Regenerate them only if the data model changes:
.venv/bin/python synth_attempts.py # seeded ~300k attempts -> attempts.parquet
.venv/bin/python build_routing_tables.py # -> routing_tables.json, routing_meta.json
.venv/bin/python backtest.py --json # -> backtest_summary.json
.venv/bin/python tests.py # engine
.venv/bin/python tests_ai.py # AI edges + HTTP contract
.venv/bin/python evals/decline_eval.py # normalizer against the golden set
Architecture
flowchart LR
subgraph offline["OFFLINE — batch, once per rebuild"]
G["synth_attempts.py<br/>seeded generator"] --> A[("attempts.parquet<br/>1 row = 1 attempt, ~300k")]
A -->|"build_routing_tables.py"| T[("routing_tables.json + routing_meta.json<br/>segment x PSP: n, approvals, p_hat, Wilson LB<br/>4-level hierarchy")]
A -->|"backtest.py — train d1-21, test d22-31"| B["backtest_summary.json<br/>out-of-sample lift"]
end
subgraph edgein["EDGE IN — language to enum"]
RAW["raw PSP decline<br/>ISO 8583 / decline_code / refusalReason / bank prose"] --> N{"decline_normalizer.py<br/>table -> LLM -> safe fallback"}
EV["evals/ — 48 golden declines<br/>accuracy by route, hallucination gate"] -.->|"scores"| N
end
subgraph online["ONLINE — pure engine, never touches raw data"]
X["txn: amount, bin6/issuer, funding,<br/>channel, attempt #, error history"] --> D{"decide(txn, config)"}
T --> D
N -->|"error_class"| D
D --> S1["1. resolve segment per PSP<br/>walk L0 to L3 until n >= min_support"]
S1 --> S2["2. score = Wilson LB x amount x (1 - fee)<br/>= expected net collected"]
S2 --> S3["3. pick PSP — cost_bias 0..1 maps to<br/>fee tolerance 0..10pp; psps_down excluded"]
S3 --> S4["4. retry state machine<br/>keyed on last error_class"]
S4 --> R["Decision: route_psp, eligible_psps with scores,<br/>retry policy, reasoning lines"]
end
R --> OPS["ops.py — route, explain, simulate,<br/>evidence, normalize, backtest"]
B --> OPS
OPS --> CLI["cli.py"]
OPS --> API["api/index.py — FastAPI on Vercel<br/>+ public/ web UI, same origin"]
OPS --> MCP["mcp_server.py — 6 MCP tools"]
Why this design
- Offline/online split.
decide(txn, config) -> Decisionis pure. It loads pre-materialized tables once and never reads raw attempts, so the decision is microseconds, testable without a database, and auditable after the fact. It is the same boundary you would draw in production between a reconciliation layer and a routing layer. - Segment hierarchy with fallback. L0 is
gateway_group × funding × issuer_bucket × amount_band; L3 isgateway_groupalone. Support is resolved per PSP, one dimension at a time (amount band → issuer → funding), until a cell clearsmin_support(default 200), and the level used is reported with the decision. The channel is first-class and never dropped: user-present and off-session are different worlds. - Wilson lower bound, not the raw rate. A segment with 3/3 approvals is not a 100% segment. The bound shrinks toward zero as support thins, so a well-evidenced 78% beats a lucky 100% without a separate confidence rule bolted on.
cost_biasas an explicit knob, in real units. The trade-off is stated as "percentage points of approval I will give up for a cheaper PSP" —tolerance = cost_bias × 10pp, and the cheapest PSP within that tolerance of the best approver wins. A blended score would let a fractions-of-a-point fee difference silently outvote a double-digit approval gap; a tolerance filter cannot. The backtest prices the knob: 72.02% / +5.89 pp approval atcost_bias=0, 71.56% / +5.43 pp at 0.5, 69.57% / +3.44 pp at 1.0.- Retry by error class, not by a blind counter.
insufficient_fundsis an account problem and retries the same PSP on the next billing window;bank_auth_requiredoff-session cannot be satisfied without the customer, so it reschedules to a user-present channel instead of burning attempts;fraud_riskstops the chain permanently;generic_declinefails over to the next PSP by score. An unrecognized class degrades to the generic failover policy and says so.
Where AI belongs — and where it does not
There is no LLM inside decide(). Money should not move on a sampled token.
The language model is confined to the two edges where natural language actually
is the problem.
In — decline_normalizer.py. Each PSP declines in its own dialect: ISO 8583
numerics, a Stripe-like decline_code, an Adyen-like refusalReason, or raw
bank prose. The retry state machine keys on one enum, so the dialects have to
collapse before the engine sees them. A deterministic table handles the codes
that carry volume — confidence 1.0, no latency, no cost. Only a table miss
reaches the model chain (Gemini, then Mistral), which answers under a
constrained enum schema. Anything outside the enum, or below 0.6 confidence, is
discarded in favour of generic_decline, which is the retry policy's own safe
default. The repo runs green with no API keys set.
Measured, not trusted — evals/. 48 golden declines: ~60% table hits, ~40%
deliberately off-table (misspellings, verbose bank text, unusual codes) plus a
few genuinely ambiguous ones where generic_decline is the correct answer. The
evals/baseline.json records two baselines. Table-only (no keys): 32/48 =
66.67%, i.e. 100% on the 28 table-route cases and the safe generic_decline
default on the 20 that fall through. LLM (keys configured, run against the
deployed API with --remote): 48/48 = 100% — 28 table, 19 answered by
gemini-3.6-flash, 1 by the low-confidence fallback where generic_decline was
the expected answer. The runner reports accuracy by route and a per-class
confusion table, asserts zero hallucinations, and fails the build if accuracy
drops more than 2 pp below the matching baseline.
Out — mcp_server.py. Six MCP tools — route_transaction,
explain_decision, simulate, segment_evidence, normalize_decline,
backtest_summary — let an agent operate the engine in English. The agent can
interrogate every decision and change none of them. See MCP.md.
Limitations
- Data is synthetic. The structure is designed to make routing decisions non-trivial, not to reproduce any real portfolio.
- No live PSP connectors: the engine decides, it does not send.
- No fraud scoring, 3DS orchestration, network tokens, or scheme retry-rule enforcement.
- The backtest is directional. Historical routing was not randomized, capacity limits are not modeled, and "expected approval" is a TRAIN-period Wilson-LB rate applied to TEST volume — not a live A/B result.
- Tables pool all attempts while the backtest trains and replays on first attempts only; first-attempt-only production tables would be the next fix.
- The LLM eval is 48 cases and one run; 100% on a golden set this small is a gate against regressions, not a claim about the long tail in production.
File map
| file | purpose |
|---|---|
synth_attempts.py |
seeded generator for attempts.parquet; channel mix, PSP fees, approval model, error mix and retry behavior documented at the top |
build_routing_tables.py |
builds routing_tables.json (segment × PSP: n, approvals, p_hat, wilson_lb, all 4 levels) and routing_meta.json (amount-band edges, issuer whitelist, bin6 → issuer map, PSP fees, error-class enum) |
orchestrator.py |
the engine: decide(txn, config) -> Decision, Config dataclass. Stdlib only |
decline_normalizer.py |
PSP decline dialects → the engine's error_class enum: table first, LLM on a miss, safe fallback |
ops.py |
the operator functions (route, explain, simulate, evidence, normalize, backtest) shared by the API and MCP |
api/index.py |
FastAPI on Vercel; contract in api/README.md |
mcp_server.py |
MCP stdio server, six tools; see MCP.md |
cli.py |
CLI front end: one transaction via flags, or a batch via --txn-file |
public/ |
static web UI, served by Vercel from the same origin as the API |
demo_transactions.json |
8 decision-boundary transactions with why_interesting notes |
backtest.py |
TRAIN days 1–21 / TEST days 22–31 replay at cost_bias 0 / 0.5 / 1.0; --json rewrites backtest_summary.json |
evals/ |
48 golden declines, the scoring runner, and the recorded baseline |
tests.py / tests_ai.py |
assert-based checks: the engine, then the AI edges and the HTTP contract |
Bryan Rodríguez Abarca · vryahn.com · Started from a technical exercise, generalized as a personal project. Synthetic data.
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。