mcp-decision-lab

mcp-decision-lab

MCP server that enables structured decision-making using weighted decision matrices, sensitivity analysis, and robust recommendations.

Category
访问服务器

README

mcp-decision-lab

tests

MCP server for structured decision-making — weighted decision matrices, criteria scoring, sensitivity analysis and a defensible recommendation.

Why

LLMs are good at listing pros and cons and then picking whatever "feels" right. That reasoning is opaque, unstable, and impossible to audit: change one adjective in the prompt and the answer flips. mcp-decision-lab is a thinking tool — the model uses it to structure its own reasoning as an explicit weighted decision matrix. Every score requires a written rationale, weights are normalized and transparent, and the final analyze step does real math: exact (closed-form, not brute-force) sensitivity analysis that tells you which criterion weight would flip the winner and at what value — so the recommendation comes with a robustness verdict instead of vibes.

Tools

Tool Arguments Returns
start_decision question: str, options: list[str] (≥2), criteria: list[dict] (≥2, each {"name", "weight", "higher_is_better"?}) New decision_id, criteria with weights normalized to sum 1.0, empty-cell list, next-step instruction
score_option decision_id, option, criterion, score: float (0–10), rationale: str (≥10 chars, required) The recorded cell, remaining missing cells, next step
get_matrix decision_id Full matrix: raw scores, effective scores, weighted scores, per-option totals, rationales, missing cells
analyze decision_id (matrix must be complete) Ranking + margin, per-criterion sensitivity (exact flip weight, direction, new winner, decisive flag), strengths/weaknesses per option, robustness verdict (robust/fragile), textual recommendation
list_decisions — All sessions with status: pending / complete / analyzed

Criteria where a high raw score is bad (cost, risk, complexity) take "higher_is_better": false — the effective score becomes 10 - score automatically.

How it works

flowchart TD
    A[start_decision<br/>question + options + weighted criteria] --> B[weights normalized to sum 1.0<br/>session dec-xxxx persisted to disk]
    B --> C[score_option x N<br/>one cell = option x criterion,<br/>score 0-10 + mandatory rationale]
    C -->|cells missing| C
    C -->|matrix complete| D[get_matrix<br/>review raw / weighted scores]
    D --> E[analyze]
    E --> F[ranking + margin<br/>weighted totals]
    E --> G["sensitivity analysis<br/>closed-form flip weight per criterion:<br/>solve gap(w') = w'·d + (1-w')·r = 0"]
    E --> H[strengths / weaknesses<br/>best and worst criterion per option]
    F --> I[robustness verdict<br/>robust vs fragile under ±50% weight shifts]
    G --> I
    H --> I
    I --> J[defensible recommendation]

The sensitivity math. With normalized weights, changing criterion c's weight from w to w′ (renormalizing the others proportionally) makes every option's total linear in w′. For the winner A vs a challenger B the gap is gap(w′) = w′·d + (1−w′)·r, where d is their effective-score difference on c and r is their weight-scaled difference on everything else. Solving gap(w′) = 0 gives the exact flip weight w* = r/(r−d) — reported only if a region of [0, 1] exists where the challenger strictly wins. If no ±50% relative change to any single weight flips the winner, the verdict is robust; otherwise fragile, naming the criteria the decision hinges on.

Quickstart

pip install -e .

Claude Desktop — add to claude_desktop_config.json:

{
  "mcpServers": {
    "decision-lab": {
      "command": "python",
      "args": ["/absolute/path/to/server.py"]
    }
  }
}

Claude Code:

claude mcp add decision-lab -- python /absolute/path/to/server.py

Sessions persist as JSON in ~/.mcp-decision-lab/decisions.json (override the directory with the DECISION_LAB_DIR environment variable).

Example session

User: Help me pick a database for the new SaaS backend — Postgres, MongoDB or DynamoDB. Cost matters most, then scalability, then how well the team knows it.

The model starts a session:

start_decision(
    question="Which database should we use for the new SaaS backend?",
    options=["Postgres", "MongoDB", "DynamoDB"],
    criteria=[
        {"name": "cost", "weight": 0.5, "higher_is_better": False},
        {"name": "scalability", "weight": 0.3},
        {"name": "team-familiarity", "weight": 0.2},
    ],
)
# → {"decision_id": "dec-34b5", "cells_total": 9,
#    "next_step": "Score each option against each criterion using score_option ..."}

Then scores all 9 cells, each with a rationale:

score_option("dec-34b5", "Postgres", "cost", 3,
    "Managed Postgres (RDS/Neon) is cheap and predictable at our scale")
score_option("dec-34b5", "DynamoDB", "scalability", 10,
    "Effectively unlimited managed horizontal scale")
score_option("dec-34b5", "DynamoDB", "team-familiarity", 4,
    "Only one engineer knows single-table design patterns")
# ... 6 more cells ...
# → {"missing_cells": [], "next_step": "Matrix complete — call analyze('dec-34b5') ..."}

And analyzes:

analyze("dec-34b5")
{
  "ranking": [
    {"rank": 1, "option": "Postgres", "total": 7.1},
    {"rank": 2, "option": "DynamoDB", "total": 5.8},
    {"rank": 3, "option": "MongoDB", "total": 5.6}
  ],
  "winner": "Postgres",
  "margin": 1.3,
  "sensitivity": [
    {"criterion": "cost", "weight": 0.5, "decisive": true,
     "flip": {"flip_weight": 0.1176, "direction": "decrease",
              "new_winner": "DynamoDB", "within_50pct_band": false}},
    {"criterion": "scalability", "weight": 0.3, "decisive": true,
     "flip": {"flip_weight": 0.4717, "direction": "increase",
              "new_winner": "DynamoDB", "within_50pct_band": false}},
    {"criterion": "team-familiarity", "weight": 0.2, "decisive": false, "flip": null}
  ],
  "decisive_criteria": ["cost", "scalability"],
  "robustness": "robust",
  "recommendation": "Choose Postgres (7.1 weighted) over DynamoDB (5.8); margin 1.3. Postgres is strongest on 'team-familiarity' (9/10) and weakest on 'scalability' (6/10). The result is robust: no ±50% relative change to any single criterion weight changes the winner. (Extreme shifts could still flip it: 'cost' would have to decrease from 0.5 to 0.1176 for DynamoDB to win; 'scalability' would have to increase from 0.3 to 0.4717 for DynamoDB to win.)"
}

Model (to user): Postgres wins at 7.1 weighted vs DynamoDB's 5.8, and the result is robust: no reasonable (±50%) shift in any single weight changes the winner. It would take valuing scalability at ~0.47 instead of 0.3 — or cost dropping below ~0.12 — for DynamoDB to overtake. If you genuinely expect hockey-stick scale, revisit; otherwise Postgres is the defensible choice.

Development

pip install -e ".[dev]"
python -m pytest

Tests exercise core.py only and run without the mcp package installed.

License

MIT

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选
mcp-server-qdrant

mcp-server-qdrant

这个仓库展示了如何为向量搜索引擎 Qdrant 创建一个 MCP (Managed Control Plane) 服务器的示例。

官方
精选
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选