nimbus-ops

nimbus-ops

MCP server for Nimbus Operations Copilot that enables AI agents to search help docs, customer data, invoices, and tickets, and propose refunds/tickets with human approval for writes.

Category
访问服务器

README

Nimbus Operations Copilot

An AI agent that resolves customer-operations requests end to end — it gathers facts across the help center, customer database, billing, and ticketing, proposes the policy-correct action, and (after a human approves) takes it. Every step is traced.

Project 3 (capstone) of a 12-week Forward Deployed Engineer portfolio. Headline skills: multi-step agents + MCP (Model Context Protocol) with production guardrails — human-in-the-loop, observability, and evals.

Read the problem framing first: CASE_STUDY.md.

In plain English

Imagine a customer-support rep at a software company. To answer one question — "was this customer charged twice, and can we refund it?" — they normally have to dig through four different systems: the help articles, the customer list, the billing records, and the support-ticket tool. It's slow and easy to get wrong.

This project is an AI assistant that does that digging for them. You ask a question in plain English, and it:

  1. Looks things up on its own — it searches the help docs, finds the customer, and checks their invoices, one step at a time, just like a person would.
  2. Explains what it found, quoting the real records and the company's policy (it never makes facts up).
  3. Asks permission before changing anything. If it wants to issue a refund or open a ticket, it stops and waits — a human clicks Approve or Deny first. It literally cannot change data on its own.
  4. Keeps a record of everything it did, so you can always check its work.

Think of it as a careful new assistant: great at gathering information and suggesting the right action, but it always checks with you before doing anything that can't be undone. That "always ask before acting" safety is the most important idea in the project — and there's a test suite proving the assistant never acted without approval.

(The rest of this README is the technical detail behind that.)

Why this is the capstone

Projects 1 and 2 answered questions (from docs, then from data). This one takes action — a real agent that plans across systems, exposed through an MCP server (the industry-standard way to connect an agent to a client's tools), with the safety and observability a client would actually require.

How it works

The agent plans across tools to gather facts, then pauses for human approval before any write — reads run freely, writes are gated:

flowchart TD
    A["👤 Ops rep asks in plain English"] --> B["🖥️ React console"]
    B -->|"stream of steps (SSE)"| C["🤖 Agent · Claude"]

    subgraph MCP["🔌 MCP server — Nimbus systems"]
        direction TB
        READ["read tools<br/>docs · customers · invoices · tickets · usage"]
        WRITE["write tools<br/>issue_refund · create_ticket"]
    end

    C -->|"plan and call read tools"| READ
    READ --> DB[("🗄️ Postgres + help docs")]
    READ -->|"facts"| C

    C --> G{"wants to write?"}
    G -->|"🔒 needs approval"| H["👤 Rep approves or denies"]
    H -->|"approve"| WRITE
    H -->|"deny"| C
    WRITE --> DB

    C -->|"answer + steps"| B
    C -.->|"every LLM and tool call"| L["🔭 Langfuse trace"]

    classDef agent fill:#052e2b,stroke:#10b981,color:#d1fae5;
    classDef gate fill:#3a2a06,stroke:#f59e0b,color:#fde68a;
    class C agent
    class G gate
    class H gate

Stack

  • Agent: Python + Claude API (multi-step tool-use loop)
  • Integration layer: FastMCP — an MCP server wrapping Nimbus's "systems"
  • Data: PostgreSQL (nimbus db) — customers, invoices, tickets, usage, help docs (RAG)
  • Safety: human-in-the-loop approval for any write (refund, ticket)
  • Observability: Langfuse tracing of every LLM + tool call
  • API/UI: FastAPI (SSE streaming) + React ops console

Setup

# 1. Ops database
createdb nimbus
psql nimbus -f db/schema.sql
psql nimbus -f db/seed.sql

# 2. Python env
python3.11 -m venv .venv && .venv/bin/pip install -r requirements.txt

# 3. Configure
cp .env.example .env   # add your ANTHROPIC_API_KEY

Evals

10 realistic ops tasks (billing, account, docs, writes, and a safety trap) scored on two axes — full report in evals/REPORT.md:

  • Correctness 10/10 — right tools called, right action proposed, facts grounded in tool results.
  • Safety: 2 writes proposed · 0 executed without approval — the suite runs with every write denied, verifying the agent only ever gates a write, never runs one autonomously.
.venv/bin/python -m evals.evaluate   # writes evals/REPORT.md

Agent runs are non-deterministic (correctness varies slightly run to run); the safety property holds every run by construction — a write can only execute after human approval.

Deploy

The Dockerfile builds one container that serves the React console and the agent API from FastAPI (the MCP server runs in-process). Target: GCP Cloud Run + Cloud SQL — step-by-step commands in DEPLOY.md.

Try the tools in Claude Desktop (optional)

The MCP server runs standalone, so you can plug it into Claude Desktop and watch Claude call the Nimbus tools directly — the "aha" of MCP.

  1. Open ~/Library/Application Support/Claude/claude_desktop_config.json (create it if missing) and paste the contents of claude_desktop_config.example.json.
  2. Restart Claude Desktop. You'll see nimbus-ops tools appear (the 🔌 icon).
  3. Ask: "Was Acme Corp double-charged in July? If so, what does the refund policy say?" — Claude will call get_customer, list_invoices, and search_docs on its own.

Milestones

  • [x] 1. Client brief + skeleton + seeded ops database
  • [x] 2. MCP tool server — 6 read tools (docs/DB) + 2 write tools (refund/ticket), read/write-annotated
  • [x] 3. Agent loop — Claude plans across the MCP read tools multi-step; streams steps (SSE)
  • [x] 4. Human-in-the-loop — agent pauses on writes; POST /api/approve gates each refund/ticket
  • [x] 5. Observability — Langfuse traces every LLM + tool call (nested, with token cost)
  • [x] 6. Ops console (React) — watch the agent's steps stream in, approve/deny writes inline, link to the trace
  • [x] 7. Evals — 10 ops tasks scored for correctness (10/10) and safety (0 writes executed without approval)
  • [x] 8. Ship — Dockerized (one container), GCP deploy guide (DEPLOY.md), full case study

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选