agent-telemetry-mcp

agent-telemetry-mcp

MCP server that reports AI agent costs (tokens, latency, dollars) with per-tenant isolation via OAuth, and includes an ADK agent for natural-language queries.

Category
访问服务器

README

agent-telemetry-mcp

An MCP server that reports what your AI agents cost: tokens, latency, dollars per request. Each caller sees only their own rows, because the customer id comes from their OAuth token rather than from a tool argument.

There is also an ADK agent that asks it questions in English.

MCP Python SDK 2.0 (protocol revision 2026-07-28), ADK 2.6, Auth0, BigQuery, Cloud Run.

Try it without installing anything

It is deployed. An unauthenticated call tells you where to authenticate:

$ curl -si https://agent-telemetry-x62fjiecda-ew.a.run.app/mcp \
    -H 'Content-Type: application/json' \
    -H 'Accept: application/json, text/event-stream' \
    -d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}'

HTTP/2 401
www-authenticate: Bearer error="invalid_token", error_description="Authentication required",
  resource_metadata="https://agent-telemetry-x62fjiecda-ew.a.run.app/.well-known/oauth-protected-resource/mcp"

That last URL is the RFC 9728 document, and it is public — it names the authorization server and the scopes this resource accepts.

What it looks like

two tenants, one binary

Two logins against the same deployment. demo@acme.com sees about $0.29; demo@globex.com sees about $0.49 — a bigger bill from a third as many calls, because that tenant is seeded pro-heavy. The totals creep up as you use it, since CostPlugin records each run. Then the model is taken out of the loop and the tool is called directly with a tenant named in the arguments:

{'days': 7}                       -> tenant=globex  cost=$0.4916
{'days': 7, 'tenant': 'acme'}     -> tenant=globex  cost=$0.4916
{'days': 7, 'tenant_id': 'acme'}  -> tenant=globex  cost=$0.4916

Both accounts are on a throwaway Auth0 dev tenant, and both can write: demo@acme.com / AcmeDemo-20b9c40a64-Xq7, demo@globex.com / GlobexDemo-6d9b1808cf-Kt3.

Why a server and not just ADK's BigQueryToolset

BigQueryToolset accepts an end user's token:

credentials = Credentials(token=user_token)
BigQueryToolset(credentials_config=BigQueryCredentialsConfig(credentials=credentials))

That is fine when you trust the agent. It is not enough when one agent serves many customers, because nothing checks the token — the agent uses whatever it was handed.

So the server does the checking. It validates issuer, audience and scopes, reads the customer id from the validated token, and binds that id as a query parameter. ADK's own docs say the same about their nearest equivalent:

job_labels: Note: These labels are for usage discovery and tracking purposes only and should not be used for security-sensitive decisions.

How it fits together

telemetry_analyst              picks which sub-agent answers
├── explorer_agent    ──→  BigQueryToolset   public datasets, agent's identity
└── telemetry_agent   ──→  this MCP server   your own data, your identity

CostPlugin  ──→  this MCP server  ──→  BigQuery table

CostPlugin runs on the ADK Runner, so it sees every model call from every agent in the tree. After each call it reads the token counts, prices them, and sends a row through the same authenticated connection the reads use. That is where the data in the table comes from — the system measures itself.

explorer_agent is four lines of BigQueryToolset. Read-only mode and the byte ceiling are already BigQueryToolConfig fields, so there was nothing to write.

The one rule

No tool takes a tenant argument. Not a validated one — there is no such parameter:

usage_summary          days
cost_by_model          days
slowest_invocations    days, limit
record_invocation      agent, model, prompt_tokens, output_tokens, ...

record_invocation needs telemetry:write, the reads need telemetry:read, and a token that lacks a scope does not see the tool it would have unlocked. Asking anyway returns 403 with an RFC 6750 scope hint, which the client uses to step up in one round trip. tests/test_tenancy.py fails the build if a tenant argument ever appears.

Run it

uv venv && uv pip install -e ".[dev,agent]"
cp .env.example .env          # AUTH0_DOMAIN, MCP_CANONICAL_URI, GOOGLE_CLOUD_PROJECT
                              # plus GOOGLE_API_KEY for the model

uvicorn agent_telemetry.server:create_app --factory --reload
python -m agent.main

deploy/deploy.sh provisions and deploys; the Auth0 steps that have no API are in deploy/README.md. deploy/seed.py fills a demo tenant, priced from agent/pricing.py so the demo cannot drift from the code it is demonstrating.

Tests

pytest        # 43 tests, no credentials needed
file what it covers
test_tenancy.py two tokens get two different slices; no tenant claim is refused
test_scope_gate.py per-tool scopes, through the real ASGI stack
test_agent_auth.py the client half: loopback callback, cached tokens, discovery
test_cost_plugin.py pricing, and not double-counting streamed responses
test_wiring.py the plugin and the server still agree on field names

Things that were not obvious

The tests all passed while the deployed server could not answer a single authenticated request. An in-memory Client(server) hands the server an object graph — no HTTP, no Host header, no ASGI receive channel. Four separate faults lived in that gap. The one I would not have guessed: streamable_http_app() defaults to a localhost-only Host allowlist, so every Cloud Run request came back 421, and only after authentication had succeeded.

scopes_supported and required_scopes are different things. The SDK builds the RFC 9728 document from required_scopes, which is the floor a token must already clear, not what a client may ask for. So the document advertised telemetry:read alone, and a client that follows the spec's scope-selection order was told telemetry:write did not exist.

The client discovers metadata on the 401 path only. A process that starts with a cached token never sees a 401, so a later 403 step-up re-authorized with no metadata in hand and fell back to {resource_server}/authorize — a URL this server does not serve. The browser opened on a 404 and the flow waited forever.

ADK's McpToolset does not work with MCP SDK 2.x. It imports mcp.shared.session, which 2.0 removed, and the package swallows the ImportError and logs it at DEBUG. What you see is cannot import name 'McpToolset'. That is why agent/toolset.py exists.

Not included

Approval prompts, the Tasks extension, MCP Apps, A2A, ADK evals. All doable, none of them make the access control any better.

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选