WorkloadTruth MCP Server

WorkloadTruth MCP Server

Enables classification of GPU workloads as training, inference, or idle from telemetry data, with tools for one-shot classification, benchmarking, and audit log verification.

Category
访问服务器

README

WorkloadTruth

Classify a GPU workload as TRAINING, INFERENCE, or IDLE from telemetry alone. No code changes to the workload, no self-reported job labels.

CI PyPI License: Apache 2.0 Python 3.9+

WorkloadTruth classifying a synthetic training workload, then running the evasion-robustness benchmark

Every GPU scheduler in common use today, including run:ai, Slurm, and Kubernetes GPU operators, asks you to declare whether a job is training or inference at submission time. None of them check. WorkloadTruth reads GPU telemetry (utilization, memory pattern, power draw) and answers the question independently, so a mislabeled or misbehaving job doesn't go unnoticed.

Table of contents

Quick summary

  • Install: pip install "workloadtruth-cli[nvml]" for real GPU access, or pip install workloadtruth-cli to try it with the synthetic backend, no GPU needed
  • Use it for: catching cost-misallocated GPU jobs (a job billed as low-priority "inference" that's actually running full training) and unauthorized workload changes (an inference endpoint that starts training on live traffic without sign-off)
  • What it's not: a compliance or regulatory-audit tool. No regulation currently requires this kind of monitoring, see What WorkloadTruth is not below
  • Prior art: builds on and cites arXiv:2606.19262 (ICML 2026), see Relationship to prior research

Install

# Real NVIDIA GPU telemetry (requires an NVIDIA driver on the host)
pip install "workloadtruth-cli[nvml]"

# Try it without a GPU, using the synthetic backend
pip install workloadtruth-cli

# npm launcher (thin wrapper around the PyPI package, see "Why two registries")
npx workloadtruth-cli --help

Quickstart

# No GPU required. Classify a synthetic "training" telemetry trace.
$ workloadtruth classify --backend synthetic --profile training --samples 10 --interval 0
workload_type : TRAINING
confidence    : 1.00
gpu_index     : 0
samples       : 10 over 9.0s
reasons:
  - avg GPU utilization 87.8% >= training threshold 65.0%
  - low GPU utilization variance (std=3.4) <= training ceiling 15.0
  - memory growing 120.0 MiB/sample >= training threshold 5.0
  - low power-draw variance (std=9.4W) <= training ceiling 25.0W

# Real hardware
$ workloadtruth classify --backend nvml --samples 10 --interval 1 --json

--json on every command switches to machine-readable output for scripts and agents.

How classification works

WorkloadTruth ships a rule-based classifier only in v0.1: a set of documented, inspectable thresholds over four features extracted from a telemetry window (average and variance of GPU utilization, memory-growth slope, average and variance of power draw). Every threshold lives as a named constant in src/workloadtruth/classifier/rules.py with a comment explaining its intuition. Nothing is a black box.

--experimental (an ML-based classifier) is present as a flag but fails loudly with an explanation rather than shipping a fake result. arXiv:2606.19262's trained model and dataset were never published, and this project has no NVIDIA GPU in its build environment to collect real training data. An ML classifier ships here only once it's trained on a real, disclosed dataset and independently shown to beat the rule-based baseline, not before.

Benchmark

workloadtruth benchmark measures the classifier's accuracy against documented synthetic GPU telemetry traces (see synthetic_backend.py for the exact profile parameters), both clean and under a deliberate evasion transform that mimics an operator disguising a training job as inference (capping utilization swings, flattening memory growth, injecting artificial idle gaps).

This is run on synthetic data, not live NVIDIA hardware, so it is not directly comparable to arXiv:2606.19262's real-hardware numbers. Both are reported below, side by side, never blended into one figure.

arXiv:2606.19262 (real hardware, NVML) WorkloadTruth v0.1 (synthetic traces, 300 trials/cell)
Clean accuracy 98.2% 100.0%
Evasion accuracy 43-87% 66.7% overall

The overall evasion number hides the real finding. Broken out by workload type, WorkloadTruth's rule-based classifier is currently completely fooled on evasive training workloads (0% accuracy), while staying accurate on evasive inference and idle traces (100% each). Capping GPU utilization alone is enough to make a training job read as inference to this classifier today.

Profile Clean accuracy Evasion accuracy
Training 100.0% 0.0%
Inference 100.0% 100.0%
Idle 100.0% 100.0%

This is the concrete, disclosed gap the benchmark suite exists to surface and track, not a footnote. A rule-based classifier that only looks at utilization variance is trivially evadable. Closing this gap, through multi-signal fusion or a real trained classifier once data exists, is the roadmap, not a solved problem. Reproduce it yourself:

workloadtruth benchmark --trials 300 --window 30 --json

CLI reference

$ workloadtruth --help
Usage: workloadtruth [OPTIONS] COMMAND [ARGS]...

  Classify a GPU workload as INFERENCE, TRAINING, or IDLE from telemetry
  alone.

Options:
  --version  Show the version and exit.
  --help     Show this message and exit.

Commands:
  benchmark   Run the evasion-robustness benchmark against synthetic...
  classify    One-shot classification of the current GPU workload.
  mcp         Start an MCP server exposing classify/benchmark/verify-log...
  verify-log  Verify the hash chain of a local audit log has not been...
  watch       Continuously classify and append hash-chained entries to...
Command Purpose
classify One-shot classification. --backend synthetic|nvml, --json for machine-readable output.
watch Continuous classification; appends a hash-chained entry to a local audit log on every window.
benchmark Runs the evasion-robustness benchmark (see above).
verify-log Re-derives every audit-log entry's hash and confirms the chain hasn't been tampered with.
mcp Starts an MCP server (stdio) exposing classify_workload, run_benchmark, verify_audit_log as agent-callable tools. Requires pip install "workloadtruth-cli[mcp]" on Python 3.10+ (see below).

Every command supports --json. Full flag reference: workloadtruth <command> --help.

Agent-native (MCP + A2A)

pip install "workloadtruth-cli[mcp]"
workloadtruth mcp

The mcp package itself requires Python 3.10+, stricter than WorkloadTruth's own 3.9 floor. Every other feature (classify, watch, benchmark, verify-log) works on Python 3.9.

Exposes three tools over stdio MCP: classify_workload, run_benchmark, verify_audit_log. A .well-known/agent.json manifest is shipped at the repo root for A2A-style discovery, listing both the CLI and MCP interfaces and the packages that provide them.

Audit log

workloadtruth watch appends a hash-chained entry to workloadtruth.log.jsonl on every classification window. Each entry's hash covers its own content plus the previous entry's hash, so any edit, reorder, or deletion after the fact breaks the chain from that point forward. workloadtruth verify-log re-derives every hash and reports the first broken link, if any.

This proves what was classified, when, and that the local record hasn't been silently altered afterward. It does not prove the classification itself was correct, and it is not evidence of regulatory compliance. See below.

Why two registries

WorkloadTruth's implementation is Python. NVML access (pynvml/nvidia-ml-py) is the mature, official way to read NVIDIA GPU telemetry, and it's also what the closest prior art (arXiv:2606.19262) uses. The npm package (workloadtruth-cli) is a thin launcher, not a reimplementation. It locates and execs the real workloadtruth binary installed from PyPI, so npx workloadtruth-cli works for npm-first agent tooling without duplicating the classifier in two languages.

Comparison

WorkloadTruth NVIDIA DCGM / dcgm-exporter run:ai Weights & Biases
Reads GPU telemetry Yes (via NVML) Yes (source) Yes Yes
Classifies workload type automatically Yes No, exposes raw metrics only No, workload type is user-declared at job submission No, training-run-scoped by design, no classification
Local, hash-chained audit trail Yes No No No
Evasion-robustness benchmark Yes (documented, reproducible) N/A N/A N/A
Requires an NVIDIA GPU Only for the nvml backend; the synthetic backend works without one Yes Yes No (general system metrics)

Checked directly against each project's own documentation: DCGM exporter docs, run:ai inference overview, W&B system metrics docs. None of these classify workload type from telemetry alone. That gap is what WorkloadTruth fills.

What is WorkloadTruth, and why does it exist

WorkloadTruth is an open-source command-line tool and MCP server that classifies a running GPU workload as TRAINING, INFERENCE, or IDLE using only GPU-level telemetry (utilization, memory pattern, power draw), with no changes to the workload's own code and no reliance on a self-reported job label.

It exists because every mainstream GPU scheduler asks the job's owner to declare its type at submission time and never checks that declaration against what the hardware is actually doing. That gap has two real consequences: cost misallocation (a job scheduled at low-priority "inference" pricing that is actually running full training) and unauthorized workload changes (an inference endpoint that quietly starts training on live traffic). WorkloadTruth closes that verification gap today, and doubles as the first open, installable implementation of a real academic research thread on verifying AI training runs from hardware telemetry (see below).

Relationship to prior research

WorkloadTruth's core technique, classifying training vs. non-training GPU activity from telemetry, is not novel. It's the direct application of a real, active research thread:

  1. Yonadav Shavit (Harvard), "What does it take to catch a Chinchilla?" (2023): proposed hardware-level "training transcripts" for verifying large training runs.
  2. GovAI, "Computing Power and the Governance of AI" (2024): surveyed compute-governance mechanisms, explicitly framed as exploratory, not endorsed policy.
  3. "Hardware-Enabled Mechanisms for Verifying Responsible AI Development" (2025): hardware-security researchers proposing on-chip attestation.
  4. Rahman & Tajdari, "Detecting Hidden ML Training With Zero-Overhead Telemetry" (ICML 2026 Technical AI Governance workshop): a working NVML-telemetry classifier, 98.2% accurate on unobfuscated workloads, the direct prior art for this project's core classification technique.

What WorkloadTruth adds: as of this project's own research (2026-07-19), no open-source, installable implementation of this research thread existed, only academic prototypes. WorkloadTruth is that packaging: a real CLI, an MCP server, a hash-chained audit log, and a reproducible evasion-robustness benchmark, built in the open. It does not claim to improve on the paper's classification technique. See the benchmark section above: the current rule-based classifier is considerably more evadable than the paper's ML approach on the one axis it measures.

What WorkloadTruth is not

  • Not a compliance or regulatory-audit tool. No law currently requires inference/training classification or reporting. Any future claim otherwise will name the specific enacted regulation; none exists as of this writing.
  • Not a content inspector. WorkloadTruth reads GPU-level signals only (utilization, memory, power). It never inspects model weights, training data, prompts, or completions.
  • Not proof of "compliant" or "safe" operation. The audit log proves what was classified and when, and that the record wasn't altered afterward, not that the classification was correct or that any policy was followed.

FAQ

Does this need an NVIDIA GPU? Only for the nvml backend. --backend synthetic runs the full classifier and CLI against documented synthetic traces, no GPU required. Useful for trying the tool or for CI.

Can it classify AMD or Intel GPU workloads? Not in v0.1. The telemetry layer is a pluggable interface (TelemetryBackend) specifically so a new vendor backend (AMD ROCm, Intel Level Zero) can be added without touching the classifier. See CONTRIBUTING.md.

Is the classifier accurate enough to bill or penalize someone based on its output? Not yet, and the benchmark section above is the honest reason why: 0% accuracy on evasive training workloads today. Treat workload_type as a signal to investigate, not a verdict.

Why not just use the ML classifier from the paper? Its trained weights and dataset were never published. Reimplementing an ML classifier without real training data would produce an unvalidated accuracy claim, not a measured one. See How classification works.

How is this different from run:ai or NVIDIA DCGM? run:ai and DCGM both expose or use GPU telemetry, but neither classifies workload type from that telemetry. run:ai relies entirely on the label the job's owner declares at submission; DCGM just exposes raw utilization and memory metrics for something else to interpret. WorkloadTruth is the layer that actually looks at the telemetry and answers the question. See the comparison table.

Does this work on Windows, macOS, and Linux? The synthetic backend runs anywhere Python 3.9+ runs, including this project's own macOS build environment (which has no NVIDIA GPU). The nvml backend requires an NVIDIA GPU and driver, which in practice means Linux or Windows with NVIDIA hardware; NVML itself is not available on macOS.

What license is this under, and can I use it commercially? Apache 2.0. Commercial use, modification, and redistribution are all permitted under its terms; see LICENSE.

Contributing

See CONTRIBUTING.md. Security issues: see SECURITY.md.

License

Apache 2.0

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选