timeseries-mcp
Deterministic time-series statistics for AI agents. This MCP server gives any LLM agent unit-tested statistical tools — anomaly detection, changepoint detection, seasonal decomposition, stationarity/trend tests, data-quality audits, baseline forecasts — with schema-validated structured output and no arbitrary code execution.
README
timeseries-mcp
Deterministic time-series statistics for AI agents. An MCP server that gives any LLM agent unit-tested statistical tools — anomaly detection, changepoint detection, seasonal decomposition, stationarity/trend tests, data-quality audits, baseline forecasts — with schema-validated structured output and no arbitrary code execution.
Agent: "Is anything wrong with the server room this week?"
load_csv(server_room_temp.csv) → ts1: 1992 points @ 5min
data_quality(ts1) → 1 sampling gap (2h, ~24 points missed)
detect_anomalies(ts1, stl_residual, → 3 spikes flagged, seasonal-context aware
period=288)
detect_changepoints(ts1_daily) → level shift on day 7: 21.4°C → 23.4°C
forecast_baseline(ts1, seasonal_naive) → next hour ± honest backtest error
Agent: "There's a 2-hour telemetry gap on June 4, three temperature spikes,
and a sustained +2°C shift starting June 7 — likely HVAC degradation.
Baseline forecast error is MAE 2.2°C, so alert thresholds under 3°C
will false-positive."
<picture> <source media="(prefers-color-scheme: dark)" srcset="docs/charts/detections_dark.png"> <img alt="Real detections on the bundled sample data: STL-residual anomalies and a telemetry gap on server-room temperature; CUSUM level shifts bracketing a bad deploy on daily CPU means" src="docs/charts/detections_light.png"> </picture>
Both panels are generated by the library itself — the anomaly markers, gap band, and changepoint segments are real outputs of detect_anomalies, data_quality, and detect_changepoints on the seeded sample datasets (regenerate them).
Why this exists
LLMs are unreliable at arithmetic over long arrays, and the common workaround — handing the model a Python sandbox — is a non-starter in locked-down environments and unauditable everywhere else. The existing "data analysis" MCP servers are mostly run_script shims: the model writes pandas code, executes it server-side, and hopes.
This server takes the opposite position:
- Deterministic — same input, same output, every time. Every number comes from a unit-tested routine (57 tests), not model-generated code.
- No code execution — the tool surface is 17 typed functions. There is nothing to inject into. Safe for enterprise hosts that cannot allow
exec(). - Schema-validated — every tool returns a Pydantic model published as an MCP
outputSchema, so hosts get structured content they can verify, log, and post-process. - Token-frugal by design — data loads once into a server-side registry and gets a handle (
ts1). A million-point series never enters the model's context; every response is capped and previews are evenly thinned.
Tools
| Tool | What it does |
|---|---|
load_csv / load_values / load_sample |
Register a series, get a handle + summary stats back |
list_series / describe / get_window |
Catalog, distribution summary, capped raw windows |
resample / rolling_stats |
Regularize onto a grid; rolling mean/std/min/max/median |
data_quality |
Gaps, duplicate timestamps, missing values, sampling regularity |
detect_anomalies |
zscore, mad (robust), iqr, stl_residual (seasonal-context) |
detect_changepoints |
Level shifts via CUSUM binary segmentation, MAD-robust noise scale |
decompose |
STL / classical split + Hyndman trend/seasonal strength (0–1) |
stationarity |
ADF + KPSS read together, combined verdict + differencing hint |
autocorrelation |
ACF/PACF, significance bounds, seasonal-period suggestion |
trend_test |
OLS + robust Theil-Sen + Mann-Kendall (tie-corrected) |
compare_series |
Pearson/Spearman on shared timestamps + best lead/lag scan |
forecast_baseline |
naive / seasonal-naive / drift / SES, 95% intervals, holdout backtest included |
Plus MCP resources (timeseries://catalog, timeseries://{id}/summary) and a guided analyze_series prompt.
Statistical choices worth noting: anomaly scores are method-honest (MAD falls back with an explanation when 50%+ of values tie); changepoint noise is estimated from first differences so the shifts being hunted don't inflate their own denominator; every forecast ships with a real holdout backtest because a baseline you can't beat is information.
Install
Requires Python 3.11+ and uv.
Claude Code
claude mcp add timeseries -- uvx --from git+https://github.com/Lkhanaajav/timeseries-mcp timeseries-mcp
Claude Desktop / Cursor (claude_desktop_config.json / mcp.json)
{
"mcpServers": {
"timeseries": {
"command": "uvx",
"args": ["--from", "git+https://github.com/Lkhanaajav/timeseries-mcp", "timeseries-mcp"],
"env": { "TIMESERIES_MCP_DATA_ROOT": "/path/to/your/csv/files" }
}
}
}
Streamable HTTP (remote / multi-client)
uvx --from git+https://github.com/Lkhanaajav/timeseries-mcp timeseries-mcp --transport http --port 8000
Try it without an MCP host — the example walkthrough runs the full agent workflow over the in-memory transport, no API key needed:
git clone https://github.com/Lkhanaajav/timeseries-mcp && cd timeseries-mcp
uv sync && uv run python examples/demo.py
Architecture
MCP host (Claude Code / Desktop / Cursor / any client)
│ stdio or Streamable HTTP
▼
FastMCP server — 17 typed tools, 2 resources, 1 prompt
│ series handles (ts1, ts2, ...) — raw data never re-enters context
▼
SeriesStore ── path-sandboxed CSV loader (TIMESERIES_MCP_DATA_ROOT)
│
▼
analysis/ — pure, deterministic, unit-tested routines
anomalies · changepoints · decompose · stationarity
correlation · trend · quality · baselines
(numpy / scipy / statsmodels underneath)
Tool logic is transport-agnostic and per-session state is a single registry object — aligned with where the MCP spec is heading (stateless Streamable HTTP core in the 2026-07-28 revision).
Security posture
- No code execution. No
eval, noexec, no model-written scripts. - Filesystem sandbox.
load_csvresolves paths againstTIMESERIES_MCP_DATA_ROOT(default: the server's working directory) and refuses traversal outside it — tested, including absolute-path escapes. - No network access. The server reads local CSVs and inline arrays only; no URL fetching, no SSRF surface.
- Bounded everything. Row caps on ingestion, point caps on every response, series-count caps on the registry.
- Self-correcting errors. Invalid inputs return actionable tool errors (
Unknown series_id 'ts9'. Known ids: ts1, ts2.) so agents recover instead of hallucinating.
Testing
uv run pytest # 57 tests, ~2s
- Golden statistical tests — injected spikes are found, known slopes are recovered within tolerance, random walks fail stationarity, seasonal-naive beats naive on seasonal data.
- Behavioral contrasts — a value that is globally unremarkable but wrong for its phase of the daily cycle is caught by
stl_residualand correctly not caught by global z-score. - Protocol tests — the full workflow runs over the real MCP transport in memory; every tool is asserted to publish an
outputSchema; error paths surface as MCP tool errors, not crashes.
Honest limitations
- Changepoint detection assumes shifts-plus-noise; on strongly seasonal or trending series,
decomposeorresamplefirst (the sample demo shows this workflow). - Forecasts are reference baselines, deliberately. If your ARIMA can't beat
seasonal_naive's backtest here, it's not adding value. - The series registry is in-process memory: restart = clean slate, and horizontal HTTP scaling would need a shared store (roadmap).
- No multivariate methods yet beyond pairwise comparison.
Related work
mcp-server-data-exploration and pandas-mcp-server take the code-execution route — maximum flexibility, minimum auditability. Vendor servers like InfluxDB MCP front their own databases. This server is the deterministic, self-contained middle: bring a CSV, get defensible statistics.
An agent-facing evaluation suite for this server — scoring whether agents pick the right tools with the right arguments — lives at mcp-trajectory-evals.
Development notes
Built with AI assistance (Claude Code) for scaffolding and test generation; statistical method selection, API design, parameter defaults, and final review are mine. Notable choices I'd defend in review: MAD-of-differences noise estimation for CUSUM (a global σ is inflated by the shifts being detected), reading ADF and KPSS jointly rather than either alone, and refusing to ship forecasts without a holdout backtest.
MIT © Lkhanaajav Mijiddorj
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。