malaria-forecast-mcp
MCP server enabling AI agents to access provincial malaria surveillance and outbreak forecasts for Angola, with guardrails to ensure safe and validated outputs.
README
malaria-forecast-mcp
An MCP server that gives an AI agent access to provincial malaria surveillance and short-horizon outbreak forecasting for Angola — with the guardrails that make model output safe for an agent to act on.
Built by Joaquim Timóteo. The forecasting work it wraps is described in Operational Malaria Forecasting in Angola Using Ensemble Models, Regional Clusters, and Epidemiological Memory Features (ResearchGate, Feb 2026).
Why MCP instead of a REST API
A REST endpoint gives a model a URL and hopes the prompt explains the rest. MCP ships the contract alongside the capability, and three consequences follow that matter for anything forecasting-shaped:
Discovery is dynamic. Tool schemas are read at connect time. Adding compare_provinces made it available to every connected client without a single prompt being rewritten.
Provenance travels with the capability. malaria://model-card is a resource the model can read before quoting a number — validation method, measured skill, known failure modes. With a REST API that context lives in a PDF somewhere, which is to say it does not reach the model at all.
Refusals are structured. Ask for a 20-week horizon and you get a typed error naming the validated range, not a plausible-looking wrong number:
{
"error": "horizon_out_of_range",
"detail": "horizon_weeks must be between 1 and 8; got 20. The model was validated only to 8 weeks and will not extrapolate beyond it.",
"max_validated_horizon_weeks": 8
}
That last one is the whole argument. A forecasting model wired to an agent without guardrails will answer any question it is asked, including the ones it has no business answering.
What it exposes
Tools
| Tool | Purpose |
|---|---|
list_provinces |
All 18 provinces with epidemiological stratum (K-means burden clustering) |
get_incidence_history |
Weekly incidence and the rainfall driver, filtered by date range |
forecast_incidence |
1–8 week forecast with empirical 80% intervals |
detect_outbreak_signals |
Weeks running above the same-calendar-week seasonal baseline |
compare_provinces |
Ranked forecast across provinces, for resource prioritisation |
Resources
malaria://model-card— architecture, validation method, measured metrics, limitations, guardrailsmalaria://provinces— province directory for grounding
Prompts
outbreak_briefing— walks the agent through model card → history → signals → forecast, then writes a briefing that always states intervals rather than point estimatescompare_and_prioritise— ranks provinces and requires the agent to say when two are not meaningfully separable
Guardrails
- Horizons outside 1–8 weeks are refused, with the reason, rather than extrapolated.
- Provinces with under 52 weeks of history are refused rather than forecast on a season the model has never seen.
- Every point carries an empirical 80% interval from rolling-origin residuals — no distributional assumption.
- Anomaly flags are seasonal. A flag means "high for this week of the year" against prior years, not "high in absolute terms" — which in a seasonal disease is the difference between a signal and a calendar.
Evaluation
The harness was written before the tools, and it earns its place: it caught a real defect.
python evals/backtest.py
Rolling-origin backtest, 26 origins per province per horizon — 468 scored forecasts at each horizon:
h origins MAE baseline skill cov80
--------------------------------------------------
1 468 0.5677 0.8666 0.3450 79.70%
2 468 0.5992 0.8666 0.3085 80.13%
3 468 0.6137 0.8666 0.2919 80.77%
4 468 0.6100 0.8666 0.2961 82.69%
5 468 0.6051 0.8666 0.3018 82.69%
6 468 0.6299 0.8666 0.2732 85.26%
7 468 0.6461 0.8666 0.2544 86.11%
8 468 0.6538 0.8666 0.2456 86.11%
skill is 1 − (model MAE / seasonal-naive MAE). The script exits non-zero if any horizon stops beating the baseline, so this is a gate rather than a report.
Two findings worth stating plainly, because they are the reason the harness exists:
Fixed ensemble weights lost to the baseline at 7–8 weeks. Local trend and climate signal decay with range while seasonal structure survives. Weights are now horizon-dependent, and skill is positive across the full range. Intuition said the ensemble was fine; the backtest said otherwise.
The intervals were miscalibrated. The textbook 0.80 quantile of absolute residuals produced 90–95% measured coverage — too wide, because residuals estimated on recent origins are systematically harder than the weeks being forecast. The quantile was calibrated down to 0.60, which measures at ~80% at short horizons and stays conservative (~86%) at long ones. Coverage is reported on every run so it cannot drift silently.
Data
The bundled dataset is synthetic. Provincial surveillance records are not redistributable, so the series reproduces the statistical shape of the real thing — rainy-season seasonality, burden strata, interannual variability, outbreak excursions — without exposing restricted data.
Every metric in this README describes this reimplementation on synthetic data. The published research model reports R² 0.985, MAE 6.9 per 1,000 and an 87.5% skill score on real surveillance across all 18 provinces, 2000–2024. Those are different numbers about a different artefact and the model card keeps them clearly separated.
To run against real data, implement the SurveillanceStore interface in data.py. No tool signature changes.
Install and run
git clone https://github.com/joaquimtimoteo/malaria-forecast-mcp
cd malaria-forecast-mcp
pip install -e .
python -m malaria_forecast_mcp # stdio server
python scripts/smoke_check.py # 26 end-to-end protocol checks
python evals/backtest.py # evaluation gate
pytest tests/ # full suite
Claude Desktop
Add to claude_desktop_config.json:
{
"mcpServers": {
"malaria-forecast": {
"command": "python",
"args": ["-m", "malaria_forecast_mcp"],
"env": { "PYTHONPATH": "/absolute/path/to/malaria-forecast-mcp/src" }
}
}
}
Then ask: "Which three provinces should we prioritise six weeks out, and how confident are you?" — the agent reads the model card, ranks provinces, checks each against seasonal baselines, and reports intervals rather than point estimates.
Layout
src/malaria_forecast_mcp/
server.py MCP tools, resources, prompts
forecasting.py ensemble, intervals, guardrails
data.py surveillance store + synthetic generator
model_card.py machine-readable provenance
evals/backtest.py rolling-origin evaluation gate
scripts/smoke_check.py
tests/
Roadmap
- RAG over published epidemiological literature, so briefings cite evidence
- Real-data adapter for DHIS2 surveillance exports
- Intervention-effect handling (bed-net campaigns, IRS rounds)
Licence
MIT
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。