worldmodel-mcp
Provides an external, validated state database for LLM agents to manage long-horizon tasks, with tools for defining schemas, invariants, actions, procedures, and branching, preventing state drift and compounding errors.
README
worldmodel-mcp
An external world model for LLM agents: a validated state database that lives
outside the context window, exposed as an MCP server and a standalone wm CLI.
Why
Apple's The Illusion of Thinking shows reasoning models collapsing on compositional tasks (Tower of Hanoi fails around 8 disks) — even when handed the explicit algorithm. The failure pattern points at execution fidelity, not insight: a transformer cannot reliably maintain exact state across hundreds of steps of its own token stream, and one silent tracking error compounds forever.
This tool changes the model's job. Instead of being the state-keeper, the executor, and the policy all at once, the model keeps only the part it's good at (judgment, decomposition, interpretation) and delegates the rest:
| Paper failure mode | Tool answer |
|---|---|
| State drift over long horizons | State lives in SQLite tables; the model queries only the working set it needs per step (query) |
| Silent, compounding errors | Every mutation is a transaction checked against declared invariants; illegal transitions are rejected with the broken rule named (apply_action, define_invariant) |
| Can't execute known algorithms | Register the algorithm once and let the tool run it — a thousand exact, validated steps for zero tokens (define_procedure, run_procedure) |
| Can't backtrack cleanly | Snapshots and branches; abandoning a branch truly forgets it (snapshot, branch, checkout, restore_snapshot) |
| Work lost to context limits | The world outlives the window; describe_world + history re-orient a fresh context in one call |
What it does not fix: choosing the right next action is still the model. This moves the collapse threshold out substantially; it does not remove it.
Acceptance test
Straight from the paper: 12-disk Tower of Hanoi (unaided models collapse ~8).
Through the world model: schema + 2 invariants + a 12-line recursive procedure →
4095 moves, every one invariant-checked, in 0.26 s, final state verified by
query. See src/core.test.ts ("Tower of Hanoi") for the runnable version.
Documentation
- AGENTS.md — the agent operating manual. If you are an LLM agent (or configuring one), start there: when to reach for a world, the five habits, recipes, the procedure contract, how to write good invariants, gotchas.
- Over MCP the same manual is available without reading files: call the
guidetool (topic:all|quickstart|recipes|procedures), or read theworldmodel://guide/worldmodel://readmeresources. The server also sends a summary as MCP instructions at connection time, so a client model knows what this is before calling anything. wm helpdocuments every CLI command.
Install
npm install
npm run build
npm test
Optionally npm link to get wm and worldmodel-mcp on your PATH.
Register as an MCP server (Claude Code)
claude mcp add worldmodel -- node /path/to/worldmodel-mcp/dist/server.js
Worlds are stored under ~/.worldmodel/ (override with WORLDMODEL_DIR).
Model of operation
A world = one problem domain (a puzzle, an investigation, a migration plan). Each world has:
- Tables — ordinary SQLite tables you define; one fact per row.
- Invariants — named SELECTs that return violating rows (empty = healthy). Checked inside every action's transaction; violations roll the action back.
- Actions — named, logged, transactional mutations (
apply_action, or theassert_facts/retract_factsconveniences). - Procedures — deterministic JS function bodies
(query, act, log, args, lib)run in anode:vmsandbox with step/time budgets; eachact()is a full validated transition, and a pre-run snapshot makes any run onerestorefrom undone. (The sandbox is an isolation convenience, not a security boundary — procedures run with local trust, same as the agent's shell access.) lib.search— generic BFS/DFS over in-memory states:lib.search({start, moves, goal, key?, mode?, maxNodes?})returns{found, path, finalState, nodesExpanded}. The intended pattern is plan in memory, execute validated: search over lightweight state objects (no DB copies per node), then replay the winning path viaact()so every real step is still invariant-checked. See the river-crossing test — the paper's other benchmark family — solved this way insrc/core.test.ts.- Templates —
create_worldwithtemplate: "facts"scaffolds an evidence ledger for investigations:facts(entity, attribute, value, source, confidence, status, superseded_by, ...)plus invariants that force contradictions to be resolved explicitly (supersede or markdisputed) instead of letting incompatible beliefs silently coexist. - Branches & snapshots — cheap file-level copies (
VACUUM INTO) for exploration and undo. - Event log — every schema change, action, and procedure run, queryable
via
history.
CLI quickstart
wm create hanoi "Tower of Hanoi"
wm schema hanoi "CREATE TABLE disks (disk INTEGER PRIMARY KEY, peg TEXT NOT NULL, height INTEGER NOT NULL, UNIQUE (peg, height))"
wm invariant hanoi no-disk-on-smaller "No disk on a smaller disk" \
"SELECT u.disk, l.disk FROM disks u JOIN disks l ON u.peg = l.peg AND u.height = l.height + 1 WHERE u.disk > l.disk"
wm assert hanoi disks '[{"disk":1,"peg":"A","height":3},{"disk":2,"peg":"A","height":2},{"disk":3,"peg":"A","height":1}]'
wm act hanoi move "UPDATE disks SET peg=@dst, height=COALESCE((SELECT MAX(height) FROM disks WHERE peg=@dst),0)+1 WHERE disk=(SELECT disk FROM disks WHERE peg=@src ORDER BY height DESC LIMIT 1)" '{"src":"A","dst":"B"}'
wm query hanoi "SELECT * FROM disks ORDER BY peg, height"
wm check hanoi
wm help lists everything; long SQL/code args accept - for stdin;
wm serve runs the MCP server.
MCP tools
list_worlds · create_world · describe_world · define_schema ·
define_invariant · drop_invariant · apply_action · assert_facts ·
retract_facts · query · check_invariants · snapshot ·
restore_snapshot · branch · checkout · history ·
define_procedure · run_procedure
Agent usage discipline
The tool only helps if state is routed through it. The intended habit, for any task with evolving state plus rules:
create_world,define_schema, and encode the domain rules as invariants first — rules in the database, not in your head.- Mutate only via actions; treat a rejection as information, not failure.
- Query the working set per step; never re-derive state from the transcript.
- The moment you know the algorithm, write a procedure instead of narrating steps.
- Snapshot before anything exploratory; branch to compare alternatives.
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。