Clonst
Enables adversarial code review by connecting Claude Code to OpenAI Codex, iterating until consensus is reached between the two AI models.
README
Clonst - AI code review for Claude Code, powered by Codex
Get a second AI opinion on your code before it ships. Clonst is a Model Context Protocol (MCP) server that connects Claude Code to OpenAI Codex for adversarial code review: Claude writes the plan or the code, Codex critiques it, Claude revises, and the loop repeats until both models reach consensus. It runs on your existing ChatGPT subscription through the official Codex CLI. No API keys, no extra billing.
Independent project. Not affiliated with OpenAI ("Codex", "ChatGPT") or Anthropic ("Claude").
Why a second model?
An LLM reviewing its own output shares its own blind spots. A second model, from a different provider, with its own training and its own memory, catches what the first one missed: wrong assumptions, missing edge cases, fragile migrations, race conditions, security holes. Clonst turns that into a structured review loop with a hard exit criterion: consensus, not politeness.
Features
- Zero ritual. Install it and forget it: Claude triggers the review by itself, and only when the change is worth your quota.
- Ping-pong until consensus. Unlimited rounds by default. Structured verdicts (APPROVED / CHANGES_NEEDED), required changes, suggestions, risks. Claude can reject a critique with justification; Codex re-evaluates the rejection next round.
- Real session memory. Codex resumes the same CLI session on every round and remembers its earlier critiques. No context re-sending, no goldfish reviewer.
- Your ChatGPT subscription, no API keys. Reviews go through the official
codexCLI and its existing login. - Pick your reviewer. By default, reviews use the model and reasoning
effort your codex CLI is already configured with. One line in
~/.clonst/config.jsongives reviews their own setting (e.g. a faster, cheaper effort) without touching your Codex extension - see Configuration. - Cost transparency. The final report shows the reviewer model, rounds, total duration and tokens consumed.
- Your project's own review rules. Drop a
CLONST.mdat the project root and the reviewer checks your conventions on top of its own standards. CLAUDE.md guides the writer; CLONST.md guides the reviewer. - Reviews in your language. Critiques come back in the language you work
in (per call, or once for all with
default_languagein the config), while the protocol stays machine-readable English. - Nothing is ever lost. Every raw reviewer response is saved to disk before any parsing.
- Cross-platform. Windows, macOS and Linux, all three covered by the CI.
- Hardened. Prompt-injection guards, read-only sandbox, whitelisted CLI arguments, hermetic test suite (no LLM call, no quota).
When does it trigger?
You never have to ask. Claude decides when a review is worth your quota, and the rule is stakes, not size:
| Reviews by itself | Stays silent |
|---|---|
| Business logic, computations | Pure presentation (HTML/CSS, copy) |
| Data flows, models, migrations | Documentation, comments |
| Routes, APIs, integrations | Renames without behavior change |
| State, error handling, concurrency | Local config tweaks |
| Security, authentication | Throwaway scripts and prototypes |
| Plans and architecture, before coding | (when in doubt, it asks you) |
The ping-pong is unlimited by default: it runs until consensus, checking in with you every 5 rounds (configurable). And you can cap any review in plain language: "review this, 3 rounds max".
How it works
You ── conversation ── Claude (reviser, keeps the conversation context)
│
│ clonst_review (MCP, one call = one critique)
▼
Clonst ── spawn ── codex exec [resume <thread_id>]
(reviewer, keeps its session)
The loop lives on Claude's side: it submits, Codex critiques, Claude revises
in the conversation (in front of you), resubmits with the returned
thread_id, until consensus: true. The loop is driven through the caller (the
thread_id travels with each call), while the server keeps per-session
records on disk: full logs, the previous round's verdict for exact recall,
and the running duration/token totals. Codex sessions persist on the CLI side.
Quick start
Requirements:
- Claude Code - the terminal CLI or the VS Code / JetBrains extensions, which share the same MCP configuration
- Node.js 22+
- The Codex CLI, logged in with a ChatGPT plan:
npm install -g @openai/codex
codex login
Install Clonst (recommended, via npm):
claude mcp add clonst --scope user -- npx -y @clonst/clonst
Or from source:
git clone https://github.com/capritora/clonst.git && cd clonst
npm install && npm run build
claude mcp add clonst --scope user -- node /absolute/path/to/clonst/dist/index.js
Check it works: in a new Claude Code conversation, say "ping clonst".
Expected: codex_available: true, codex_logged_in: true.
What a review looks like
Reviews happen on their own, but you can also steer them:
Propose a plan for X, then have it reviewed by Clonst until consensus.
Have this migration reviewed by Clonst, 3 rounds max.
Ask Clonst for a security-focused review of this auth flow.
You see each revision happen in the conversation, and at consensus Claude ends with a short report, for example:
Clonst review: GPT-5.5 (high effort), 2 rounds, 5 min 30 s, ~500k tokens. The reviewer required an anti-double-correction bound on the migration and a timeout on the API call, both applied. I rejected one suggestion (out of MVP scope) and the reviewer accepted the justification.
Tools
clonst_ping
Server health: codex CLI availability, version, login status, loaded config, logs directory. Consumes no quota.
clonst_review
One structured critique per call. Parameters (all drive the calling LLM; you normally never write these yourself):
| Parameter | Default | Role |
|---|---|---|
content (required) |
- | The plan/code to review, complete (later rounds: the full revised version, never a diff) |
context |
none | Round 1: goal, constraints, decisions already made |
project_path |
none | ABSOLUTE project path: Codex runs there and reads the real files (read-only sandbox). See Privacy below |
thread_id |
none | Later rounds: the identifier returned by the previous call (resumes the Codex session) |
round |
1 (2 with thread_id) | Round number; hard safety cap at 50 |
max_rounds |
unlimited | Hard round limit for this review; at the limit, disagreement goes to the user |
language |
language of the content | Code like "fr" or "pt-BR": the reviewer writes critiques in that language. Resolved server-side; the raw value never reaches the prompt |
review_focus |
all | bugs, architecture, performance, security, or all |
changes_made / changes_rejected |
none | Later rounds: what was changed / rejected with justification |
Result: verdict, consensus (true only on a proven APPROVED: clean JSON,
zero required changes, no fallback parsing), critique, required_changes,
suggestions, risks_identified, thread_id, per-round and whole-review
duration and token usage, reviewer_model / reviewer_reasoning_effort
(best-effort resolution: override, else codex config, else null), and a
next_action instruction (text + typed fields) that drives the loop.
Configuration
Optional file, absent by default: create ~/.clonst/config.json yourself to
change any key. It is re-read on every call; invalid values fall back to the
default with a warning.
| Key | Default | What it does |
|---|---|---|
codex_model |
null = inherit ~/.codex/config.toml |
Model used for reviews only (e.g. "gpt-5.5"). Your Codex VS Code extension keeps its own setting |
codex_reasoning_effort |
null = inherit |
Reasoning effort for reviews only (e.g. "medium", "high", "xhigh"). Lower = faster, cheaper rounds |
default_language |
null = language of the reviewed content |
Language of the critiques when the caller does not pass one, as a code like "fr" or "pt-BR" |
suggested_max_rounds |
5 |
Without an explicit limit, Claude checks in with you every N rounds (5, 10, 15...) before continuing. NOT a limit |
timeout_per_call_seconds |
600 |
Timeout of one Codex call (reasoning models take minutes) |
Example, reviews at high effort while your extension stays at xhigh:
{
"codex_reasoning_effort": "high"
}
Overrides are passed as root -c flags to the codex CLI (contract verified on
codex 0.142.5 for both exec and resume); a value unknown to codex fails the
review with codex's own error message.
Project review guidelines: CLONST.md
Put a CLONST.md at the project root and, whenever a review runs with
project_path, its content is handed to the reviewer as project-specific
guidelines - your conventions, your red lines, checked on top of the
reviewer's own standards. Example:
# Review guidelines
- SQL must stay compatible with BOTH SQLite (dev) and PostgreSQL (prod).
- Every new route needs rate limiting.
- LLM results must be matched by ID, never by list position.
Guidelines can only ADD checks: a guideline trying to lower the bar or force a verdict is ignored and reported as a risk.
Round limits: unlimited by default
Say nothing and the ping-pong continues until consensus, with a check-in every
suggested_max_rounds rounds. Or ask in natural language ("review this,
3 rounds max"): the reviewer is told about the counter (exhaustive from
round 1, maximum effort on the final round, never approving just to close),
and at the limit the disagreement goes to you for arbitration.
Privacy and quota
- Every review round consumes your ChatGPT subscription quota (Codex does
the reviewing). When the quota window is exhausted, Clonst detects it and
tells Claude to continue without review; the session stays resumable later
through the same
thread_id. - With
project_path, Codex reads the whole project read-only (including.env) and that content goes to OpenAI - the same exposure as using the Codex VS Code extension directly. Default strategy: review the content passed incontentonly; reserveproject_pathfor reviews that must verify real APIs, contracts or files.
Troubleshooting
| Symptom | Cause | Action |
|---|---|---|
kind: "cli_not_found" |
codex CLI missing from PATH | npm install -g @openai/codex |
kind: "exec_failed" + login hint |
ChatGPT session expired | codex login |
kind: "timeout" |
Review too long | Raise timeout_per_call_seconds in ~/.clonst/config.json |
kind: "exec_failed" + quota hint |
ChatGPT usage limit reached (rolling window) | Continue without review; relaunch when the window resets |
codex_available: false on ping |
CLI missing or broken | codex --version in a terminal |
| Long reviews fail while short ones pass | The MCP CLIENT's timeout (not Clonst's) | Launch Claude Code with MCP_TOOL_TIMEOUT=600000 |
| Odd behavior after changing the source | dist/ is gitignored |
npm run build |
Each ping-pong is fully logged under ~/.clonst/logs/<thread_id>.jsonl, with
complete raw responses in ~/.clonst/logs/raw/<thread_id>/.
Development
npm test # build + hermetic test suite (no LLM calls)
npm run smoke # full MCP protocol smoke test
The scripts/probe-*.ps1 scripts pin the real codex CLI contract and consume
ChatGPT quota: manual execution only. The ReviewerProvider interface is ready
for other reviewer CLIs (e.g. Gemini).
License
MIT. Provided as is, without support guarantees. Used daily by its author.
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。
mcp-server-qdrant
这个仓库展示了如何为向量搜索引擎 Qdrant 创建一个 MCP (Managed Control Plane) 服务器的示例。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。