The Box
A local MCP daemon that turns an agentic coding client into a bug bounty operator, with 103 tools for offensive security testing including MITM proxy, traffic analysis, and OOB callbacks.
README
<p align="center"> <img src="./assets/the-box-chan.png" alt="the-box" width="260"> </p>
<h1 align="center">The Box</h1>
<p align="center"> <b>A local MCP daemon that turns an agentic coding client into a bug bounty operator.</b> </p>
<p align="center"> <img src="https://img.shields.io/badge/tools-103%20MCP-black" alt="103 MCP tools"> <img src="https://img.shields.io/badge/runtime-Node%2022-black" alt="Node 22"> <img src="https://img.shields.io/badge/platform-WSL2%20%2F%20Linux-black" alt="WSL2 / Linux"> <img src="https://img.shields.io/badge/license-MIT-black" alt="MIT"> </p>
The Box is not a chat wrapper around nmap. It is a stateful daemon that owns a
MITM proxy, a traffic database, an application graph, an evidence vault, an
out-of-band callback server and a local CVE/report knowledge base — and exposes
all of it as 103 MCP tools. You point Claude Code (or any MCP client) at it,
and the model drives real offensive tooling against targets you authorized,
while the daemon enforces scope, rate limits and reporting discipline in code.
The design premise: an LLM is good at deciding what to try next and terrible at not making things up. So every claim the model wants to make has to survive a gate that the daemon controls, not the model.
┌─────────────────┐ MCP over StreamableHTTP ┌──────────────────────────┐
│ Claude Code │ ──────────────────────────> │ boxbridge daemon │
│ (orchestrator) │ <────────────────────────── │ 127.0.0.1:7070/mcp │
└─────────────────┘ 103 native tools └───────────┬──────────────┘
│
┌──────────────────────┬───────────────────────┬───────┴───────────┬─────────────────┐
▼ ▼ ▼ ▼ ▼
MITM proxy SQLite traffic + OOB server Intel RAG Sandboxed
(CA, HTTP/1+2, app graph + (HTTP + DNS, (CVE, ExploitDB, tool runner
TLS impersonate) evidence vault ACME certs) H1 corpus) (rate-limited)
Table of contents
- What it actually does
- Feature tour
- Install
- Running it
- Connecting a client
- Configuration
- Tool reference
- Project layout
- Development
- Status and limitations
- Scope of use
- License
What it actually does
Three things a plain MCP tool wrapper does not:
1. It keeps state across the whole hunt. Every request that goes through the
proxy, the browser or a replay lands in traffic.db. From that traffic the
daemon builds an application graph (endpoints, parameters, edges, staleness),
tracks which of 18 attack techniques were already tried against each endpoint,
and lets the model ask "what have I not tested yet?" instead of re-scanning
blindly.
2. It enforces the rules in code, not in the prompt. A scope guard checks every offensive request against the authorized target list before it leaves the machine. A rate governor caps requests per second per target. A report gate lints drafts against banned hedging phrases and hard-stops any finding that cannot complete the sentence "As an intruder I can ___ by ___, demonstrated by ___". The model cannot talk its way past these — they run in the tool call path.
3. It closes the loop on evidence. Findings attach to a hash-chained evidence vault (Merkle-verifiable), so a report references artifacts that can be proven unmodified since capture.
Feature tour
Recon
Passive and active discovery, each wrapped as a native tool with parsed structured output rather than raw stdout.
- Subdomains —
subfinder_scan,amass_passive(CT logs, DNS providers, archives) - Historical URLs —
waybackurls_discovery,gau_discovery,urlfinder_discovery - Crawling —
hakrawler_crawl - Ports —
nmap_scan(with-sV),naabu_scan,rustscan_scan,masscan_scan - HTTP probing —
httpx_probe(status, title, server, tech fingerprint) - DNS —
dnsenum_scan(records + zone transfer attempt) - Content discovery —
ffuf_discovery,feroxbuster_discovery,gobuster_discovery,dirsearch_discovery, all with soft-404 baselining - Parameter discovery —
arjun_discovery,x8_discovery,paramspider_discovery,param_extract,qsreplace_transform - Fingerprinting —
whatweb_fingerprint,wafw00f_detect,identify_stack(aggregates both) - OSINT —
censys_host_lookup,censys_webproperty_lookup
Intercepting proxy
A full MITM proxy written from scratch, not a shell-out to Burp.
- Own root CA, per-host leaf generation with a cache
- HTTP/1.1 and HTTP/2 servers, WebSocket passthrough
- TLS impersonation — egress through
curl-impersonateso the JA4 fingerprint matches a real Chrome build instead of Node's TLS stack - Everything captured into
traffic.db, searchable viahistory_search - Configurable passthrough host list, client-certificate identities
Replay and fuzzing
repeater_send— replay any captured request with method/path/header/body overridesintruder— positional payload attacks over a captured request, with draft/run/poll verbschain_send— a sequence of requests in one call, feeding each response into the next (CSRF token flows, multi-step state machines)comparer_diff— side-by-side diff of two captured artifactsreflected_find— scan captured traffic for a canary, classified by injection context (html_body, html_attribute, js_string, …)probe_matrix— fire the full 18-technique catalog at one parameter in one calldifferential_send— replay one request as several registered identities and emit a mechanical authorization verdict
Authorization testing
Testing authz properly needs more than one session, so identities are first-class:
identity_register/identity_list— victim, attacker, unauth, low_priv, admin- Credentials live in an encrypted vault (argon2 + passphrase), never in tool output
identity_capture_denial— record how the target should deny a request, so a later bypass is provabledifferential_sendcompares responses across identities mechanically
Business logic
- Invariants —
invariant_register,invariant_check,invariant_mine. Register a property that must always hold ("balance never goes negative", "status only moves forward"), then have the daemon check it against live traffic. Mining proposes candidates from captured history. - Workflows — record a multi-step flow from the browser (
workflow_record_start/workflow_record_stop, sliced into steps by page-load events), or reconstruct one retroactively from captured traffic (workflow_from_history). Replay it pure (byte-literal, as a regression baseline) or bound (substituting tokens and cookies from earlier steps) viaworkflow_replay. - State machine abuse —
workflow_mutateruns the recorded flow with a mutation plan: skip a step, repeat one 2–50 times, reorder the permutation, or fire several in parallel.workflow_mutation_analysisclassifies what came back into hypotheses. - Race conditions —
workflow_racefires a single-packet race on a chosen step to detect TOCTOU and deduplication bypasses - Semantic parameter fuzzing —
workflow_semantic_fuzzmutates a parameter using a 50-kind catalog driven by the parameter's inferred meaning, not a generic wordlist
Browser
- Chromium via CDP, launched by the daemon or attached to an existing instance
- Routed through the MITM proxy, so browser traffic joins the same history
- Fingerprint parity work (navigator, WebGL, fonts) to survive bot detection
- Evidence capture: screenshots, console, network, DOM snapshots
Out-of-band
A self-hosted Collaborator equivalent:
oob_issue_tokenmints a callback FQDN;oob_check_hitsretrieves callbacks- HTTP and DNS listeners, ACME-provisioned certificates for HTTPS callbacks
- A public panel to inspect hits
- Detects blind SSRF, blind RCE, blind XXE, blind SQLi
Static analysis
- JS bundles —
js_analyze: fetch (8 MB cap) → prettier beautify → source-map recovery → AST + regex scanners for endpoints, secrets and dangerous sinks.js_subagent_minespawns a focused subagent to find what regex misses (constructed strings, base64 blobs, env-gated code). - Mobile —
apk_extract(apktool + jadx),ipa_extract(Mach-O encryption info, plists).mobile_endpoint_mineandmobile_subagent_minepull endpoints and findings out of the artifacts. Optional Frida hooking on a connected device (frida_inventory,frida_trace_class). - GraphQL —
graphql_introspect,graphql_take_snapshot(sha256-deduped for temporal diff),graphql_authz_fuzz(synthesizes minimal mutations with placeholder args to probe authorization without executing destructive writes)
Scanners
Wrapped with parsed output and scope enforcement: nuclei_scan (inline custom
YAML templates supported), sqlmap_scan, dalfox_scan, xsstrike_scan,
wpscan_scan.
Intel RAG
A local knowledge base, hybrid BM25 + vector retrieval with a reranker:
- CVE data (NVD), ExploitDB, nuclei templates, CWE catalog
- A corpus of disclosed HackerOne reports
- ONNX embeddings (bge-m3) served locally, LanceDB vector store, optional GPU
cve_search— keyword search over CVE, ExploitDB and nuclei templates with vendor/product/year/severity filtersh1_similar_reports— find disclosed reports resembling the endpoint being triagedpattern_match_by_endpoint— cross-source union of historical findings and CWE patterns matching an endpoint signature
Queried mid-hunt, this answers "what has already worked against this exact stack" instead of guessing.
Hunt management
- Targets —
create_targetregisters in DRAFT (passive only) until the operator promotes it to active - Hypotheses — a ranked queue of leads;
hypothesis_create_suggestionmakes the model file a lead instead of asserting a vuln - Coverage ledger —
coverage_reportanswers "what has not been tried", per endpoint and technique - Verdicts —
endpoint_verdict_emitrecords defended / inconclusive / blocked / vulnerable per endpoint - Blockers —
blocker_declareprices a blocker (2FA, CAPTCHA, missing credential) instead of silently giving up;blocker_resolvereopens everything that was closed because of it - Notebook — per-target notes with typed mentions
- Diff watcher — scheduled snapshots of subdomains, URLs, graph nodes and bundle findings, alerting on change
- ROI tracker — imports HackerOne CSV to track what actually paid
Reporting discipline
This is the part the project cares most about.
report_create_from_hypothesisrenders from a fixed templatereport_lintruns the gate: banned hedging phrases (likely,could,almost certainly,platform-wide compromise), missing sections, unproven chain links- Intruder test — a hard stop. A finding that cannot complete "As an intruder I can ___ by ___, demonstrated by ___" does not render.
program_policy_setrecords what the program actually pays for, so out-of-scope classes get dropped before effort is spent- Adversarial panel —
subagent_spawnwith refuter lenses;adversarial_verdict_emitcollects votes trying to refute a finding before it ships session_close_reportgenerates an honest end-of-session summary from the coverage ledgers, including what was left untested
Safety substrate
- Scope guard — offensive requests are checked against the active target's authorized scope; DRAFT targets allow passive tools only
- Rate governor — per-target request/second and concurrency caps
- Sandboxed runner — external binaries run with a per-plugin working directory, hard timeouts and structured JSONL logging
- Evidence vault — hash-chained,
evidence_vault_verify_chainreturns a Merkle root or the exact sequence number where the chain broke
Install
Target platform is WSL2 Ubuntu 24.04. It runs on native Ubuntu too; the Windows host steps below are WSL-specific.
0. Host prerequisites (WSL only, PowerShell as admin)
wsl --status
wsl --install -d Ubuntu-24.04 # if absent
wsl --set-default-version 2
Edit %USERPROFILE%\.wslconfig — the daemon holds embedding models in memory
and Chromium plus parallel nuclei runs are not cheap:
[wsl2]
memory=16GB
processors=8
swap=8GB
localhostForwarding=true
Then wsl --shutdown and reopen the Ubuntu terminal.
1. Clone into the Linux filesystem
Not /mnt/c/ — the I/O penalty is severe and SQLite behaves badly across the
9p mount.
git clone https://github.com/horizonfps/the-box-mcp.git ~/the-box
cd ~/the-box
2. Bootstrap the toolchain
chmod +x infra/bootstrap.sh infra/verify.sh
./infra/bootstrap.sh
./infra/verify.sh
bootstrap.sh is idempotent — a second run is mostly skip: lines. It
installs:
| Source | Contents |
|---|---|
| apt | build tooling, nmap, masscan, dnsenum, whatweb, keyring/secret libs, Chromium runtime deps, JRE (for apktool) |
| NodeSource | Node 22 |
| Go (pinned 1.23.4, sha256-verified) | subfinder, httpx, nuclei, naabu, katana, urlfinder, amass, waybackurls, gau, hakrawler, ffuf, gobuster, qsreplace, anew, dalfox — all version-pinned |
| cargo | rustscan 2.4.1, feroxbuster 2.13.1, x8 4.3.2 |
| pipx | arjun, paramspider, wafw00f, dirsearch, uro, frida-tools |
| direct download | sqlmap, nuclei templates, wordlists |
Layout it creates:
/opt/thebox/ system-wide
├── bin/ symlinks to installed tools
├── cache/ wordlists, nuclei templates, intel corpora
└── proxy/ proxy support files
~/.thebox/ per-user
├── state/ app.db, traffic.db, sessions, bridge token
├── logs/ rotated logs (bridge, tools, intel)
├── config/ runtime config
└── scratch/ per-target working directories
verify.sh is the gate: exit 0 means the environment is good.
3. Install Node dependencies
corepack enable
pnpm install
Native modules (better-sqlite3, argon2, onnxruntime-node,
@lancedb/lancedb) compile here. This is the step most likely to need build
tooling that step 2 already installed.
4. Optional components
Each is independent — skip what you do not need.
./infra/proxy-ca-install.sh # trust the MITM root CA locally
./infra/proxy-impersonate-install.sh # curl-impersonate for Chrome JA4 parity
./infra/browser-install.sh # Chromium for the CDP browser module
./infra/oob-public-listen-install.sh # public OOB listeners (needs a domain + open ports)
./infra/oob-public-cert-sync.sh # ACME certificates for HTTPS callbacks
sudo ./infra/systemd/install-templates-timer.sh # nightly nuclei template refresh
The intel RAG needs a one-time corpus build (downloads and embeds; expect it to take a while and to want a GPU if you have one):
node engine/scripts/intel-ingest.mjs
node engine/scripts/intel-embed-setup.mjs
node engine/scripts/intel-embed.mjs
Running it
cd ~/the-box
pnpm bridge
The daemon binds 127.0.0.1:7070, writes its bearer token to
~/.thebox/state/bridge-token and its port to ~/.thebox/state/bridge-port. A
PID lockfile prevents a second instance from corrupting the SQLite databases.
For a long-lived session detached from the terminal:
setsid nohup pnpm bridge > ~/.thebox/bridge.log 2>&1 < /dev/null & disown
Production build (bundles to a single ESM file, ~50 KB plus node_modules):
pnpm build && pnpm start
Inference credentials
Some tools run their own inference inside the daemon — subagent_spawn,
js_subagent_mine, mobile_subagent_mine, the adversarial review panel. Those
need an Anthropic-compatible endpoint:
export ANTHROPIC_BASE_URL="https://api.anthropic.com"
export ANTHROPIC_AUTH_TOKEN="<your token>"
Any Anthropic-compatible proxy works in place of the official endpoint. Tools that do not spawn subagents run fine without these set.
Connecting a client
Add the daemon to your MCP client. For Claude Code, in .mcp.json:
{
"mcpServers": {
"thebox": {
"type": "http",
"url": "http://127.0.0.1:7070/mcp",
"headers": { "Authorization": "Bearer ${THEBOX_BRIDGE_TOKEN}" }
}
}
}
export THEBOX_BRIDGE_TOKEN="$(cat ~/.thebox/state/bridge-token)"
Two meta-tools bracket every session:
box_targets— list authorized targets and show which is activebox_use_target— set the active target; the scope guard and the per-target rate limit key off this
Nothing offensive runs until a target is active. A typical opening:
box_use_target { "target": "example.com" }
coverage_report {} → what has not been tried
hypothesis_list {} → the ranked lead queue
Restarting the daemon drops the MCP connection: reconnect in the client and
call box_use_target again, since the active target resets.
Configuration
Everything is env-driven. Defaults are sane; the ones worth knowing:
| Variable | Default | Purpose |
|---|---|---|
THEBOX_BRIDGE_HOST |
127.0.0.1 |
Bind address. Leave it local. |
THEBOX_BRIDGE_PORT |
7070 |
Daemon port |
THEBOX_STATE_DIR |
~/.thebox/state |
Databases, token, sessions |
THEBOX_LOGS_DIR |
~/.thebox/logs |
Bridge, tool and intel logs |
THEBOX_SCRATCH_DIR |
~/.thebox/scratch |
Per-target working directories |
THEBOX_PROXY_PASSTHROUGH_HOSTS |
— | Hosts the MITM proxy must not intercept |
THEBOX_PROXY_IMPERSONATE_BIN |
— | Path to curl-impersonate for TLS parity |
THEBOX_BROWSER_MODE / THEBOX_BROWSER_CDP_URL |
— | Launch a browser or attach to one |
THEBOX_OOB_PUBLIC_DNS_BIND |
— | Bind address for the public DNS listener |
THEBOX_OOB_PUBLIC_TRUSTED_PROXY |
— | Upstream proxy allowed to set forwarded headers |
THEBOX_INTEL_EMBED_GPU |
0 |
GPU acceleration for embeddings |
THEBOX_INTEL_LANCE_DIR |
under state dir | Vector store location |
THEBOX_ADVERSARIAL_REVIEW |
on | Adversarial refuter panel before a finding ships |
THEBOX_TOOL_LOOP_MAX_ITER |
0 (unlimited) |
Cap on the internal tool loop |
NVD_API_KEY |
— | Higher NVD rate limit during intel ingest |
WPSCAN_API_TOKEN |
— | WPScan vulnerability database |
Full list: grep -rhoE 'THEBOX_[A-Z0-9_]+' engine/src | sort -u.
Tool reference
103 tools total — 101 native plugins plus box_targets and box_use_target.
Each plugin lives in engine/src/tools/native/<name>/index.ts with a JSON
Schema validated by ajv at registration and at every call.
Adding one is three steps: create the directory with an index.ts default-
exporting a ToolPlugin, add one import and one array entry in
engine/src/tools/native/loader.ts, restart. The loader uses static imports
deliberately — plugins must share the boxbridge/persistence/db.ts singleton
with the bootstrap, which separate bundler entry points would duplicate. See
engine/src/tools/native/README.md.
The daemon prints the full roster on boot:
[boxbridge] loaded 101 tools (adversarial_verdict_emit, amass_passive, apk_extract, ...),
manifestTokens=5075 fullSchemaTokens=26269
Descriptions are what steer the model, so they carry the intent and the preconditions, not just the signature.
Project layout
engine/src/
├── boxbridge/ daemon: HTTP + MCP server, auth, sessions, persistence, routes
├── tools/
│ ├── native/ 101 tool plugins, one directory each
│ ├── runner.ts sandboxed external-binary runner (rate limit, timeout, JSONL log)
│ └── ...
├── proxy/ MITM: CA, leaf certs, HTTP/1, HTTP/2, WS, TLS impersonation
├── history/ traffic capture, search, replay
├── appGraph/ endpoint/parameter graph built from traffic
├── intruder/ positional payload attacks
├── repeater/ request replay with overrides
├── comparer/ artifact diffing
├── reflected/ canary reflection finder with context classification
├── probe/ 18-technique probe matrix
├── chain/ multi-step request chains
├── bizlogic/ race conditions, state machines, semantic fuzzing
├── invariants/ invariant registration, checking, mining
├── identities/ test identities and credential vault
├── browser/ CDP control, stealth, evidence capture
├── oob/ out-of-band HTTP + DNS server, ACME, public panel
├── jsBundle/ JS bundle download, beautify, source-map recovery, scanners
├── mobile/ APK/IPA extraction and analysis, Frida integration
├── graphql/ introspection, snapshots, authz fuzzing
├── intel/ RAG: ingest, chunk, embed, BM25 + vector retrieval, rerank
├── evidence/ capture and hash-chained vault
├── hypotheses/ lead queue, ranking, lints
├── coverage/ technique coverage ledger
├── reports/ rendering, lint gate, intruder test, export
├── programPolicy/ what the program pays for; opening gate
├── adversarial/ refuter panel and verdicts
├── subagents/ focused subagent spawning and finding emission
├── diffWatcher/ scheduled surface snapshots and change alerts
├── notebook/ per-target notes with typed mentions
├── roi/ bounty ROI tracking, HackerOne CSV import
├── ops/ scope guard, rate governor, blockers, audit log
└── prompts/ system prompt modules and rules
infra/ bootstrap, verification, optional installers, systemd units
engine/scripts/ smoke tests, intel pipeline, maintenance scripts
Development
pnpm test # full suite: 4251 tests, 697 suites, ~5 min
pnpm test:orphans # find test files the runner does not pick up
pnpm build # tsup bundle → engine/dist/bootstrap.js
node engine/scripts/smoke-<area>.mjs # per-area smoke tests
Smoke scripts exist per subsystem (smoke-browser, smoke-oob-public,
smoke-evidence-vault, smoke-graphql-deep, smoke-mobile, …) and run against
the real daemon rather than mocks.
pnpm typecheck runs tsc in strict mode over the daemon surface and currently
reports around 70 errors — mostly test doubles that do not satisfy full
playwright interfaces, unused locals, and a handful of nullability gaps. They
do not affect the build (tsup transpiles without type checking) or the test
suite, which is green. It is standing debt, not a gate.
Comments and identifiers are English. Some inline comments and prompt content are still Portuguese from earlier iterations.
Status and limitations
Honest accounting, since the project's whole premise is not overclaiming:
- Memory. The daemon starts around 4.5 GB RSS because the embedding model loads twice — a known debt. Each intel query adds roughly 1.17 GB, which saturates rather than leaks. Budget accordingly, or skip the intel corpus.
- Restart cost. Restarting drops MCP clients and resets the active target.
- Single instance. A PID lockfile enforces one daemon per state directory.
- Platform. Developed and tested on WSL2 Ubuntu 24.04. Other Linux distributions should work; macOS and Windows-native are untested.
- Optional tools soft-fail. If a binary is missing the tool returns
{ok: false, error: ...}instead of crashing the daemon.masscanneeds root orCAP_NET_RAW. Frida needs a connected device. - A prior desktop UI is not part of this repository. The Box began as a Tauri application with its own chat interface; it was rebuilt as an MCP daemon so that an existing agentic client does the orchestration. Some module names and comments still carry traces of that era.
Operational notes
Things worth knowing before you run this on your own machine.
- Keep the daemon local. It binds
127.0.0.1for a reason. The bearer token in~/.thebox/state/bridge-tokenis the only thing between a caller and the full toolchain. - The proxy has a root CA. Its private key sits under
~/.thebox/state/. Installing that CA into a trust store makes the machine trust anything the proxy signs — do it on your hunting box, not on a shared one. traffic.dbis not yours. It holds full request and response bodies from targets, authenticated sessions included. It does not belong in a repository, in a backup you do not control, or pasted into a report.- Credentials are encrypted at rest (argon2 + your passphrase) and tools return a reference, never the secret.
bash_execruns arbitrary commands as the daemon user. The sandbox bounds the working directory and the timeout, not the capability.
Scope of use
This is offensive tooling. It sends real traffic to real hosts.
Use it only against systems you own or are explicitly authorized to test — a bug bounty program whose published scope covers the target, a signed engagement, or your own lab. The scope guard, the DRAFT-by-default target state and the rate governor keep an authorized engagement inside its boundaries; they are not a substitute for the authorization itself.
You are responsible for what you point this at.
License
MIT. See LICENSE.
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。