graph-mcp-java-gen
Enables users to turn natural-language requests into validated, compilable Java test methods, grounded in a graph catalog to prevent hallucinated imports and ungrounded symbols.
README
graph-mcp-java-gen
Graph-grounded MCP server that converts natural-language requests into validated, compilable Java test methods — no hallucinated imports, no ungrounded symbols, no silent failures.
A natural-language or structured request enters an official Model Context Protocol (MCP) stdio server. A versioned graph catalog (Neo4j or JSON fixture) supplies the only symbols the generator may cite. A multi-layer validator checks syntax, framework contract, grounding, and forbidden-API rules before any source is returned. Two optional LLM agents — an intent normaliser and a post-generation reviewer — extend the pipeline to free-form input without compromising the deterministic safety envelope.
Architecture
%%{init: {"theme": "base", "themeVariables": {"primaryColor": "#1e3a5f", "primaryTextColor": "#ffffff", "primaryBorderColor": "#0d2137", "lineColor": "#4a9eca", "secondaryColor": "#2d6a4f", "tertiaryColor": "#7b2d8b"}}}%%
flowchart TD
classDef input fill:#7b2d8b,stroke:#4a1a54,color:#fff,font-size:13px
classDef mcp fill:#e07b00,stroke:#9a5700,color:#fff,font-size:13px
classDef agent fill:#1a6b8a,stroke:#0d3f52,color:#fff,font-size:13px
classDef core fill:#2d6a4f,stroke:#1b4332,color:#fff,font-size:13px
classDef graph fill:#1e3a5f,stroke:#0d1f3c,color:#fff,font-size:13px
classDef validate fill:#4a6741,stroke:#2d4026,color:#fff,font-size:13px
classDef ok fill:#155724,stroke:#0a3015,color:#fff,font-size:13px
classDef reject fill:#721c24,stroke:#3d0a0e,color:#fff,font-size:13px
NL["🌎 Free-form NL\n(generate_java_test_nlp)"]:::input
SF["📄 Structured fields\n(generate_java_test)"]:::input
TX["💬 Intent text\n(generate_java_test_from_intent)"]:::input
MCP["🔌 FastMCP stdio Server\n7 tools · zero raw Cypher"]:::mcp
A1["🤖 LLMIntentParser\nAgent 1 · gpt-4o-mini\nfield extraction"]:::agent
INT["✅ GenerationIntent\nclass · package · module\nconfig · version"]:::core
GDB["📊 Graph Catalog\nNeo4j 5.26 / JSON fixture\n8 symbols · 12 methods"]:::graph
GEN["⚙️ Template Generator\ndeterministic render"]:::core
VAL["🛡️ JavaValidator\nTree-sitter AST\ncontract · grounding\nsource-safety"]:::validate
A2["🤖 ReviewAgent\nAgent 2 · gpt-4o-mini\n6-item checklist"]:::agent
OK["✅ Accepted Java\nsource + citations\n+ review verdict"]:::ok
REJ["❌ Typed Rejection\nerror code + message\nno source returned"]:::reject
NL --> MCP
SF --> MCP
TX --> MCP
MCP -->|"NLP path"| A1
MCP -->|"direct path"| INT
A1 -->|"extracted fields"| INT
INT -->|"invalid"| REJ
INT -->|"valid"| GDB
GDB -->|"cited symbols"| GEN
GEN --> VAL
VAL -->|"any gate fails"| REJ
VAL -->|"all gates pass"| A2
A2 -->|"issues found"| REJ
A2 -->|"approved"| OK
Multi-Agent Pipeline
%%{init: {"theme": "base", "themeVariables": {"actorBkg": "#1e3a5f", "actorTextColor": "#ffffff", "actorBorderColor": "#4a9eca", "activationBkgColor": "#2d6a4f", "activationBorderColor": "#155724", "noteBkgColor": "#fff8e1", "noteTextColor": "#333", "signalColor": "#4a9eca", "signalTextColor": "#1e3a5f"}}}%%
sequenceDiagram
autonumber
actor User
participant MCP as FastMCP Server
participant A1 as LLMIntentParser<br/>(Agent 1)
participant GDB as Graph Catalog<br/>(Neo4j / Fixture)
participant GEN as Generator +<br/>JavaValidator
participant A2 as ReviewAgent<br/>(Agent 2)
User->>MCP: generate_java_test_nlp(free-form NL)
MCP->>A1: extract intent fields
Note over A1: gpt-4o-mini · temp=0<br/>strict JSON schema
A1-->>MCP: {class, package, module, config, version}
MCP->>GDB: get versioned symbols
GDB-->>MCP: 7 cited GraphSymbol objects
MCP->>GEN: render Java + validate
Note over GEN: Tree-sitter AST<br/>contract · grounding · safety
GEN-->>MCP: validated Java source
MCP->>A2: review(source, class, package)
Note over A2: gpt-4o-mini · temp=0<br/>6-item checklist
A2-->>MCP: {approved, checklist, issues}
MCP-->>User: {status, source, citations, review}
Evidence Dashboard
All measurements use independently generated CC0-licensed synthetic fixtures.
Results are from the accepted strict_graph_v2 policy on the held-out confirmation split.
| Surface | Result | Artifact |
|---|---|---|
| Benchmark scale | 96 CC0 intents — 32 dev / 32 val / 32 confirmation | task_evaluation.json |
| Confirmation task success | 32 / 32 bounded tasks | evaluation_trace.json |
| Generated-source validation | 24 / 24 supported intents — syntax + contract + grounding + safety | task_evaluation.json |
| Safe adversarial rejection | 8 / 8 — zero false accepts | task_evaluation.json |
| Citation precision | 100% — only graph-cited symbols imported | task_evaluation.json |
| Required-symbol recall | 100% — every required symbol present | task_evaluation.json |
| Live Neo4j integration | Neo4j 5.26.29 — 8 symbols, 12 methods materialized | neo4j_integration.json |
| Official MCP benchmark | 120 / 120 expected outcomes — zero protocol errors | mcp_benchmark.json |
| MCP warm latency (p50 / p95 / p99) | 29.13 / 48.61 / 54.23 ms at concurrency 1 | mcp_benchmark.json |
| Java compilation | 8 / 8 class files via Eclipse ECJ 3.21 | java_compile.json |
| External model calls (deterministic path) | 0 calls · $0.00 | mcp_benchmark.json |
Latency figures are single-process local Windows measurements, not production SLOs.
Policy Selection
Four generation policies were evaluated. The selection objective was declared before opening the confirmation split: maximise validation task success among candidates passing all safety gates. Confirmation was opened exactly once for the selected candidate.
%%{init: {"theme": "base", "themeVariables": {"quadrant1Fill": "#155724", "quadrant2Fill": "#856404", "quadrant3Fill": "#721c24", "quadrant4Fill": "#856404"}}}%%
xychart-beta
title "Validation: task success vs safe-rejection recall (%)"
x-axis ["no_graph_v0", "lenient_repair_v1", "strict_graph_v2 ✓", "wide_context_v3"]
y-axis "Task success (%)" 0 --> 105
bar [21.9, 75.0, 100.0, 96.9]
line [87.5, 0.0, 100.0, 87.5]
| Candidate | Task success | Gen valid | Safe reject | Citation prec | Decision |
|---|---|---|---|---|---|
no_graph_v0 |
21.9% | 0% | 87.5% | 0% | Rejected — no grounding |
lenient_repair_v1 |
75.0% | 100% | 0% | 100% | Rejected — 8 false accepts |
strict_graph_v2 |
100% | 100% | 100% | 100% | Selected |
wide_context_v3 |
96.9% | 100% | 87.5% | 87.5% | Rejected — irrelevant context + 1 false accept |
MCP Tools
| Tool | Type | Behaviour |
|---|---|---|
get_fixture_metadata |
Read | Returns fixture identity, provenance, license, backend, symbol count |
search_graph |
Read | Parameterised name/method search; max 20 results |
generate_java_test |
Generate | Typed fields → graph lookup → Java → all validation gates |
generate_java_test_from_intent |
Generate | Bounded 3-form grammar → same strict policy |
validate_java_source |
Validate | Checks up to 20 000 chars; never writes or executes source |
generate_java_test_nlp |
Multi-agent | LLM intent parser → generator → LLM reviewer; requires OPENAI_API_KEY |
The Neo4j adapter uses fixed parameterised Cypher, rejects credentials in URIs, and refuses fixture identity collisions.
Quick Start
python -m venv .venv
# Windows
.\.venv\Scripts\Activate.ps1
# Linux / macOS
source .venv/bin/activate
pip install -r requirements-dev.txt
pip install --no-deps -e .
# Run the offline smoke test (no database needed)
python scripts/container_smoke.py python -m graph_mcp.server
MCP client configuration (VS Code / Claude Desktop)
{
"mcpServers": {
"graph-java-gen": {
"command": "/absolute/path/to/.venv/bin/python",
"args": ["-m", "graph_mcp.server"],
"cwd": "/absolute/path/to/repo"
}
}
}
Enable the multi-agent NLP tool
# Add to your environment or .env file
OPENAI_API_KEY=sk-...
GRAPH_BACKEND=neo4j # optional; defaults to local JSON fixture
Reproduce Evidence
# Build the CC0 benchmark fixture
python scripts/build_evaluation_fixture.py
# Run all four candidate policies and select strict_graph_v2
python scripts/evaluate_workflow.py
# Validate the claims ledger and evidence privacy rules
python scripts/validate_evidence.py
# Full test suite
pytest --cov=src --cov-report=term-missing --cov-fail-under=75
# Lint and security
ruff check src tests scripts
bandit -r src scripts -q -ll
pip-audit -r requirements.txt --progress-spinner off
Live Neo4j path
# Start a local Neo4j Community instance (Docker)
docker compose up -d neo4j
python scripts/wait_for_neo4j.py
# Seed the synthetic graph fixture and verify retrieval
python scripts/seed_graph.py
python scripts/verify_neo4j.py # writes evidence/neo4j_integration.json
# Full MCP benchmark over stdio with live graph
python scripts/benchmark_mcp.py # writes evidence/mcp_benchmark.json
Java compilation
# Requires JDK 21 on PATH
python scripts/compile_generated.py --require-compiler
# Writes evidence/java_compile.json
Security Design
- No raw Cypher on the MCP surface — all graph queries are parameterised.
- Strict field allowlists — class names, package names, module names, versions, and config paths are checked against compiled regex patterns before any graph lookup.
- Source safety scanner — generated Java is rejected if it references
Runtime.getRuntime,ProcessBuilder,System.exit,java.io,java.nio.file, orjava.net. - Path traversal prevention — absolute paths and
..segments are rejected in config path fields. - Grounding enforcement — every import in generated source must correspond to a symbol retrieved from the graph for that exact version.
- LLM output re-validated — fields extracted by the LLM intent parser pass through the same
GenerationIntent.from_mapping()validation as direct API calls. - Neo4j credentials — loaded only from environment variables; never logged or returned in evidence artifacts.
- XML preflight —
defusedxmlprevents entity-expansion attacks in project-structure scanning. - Container — pinned Chainguard Linux image, non-root UID/GID 65532; CI performs an MCP-over-container stdio smoke test.
See SECURITY.md for the full threat boundary.
Repository Map
src/graph_mcp/
workflow.py intent parsing · graph lookup · Java generation · validation
graph_store.py Neo4j catalog adapter (parameterised Cypher)
llm_intent_parser.py Agent 1 — LLM free-form NL → GenerationIntent
review_agent.py Agent 2 — LLM post-generation checklist reviewer
server.py FastMCP stdio server (7 tools)
evaluation.py candidate scoring and selection harness
fixtures/
synthetic_graph.json CC0 versioned framework symbol catalog (SHA-256 bound)
evaluation_cases.json 96 CC0 natural-language intents (32/32/32 split)
java_framework/ 7 independently generated Java stub classes
evidence/
claims.json machine-readable claims ledger (14 public claims)
evaluation_protocol.json pre-declared selection rules and safety gates
task_evaluation.json per-candidate, per-split, per-case results
evaluation_trace.json confirmation case-level trace
neo4j_integration.json live Neo4j integration result
mcp_benchmark.json MCP protocol benchmark (120 calls)
java_compile.json ECJ compilation result
scripts/
build_evaluation_fixture.py generate benchmark from seed
evaluate_workflow.py run and score all four candidates
validate_evidence.py verify claims ledger and privacy rules
benchmark_mcp.py official MCP stdio latency benchmark
verify_neo4j.py live graph integration check
compile_generated.py ECJ compile gate
seed_graph.py materialise fixture into Neo4j
tests/
test_generation_loop.py generation + validation unit tests
test_graph_store.py Neo4j adapter unit tests
test_mcp_protocol.py official MCP protocol conformance
test_evaluation.py evaluation harness tests
test_evidence.py claims ledger integrity tests
test_neo4j_live.py opt-in live graph tests (NEO4J_* env required)
docs/
ARCHITECTURE.md component design and data flow
POLICY_CARD.md candidate selection details
DATA_CARD.md fixture provenance and license
MCP_INTEGRATION.md client configuration guide
DEPLOYMENT.md Docker and container notes
templates/ MCP prompt templates for VS Code Copilot
examples/ sample project preflight scanner
Boundaries
The following are not claimed by this repository:
- Free-form intent parsing quality independent of model version — the LLM pipeline is opt-in and its results are not captured in the frozen evaluation artifacts.
- Compatibility with any proprietary or confidential Java test framework.
- Production latency SLO — all measurements are single-process local sequential benchmarks.
- Concurrent, distributed, or high-availability operation.
- Automatic execution of generated Java against hardware or a test instrument.
- Any productivity, cost, yield, or test-time saving — this repository contains only generation and validation evidence.
The full machine-readable boundary is in evidence/claims.json.
License
Repository code: MIT. Graph fixture, intent cases, and Java stubs: CC0-1.0 (labelled in fixture metadata).
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。