skeletongraph

skeletongraph

Provides function-level structural retrieval for coding agents using tree-sitter, with zero-LLM indexing and MCP serving, enabling exact function localization and cost savings.

Category
访问服务器

README

<!-- mcp-name: io.github.yashdoke7/skeletongraph --> <p align="center"> <img src="docs/paper/figures/skeletongraph_banner.png" alt="SkeletonGraph — the exact function, not a pile of files. Zero-LLM structural retrieval for AI coding agents, served over MCP." width="100%"> </p>

<p align="center"> <a href="https://pypi.org/project/skeletongraph/"><img src="https://img.shields.io/pypi/v/skeletongraph.svg?color=blue" alt="PyPI"></a> <a href="https://pypi.org/project/skeletongraph/"><img src="https://img.shields.io/pypi/pyversions/skeletongraph.svg" alt="Python versions"></a> <a href="LICENSE"><img src="https://img.shields.io/badge/license-MIT-green.svg" alt="MIT License"></a> <a href="https://modelcontextprotocol.io"><img src="https://img.shields.io/badge/MCP-server-orange.svg" alt="MCP server"></a> <a href="https://registry.modelcontextprotocol.io/v0/servers?search=skeletongraph"><img src="https://img.shields.io/badge/MCP%20Registry-listed-blueviolet.svg" alt="MCP Registry"></a> <img src="https://img.shields.io/badge/index-zero--LLM-16a34a" alt="Zero-LLM index"> </p>

<p align="center"> <strong>Works with</strong>  <img src="https://img.shields.io/badge/Cursor-1A1A1A?style=flat-square&logo=cursor&logoColor=white" alt="Cursor"> <img src="https://img.shields.io/badge/Claude%20Code-D97757?style=flat-square&logo=anthropic&logoColor=white" alt="Claude Code"> <img src="https://img.shields.io/badge/GitHub%20Copilot-000000?style=flat-square&logo=githubcopilot&logoColor=white" alt="GitHub Copilot"> <img src="https://img.shields.io/badge/Codex-412991?style=flat-square&logo=openai&logoColor=white" alt="Codex"> <img src="https://img.shields.io/badge/Antigravity-4285F4?style=flat-square&logo=google&logoColor=white" alt="Antigravity"> <img src="https://img.shields.io/badge/Windsurf-09B6A2?style=flat-square&logo=codeium&logoColor=white" alt="Windsurf"> <img src="https://img.shields.io/badge/Cline-1A1A1A?style=flat-square" alt="Cline"> <img src="https://img.shields.io/badge/+%20any%20MCP%20client-333333?style=flat-square" alt="Any MCP client"> </p>

<p align="center"> <strong>Languages</strong>  <img src="https://img.shields.io/badge/Python-3776AB?style=flat-square&logo=python&logoColor=white" alt="Python"> <img src="https://img.shields.io/badge/JavaScript-F7DF1E?style=flat-square&logo=javascript&logoColor=black" alt="JavaScript"> <img src="https://img.shields.io/badge/TypeScript-3178C6?style=flat-square&logo=typescript&logoColor=white" alt="TypeScript"> <img src="https://img.shields.io/badge/Go-00ADD8?style=flat-square&logo=go&logoColor=white" alt="Go"> <img src="https://img.shields.io/badge/Rust-000000?style=flat-square&logo=rust&logoColor=white" alt="Rust"> <img src="https://img.shields.io/badge/Java-ED8B00?style=flat-square&logo=openjdk&logoColor=white" alt="Java"> <img src="https://img.shields.io/badge/C%23-239120?style=flat-square&logo=csharp&logoColor=white" alt="C#"> <img src="https://img.shields.io/badge/C%2B%2B-00599C?style=flat-square&logo=cplusplus&logoColor=white" alt="C++"> <img src="https://img.shields.io/badge/Ruby-CC342D?style=flat-square&logo=ruby&logoColor=white" alt="Ruby"> <img src="https://img.shields.io/badge/PHP-777BB4?style=flat-square&logo=php&logoColor=white" alt="PHP"> </p>

Coding agents burn tokens reading whole files to find one function. SkeletonGraph indexes your repo with tree-sitter — no LLM — and hands the agent the exact function to edit, over MCP.

<picture> <source srcset="docs/paper/figures/skeletongraph_hero.gif" media="(prefers-reduced-motion: no-preference)"> <img src="docs/paper/figures/skeletongraph_hero_poster.png" alt="SkeletonGraph pipeline: parse a repo with tree-sitter into a function-level structural index and call graph (no LLM); at query time, fuse BM25 lexical, dense semantic, and structural-rerank signals; return the one function to edit, served to the coding agent over MCP" width="100%"> </picture>

<p align="center"><em>Index once (no LLM) → fuse lexical + semantic + structural signals → return the exact function, served to your agent over MCP.</em></p>

SkeletonGraph is a retrieval engine purpose-built for coding agents, not a general RAG library retrofitted onto code. It parses a repository into function-level structure, a cross-file call graph, and PageRank centrality with zero LLM calls — deterministic, cheap, and instant to rebuild after every edit. At query time it resolves the symbols an issue names, walks the call graph outward, and reranks a BM25 recall pool by structural confirmation so the agent lands on the right function on the first try, instead of grepping and re-reading its way there. Its leaner operating point, sg-rerank (the product default), skips the dense leg entirely and still delivers the best file and function recall of any method we benchmarked it against — at the lowest token cost of any of them.

The thesis: code-context tools have mostly been validated as a token-optimization game — how few tokens can you spend. SkeletonGraph re-centers the question on retrieval quality — did the agent land on the correct function — of which lower token cost turns out to be a consequence, measurable only end-to-end inside a real agent loop, not in an offline benchmark.

Results

All numbers below are regenerated from the released run artifacts (python -m eval.scripts.make_paper_figures). The full verified ledger, including withdrawn claims, is in docs/paper/FINDINGS.md.

1. Controlled retrieval ablation (react loop, open-weight model, 100 tasks)

Identical action space for every arm; only the retrieval backend changes. The none arm gets no code access at all and establishes the memorization floor.

arm pass@1 file recall@1 function hit tokens (k) turns $/task
sg-fusion 42.0% .737 57% 180 21.9 .052
bm25 41.0% .642 43% 264 24.6 .074
graphify (knowledge graph) 41.0% .223 9% 275 25.6 .078
grep 39.0% .647 0% 282 22.4 .079
aider (repo-map) 36.7% 1,126 18.1 .160
none (no retrieval) 35.0% 345 23.6 .066

sg-fusion is the top arm, the cheapest arm, and the only one that localizes to the function (57% vs grep's 0% — lexical search is file-granular by construction). Against the closed-book floor of 35.0%, retrieval is worth +7 points here.

sg-rerank's recall/cost profile is reported separately in the agent-free intrinsic retrieval ablation in docs/paper/skeletongraph.tex (Table 2, §5.1) — best MRR/recall@10 short of full fusion, at the lowest index cost.

2. Deployment: SkeletonGraph vs native Claude Code (MCP, Docker-verified)

The product itself — SG as an MCP server driving Claude Code (sonnet) against Claude Code on its own tools. 100 paired SWE-bench Verified tasks:

arm pass@1 file recall@1 turns $/task
native (Claude's own Grep/Read) 74/100 .663 14.5 .434
sg-fusion (SkeletonGraph MCP) 75/100 .862 11.4 .371

<sub>SG's first-search recall excludes 3 tasks where the agent never called SG at all — those are adoption events, not retrieval failures. Including them gives .836.</sub>

Equivalent solve rate at −14.6% cost and −21.4% turns. The saving is not spread evenly — it lives almost entirely in the tail:

cost percentile native +SG change
50th (median task) $0.255 $0.260 +1.9%
90th $1.010 $0.752 −25.6%
95th (worst tasks) $1.559 $0.896 −42.5%

Retrieval does nothing for the typical task and removes over 40% of the cost of the worst ones. Paired bootstrap 95% CI on the mean: [−25.3%, −1.2%]; McNemar on pass@1: p = 1.0 (no difference).

3. The ceiling: what structural retrieval cannot do

Retrieval collapses as location cues are removed

sg-fusion runs all three non-LLM retrieval paradigms at once — lexical (BM25), semantic (code embeddings), and topological (call graph). Crossing memorization (standard vs. decontaminated benchmark) against location cues (original vs. prose-stripped issue text) shows the limit:

condition file recall native → SG Δ recall cost turns
SWE-Verified, raw .661 → .861 +.200 −32.2% −38.5%
SWE-Verified, prose-only .717 → .711 −.006 −21.1% −22.6%
SWE-rebench (unseen repos), raw .500 → .639 +.139 −26.4% −35.6%
SWE-rebench, prose-only .394 → .439 +.044 −31.3% −29.3%

Read the Δ-recall and cost columns against each other. The retrieval advantage collapses — a +.200 edge becomes −.006 once the symbols are gone, and the lexical baseline falls in parallel, so this is a property of the whole non-LLM category, not of one implementation. Yet the cost saving persists in every condition (−21% to −31%), including the one where retrieval quality is statistically identical to the baseline. On the decontaminated benchmark the point sharpens: the largest cost saving (−31.3%) sits in the cell with the smallest retrieval edge (+.044).

What the agent was missing was direction, not files. Memorization and location cues look like two different factors, but they're the same quantity — knowledge of where to look, held by the model or supplied by the issue. Order the four conditions by how much of it is available and SG's recall orders exactly: .861 → .711 → .639 → .439.

And when direction isn't given, the agent acquires it by exploring. On the prose-only condition, native's cumulative recall reaches .967 against SG's .883 — given enough turns, the agent that grepped and read its way through the repo located the code better than the agent handed a ranked list. SG gets there in half the searches and for less money; it does not get further.

That inversion is the honest statement of what this tool is: a retrieval layer delivers candidate files, not knowledge of the repository — and a confident ranked list actively suppresses the agent's own acquisition of that knowledge (reads drop 3.86 → 1.55/task). The agent spends tokens until it's certain, not until it has files. That's why recall and cost decouple, and it's what the paper's title refers to — the honest ceiling is a finding here, not a caveat.

SWE-Verified rows are restricted to the same 15 tasks as the prose run so raw and prose are paired. That subset is not representative of the full 100 (SG saves −32.2% on it vs −14.6% overall) — do not compare it against the n=100 figure.

4. Deployment finding: slow MCP servers are structurally excluded

Two graph/LSP-based competitors wired as MCP servers never participated at all. Claude Code's headless (-p) mode finalizes its tool manifest within ~2 seconds of launch and never updates it; servers needing real bootstrap time (a language server, a Node CLI + index) finish their handshake just past that window and are silently absent for the entire run — confirmed via session-init transcripts and each server's own logs showing it was ready seconds later.

This is a deployment-mode result, not a retrieval-quality one, and we report no performance comparison for those systems: with zero tool calls, any such number would measure their absence rather than their retrieval. A separately wired zero-LLM graph competitor connected cleanly with all 14 tools visible, yet the agent never invoked one across 10 tasks, defaulting to native grep every time. Fast connection and actual adoption are prerequisites that retrieval quality cannot substitute for.

SkeletonGraph is wrapper-first: it returns a full context packet or exposes a retrieval index (AST skeletons + call graph + local summaries + optional embeddings) so the IDE agent or CLI can choose targets.

SkeletonGraph has two product surfaces:

  • SG IDE: MCP context server for Cursor, Claude Code, Copilot, Codex, Antigravity, Windsurf, and other agentic IDEs.
  • SG CLI: terminal pipeline for route, prepare, dry-run, provider execution, and cost-aware model selection.

Why SkeletonGraph

Most coding agents spend expensive turns discovering the repo:

search -> read file -> read neighbor -> read tests -> realize the target

SkeletonGraph moves that work into a deterministic graph pipeline:

prompt -> (optional) retrieval planner -> classify task -> find target nodes -> expand graph -> assemble packet

The goal is not only lower token cost. The useful product outcomes are:

  • fewer exploratory file reads
  • faster first useful answer
  • better target/test/blast-radius context
  • transparent routing reasons
  • lower model overkill for routine tasks
  • reusable packets for IDEs, CLIs, and other agents

Install

pip install skeletongraph           # core: indexing, MCP server, CLI (no API key needed)
pip install "skeletongraph[llm]"    # + litellm for sg run --execute / sg summarize --tier cloud
pip install "skeletongraph[all]"    # everything

Quick Start: SG IDE

Use this path when you already work inside Cursor, Claude Code, Copilot, Codex, Antigravity, or another MCP-capable coding environment.

cd your-project
sg init
sg build
sg doctor

sg init writes the MCP config and the agent instruction file for the selected IDE. SG IDE does not require an API key. Your IDE subscription/model still does the reasoning and editing; SkeletonGraph supplies the packet or retrieval signals for efficient target selection.

Supported IDE setup targets include:

IDE Integration Model switching
Cursor MCP + rules manual in IDE
Claude Code MCP + CLAUDE.md /model command
GitHub Copilot MCP + instructions manual in IDE
Codex MCP + AGENTS.md manual in agent
Antigravity MCP + rules manual in IDE
Windsurf MCP + rules manual in IDE

Quick Start: SG CLI

Use this path when you want a terminal-first context and model-routing pipeline.

cd your-project
sg build
sg route "fix the auth token validation bug"
sg prepare "fix the auth token validation bug" --out .skeletongraph/context.md
sg run "fix the auth token validation bug" --dry-run

sg route, sg prepare, and sg run --dry-run do not need an API key.

To call a provider:

sg config --cli-provider anthropic
$env:ANTHROPIC_API_KEY = "..."
sg run "fix the auth token validation bug" --execute

To test locally without a paid provider key:

ollama pull qwen3-coder:latest
ollama serve
sg config --cli-provider local
sg run "fix the auth token validation bug" --dry-run
sg run "fix the auth token validation bug" --execute

Local execution is intended for cheap pipeline testing. Use provider models for quality benchmarks unless the benchmark is specifically for local models.

Model Dependency, Prewarming, and Keeping the Index Fresh

SG downloads two small embedding models on first use, both via sentence-transformers (a hard dependency, not optional):

  • jinaai/jina-embeddings-v2-base-code (SG_DENSE_MODEL) — the semantic leg of fusion/sg_search. Loaded on sg warm or on an agent's first dense-retrieval query. Loads with trust_remote_code=True (Jina ships custom modeling code on the HF Hub) — this executes code from that model repo, same as any trust_remote_code model.
  • all-MiniLM-L6-v2 (SG_EMBED_MODEL) — a smaller, separate model used only as a confidence-score tiebreaker at index time. Downloads automatically on the first sg build, not on sg warm.

Both need internet access the very first time each is used on a machine — after that, both are cached locally (Hugging Face's model cache, plus SG's own content-hash caches: .skeletongraph/dense_cache for the dense leg, .skeletongraph/embeddings.npz for the confidence tiebreaker) — so later builds are incremental: only functions whose text actually changed get re-embedded.

Prewarm before launching an agent, so that cost lands during setup instead of on the agent's first real search:

sg build                 # parse + structural index (no LLM, fast)
sg warm --path .          # prebuild BM25 + dense caches (one-time; minutes on CPU)
sg warm --path . --mode rerank   # skip the dense leg entirely (no embedding cost)

Without this, the first sg_search call an agent makes pays the cold-encode cost inline — on a large repo this can exceed the dense retrieval leg's internal timeout (SG_DENSE_TIMEOUT_S, 20s by default), in which case it silently degrades to a 2-signal (lexical + structural) result rather than failing outright. Prewarming avoids relying on that fallback altogether.

Keeping the index current as files change — two options, pick based on how you work:

sg update --path .        # one-shot: re-index only files that changed since last build
sg watch --path .         # background daemon: auto-reindexes on save (needs `pip install "skeletongraph[daemon]"`)

sg watch is the hands-off option for active development — it debounces rapid saves and calls the same incremental update path as sg update, so editing a file is reflected in the index without a manual rebuild.

Model Routing

SkeletonGraph separates IDE-facing model labels from CLI provider model names.

For IDEs, model tiers are recommendations:

Tier Typical use
SLM docs, explanations, simple lookup
MLM normal coding, debugging, tests, review
LLM architecture, broad migrations, low-confidence tasks

For CLI execution, SkeletonGraph can route to provider model names:

sg config --cli-provider anthropic
sg config --cli-provider openai
sg config --cli-provider google
sg config --cli-provider local

Dynamic routing uses task mode, confidence, candidate count, token size, and complexity. Code-changing work keeps an MLM floor by default so cost savings do not come from making weak models edit code unsafely. Retrieval planning can use small models to propose targets over AST/summaries before the heavy model runs.

IDE Integration

After sg init and sg build, register SG as an MCP server and write IDE hooks:

sg install --ide claude-code   # Claude Code: hooks + MCP server + CLAUDE.md rules
sg install --ide cursor        # Cursor: MCP + .cursor/rules/skeletongraph.mdc + hooks
sg install --ide cline         # Cline: MCP config + rules block
sg install --ide roo           # Roo: MCP config + rules block
sg install --ide copilot       # GitHub Copilot: MCP + copilot-instructions.md
sg install --ide windsurf      # Windsurf: MCP + .windsurfrules
sg install --ide zed           # Zed: MCP config + rules block
sg install --ide continue      # Continue: MCP config + rules block
sg install                     # auto-detect all installed IDEs

codex and antigravity are accepted as aliases and currently route through the Copilot-style MCP installer.

For any other MCP-capable client, or to configure it by hand, see mcp.example.json for the raw server config (sg serve --path /path/to/your/project).

After install, restart your editor. SkeletonGraph runs as a background MCP server (sg serve --path .) that the IDE connects to automatically.

MCP Tools

Seven tools are exposed to the IDE agent. Use these instead of grep/glob/file reads:

Tool When to call Returns
sg_overview Session start — once per session Constraints + top-N functions (by PageRank) + recent turns + index stats
sg_search "query" Primary retrieval — almost every prompt Top-3 matches with body excerpts + summaries + 1-hop callers; top-4..N as signatures + summaries. One call usually enough — no need to chain.
sg_get "fqn" When the exact FQN is known Signature + summary + 1-hop callers + callees
sg_expand "target" When more body is needed than sg_search returned Full function body / file / line range (token-capped)
sg_constraint list / propose Before proposing changes Confirmed + proposed project rules
sg_log Reviewing recent session turns Last-N turn summaries with files touched
sg_decision A design/implementation choice is made (picked or rejected, and why) Recorded so it survives context compaction — recall later with sg_log(kind="decision")

Smart context routing. On each UserPromptSubmit, SG classifies the prompt (architecture / explain / decision / debug / test / review / general) and includes the matching MD file from .skeletongraph/ — e.g. architecture.md only for design/refactor queries, project.md only for "what is this codebase" queries. Constraints + session digest + relevant functions are always injected.

Cold start. If no .skeletongraph/ index exists when an MCP tool is called, SG auto-builds on first invocation (see auto_build_on_query in config).

CLI Reference

Indexing & status

Command Purpose
sg init [--agent cursor] Configure project, IDE preset, MCP, constraints
sg index Full index (alias for sg build)
sg index --incremental Only re-index changed files
sg build Full index with detailed output
sg update Incremental update
sg status Show index status
sg doctor Check index, routing, provider, Ollama readiness
sg overview Project skeleton: top functions, constraints, session
sg install [--ide <name>] Write IDE hooks + MCP config

Retrieval

Command Purpose
sg search "query" BM25 + graph search (no API key)
sg get "fqn" Get function signature, summary, callers
sg expand "target" Expand function body / file / line range

Constraints & session

Command Purpose
sg constraint list List all constraints
sg constraint propose "text" Add a proposal
sg constraint confirm <id> Promote proposal → decisions.md
sg constraint remove <id> Remove a constraint
sg constraint aggregate Import from IDE rule files
sg log [--last-n 10] Show recent session turns

Summarization

Command Purpose API key
sg summarize --tier local Ollama Tier-0.5 (free, on-device) no
sg summarize --tier cloud Cloud LLM Tier-1 provider key
sg summarize --tier cloud --force Re-summarize all functions provider key

Model routing & execution

Command Purpose API key
sg route "task" Show task mode, tier, recommended model no
sg run "task" --dry-run Plan routed execution no
sg run "task" --execute Call configured provider provider or local
sg config [--agent cursor] Configure IDE and CLI models no
sg config --cli-provider anthropic Set CLI execution provider no

Background indexing

Command Purpose
sg watch Daemon: auto-reindex files on save

Provider output from sg run --execute is written to .skeletongraph/runs/. Evaluation is currently done externally via a SWE-bench harness (see the Evaluation section below).

Python API

from skeletongraph.engine import SGEngine

engine = SGEngine(project_root=".")
result = engine.query("fix the content-length bug", delivery="cli")

print(result.context_text)
print(result.query_mode)
print(result.model_tier)
print(result.recommended_model)
print(result.routing_reason)

Architecture

src/skeletongraph/
  parser/       AST extraction
  graph/        dependency graph and ranking
  storage/      .skeletongraph persistence
  retrieval/    classification, resolution, model routing
  assembly/     context packet construction
  session/      memory and dedup
  server/       MCP server
  install/      per-IDE hook + MCP config writers (`sg install`)
  hooks/        IDE hook handlers (prompt-submit routing, etc.)
  llm/          LiteLLM wrapper for optional CLI execution
  cli/          Click commands
  engine.py     unified query pipeline

Evaluation

The full methodology, verified results, and every withdrawn/superseded claim are in docs/paper/skeletongraph.tex and docs/paper/FINDINGS.md.

SkeletonGraph should be evaluated on both quality and cost:

  • target recall and packet completeness
  • missed tests/callers
  • first useful answer latency
  • file reads after SG context
  • pass rate
  • cost per passing task
  • dynamic routing overkill/underpower rate
  • IDE compliance with SG-first context usage

Cost savings are only meaningful when reported with pass rate.

License

SkeletonGraph is released under the MIT License — free to use, modify, and distribute, for commercial and private projects alike.

Citation

If SkeletonGraph is useful in your research, please cite:

@misc{doke2026skeletongraph,
  title  = {Retrieval Is Not Reasoning: The Localization Ceiling in AI Coding Agents},
  author = {Doke, Yash},
  year   = {2026},
  note   = {SkeletonGraph — zero-LLM structural retrieval for coding agents},
  url    = {https://github.com/yashdoke7/skeletongraph}
}

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选