Corpus

Corpus

MCP server that aggregates documentation and code context across GitHub repositories, auto-generates a system map from Backstage catalog files, and provides search capabilities for docs, code, issues, and API schemas to AI agents.

Category
访问服务器

README

<div align="center"> <img src="assets/logo.jpg" width="200" height="200" alt="Corpus Logo"> <h1>📚 Corpus</h1>

<p> <em>A corpus is a large, structured collection of written or spoken texts stored on a computer and used for linguistic research, or a complete body of written work by a single author.</em> </p>

<p> <strong>Create a Corpus for your entire codebase.</strong> </p>

<p> <a href="https://modelcontextprotocol.io"> <img src="https://img.shields.io/badge/MCP-Ready-blue?style=for-the-badge&logo=modelcontextprotocol" alt="MCP Ready" /> </a> <a href="https://nodejs.org"> <img src="https://img.shields.io/badge/Node.js-22+-green?style=for-the-badge&logo=node.js" alt="Node.js 22+" /> </a> <a href="https://github.com"> <img src="https://img.shields.io/badge/GitHub-Integrated-black?style=for-the-badge&logo=github" alt="GitHub Integrated" /> </a> </p> </div>


Corpus aggregates the documentation across all your organization's repositories, builds a live system map using Spotify Backstage catalog entities, and puts it all behind a powerful Model Context Protocol (MCP) server.

It supports GitHub, GitLab, Bitbucket, and Azure DevOps out of the box, with an extensible adapter architecture to support any VCS provider.

Give your AI agents (Claude, Copilot, etc.) the holistic context they need to understand your architecture, service ownership, docs, and code—all in one place!

✨ Features

  • 🗺️ Auto-Generated Entity Graph: Fully parses Backstage catalog-info.yaml entities (Components, APIs, Systems, Users) and generates a bidirectional relationship graph using well-known relations (e.g., ownerOf/ownedBy, providesApi/apiProvidedBy).
  • 📖 Centralized Doc Search: Search across README.md, docs/**/*.md, adr/**/*.md, and AI skills across your entire org. Choose between fast Lexical (keyword) search or state-of-the-art Semantic (vector embedding) search run locally in-memory.
  • 🔍 Global Code Search: Keyword search across all organization repositories via your provider's native Code Search APIs.
  • 💬 Issue & PR Context: Proxies to your VCS's search APIs to find discussions, PRs, and issues across the org (search_issues_and_prs).
  • 📖 File Reading: Direct access to precise file contents from any repository branch or commit.
  • 🔗 API Schema Aggregation: Automatically indexes openapi and swagger files so agents can pull down endpoint contracts instantly (list_api_schemas).
  • 🧠 Automated Skill Directory: Scans all repositories for .[client]/skills/**/SKILL.md files (like .agents, .claude, or .cursor) and extracts their YAML frontmatter, allowing AI agents to dynamically discover and learn your company's standard operating procedures (list_company_skills).
  • ⚡ Zero-Config Start: Auto-runs missing builds on startup. If you have credentials, just hit npm start and the server fetches and indexes everything.
  • 🚨 Gap Reporting: Optional ability to file a central Issue / Work Item when docs fail to answer an agent's question.

🛠️ Quick Start

1. Prerequisites

  • Node.js v22+
  • VCS Authentication:
    • GitHub: Requires a PAT (Classic: repo, read:org | Fine-Grained: Contents: Read-only, Metadata: Read-only).
    • GitLab: Requires a Personal Access Token with read_api and read_repository scopes.
    • Bitbucket: Requires an App Password with repository:read and workspace:read scopes.
    • Azure DevOps: Requires a Personal Access Token with Code (Read) scope.

(Note: If you enable ENABLE_GAP_REPORTING, ensure your token also has Write permissions for Issues / Work Items).

2. Configure Environment

Create a .env file in the root directory:

Global Search Settings (Optional):

# Choose between 'lexical' (default keyword search) or 'semantic' (local embeddings)
DOC_SEARCH_METHOD=lexical

For GitHub (Default):

VCS_PROVIDER=github
GIT_ORG=your-github-org-or-username
GIT_PAT=your-github-personal-access-token

# Optional
ENABLE_GAP_REPORTING=false
GITHUB_PROJECT=your-github-org/doc-gaps-repo

For GitLab:

VCS_PROVIDER=gitlab
GIT_PAT=your-gitlab-personal-access-token
GIT_ORG=your-gitlab-group-name # Optional: Scopes discovery to a specific group
GITLAB_URL=https://gitlab.com # Optional: Change if using self-hosted GitLab

# Optional
ENABLE_GAP_REPORTING=false
GITHUB_PROJECT=your-gitlab-project-id # The Project ID where doc gaps are filed

For Bitbucket:

VCS_PROVIDER=bitbucket
GIT_ORG=your-workspace-name
GIT_PAT=your-username:your-app-password # Basic Auth or Bearer token

# Optional
ENABLE_GAP_REPORTING=false
GITHUB_PROJECT=your-workspace/your-repo-for-gaps

For Azure DevOps:

VCS_PROVIDER=azure
GIT_ORG=your-organization-name
GIT_PAT=your-personal-access-token

# Optional
ENABLE_GAP_REPORTING=false
GITHUB_PROJECT=your-project-name # The Azure Project where Doc Gaps (Work Items) are filed

3. Build & Run

Local Execution:

npm install
npm run build
npm start

Note: npm start automatically kicks off the corpus and system-map generation scripts if they haven't been run yet.

Docker Execution:

docker build -t corpus-mcp .
docker run -i -e GIT_ORG=your-github-org -e GIT_PAT=your-github-pat corpus-mcp

🤖 Registering with AI Clients

Antigravity

Antigravity natively supports MCP. Configure the server globally by adding it to ~/.gemini/config/mcp_config.json:

{
  "mcpServers": {
    "corpus": {
      "command": "node",
      "args": ["/absolute/path/to/code-context-mcp/dist/src/index.js"],
      "env": {
        "GIT_ORG": "your-github-org",
        "DOTENV_CONFIG_PATH": "/absolute/path/to/code-context-mcp/.env",
        "CORPUS_DIR": "/absolute/path/to/code-context-mcp/corpus"
      }
    }
  }
}

Claude Desktop

Add this to your claude_desktop_config.json:

{
  "mcpServers": {
    "corpus": {
      "command": "node",
      "args": ["/absolute/path/to/code-context-mcp/dist/src/index.js"],
      "env": {
        "GIT_ORG": "your-github-org",
        "GIT_PAT": "your-github-pat",
        "CORPUS_DIR": "/absolute/path/to/code-context-mcp/corpus"
      }
    }
  }
}

Claude Code

Run the following in the project root:

claude mcp add corpus "node $(pwd)/dist/src/index.js"

🏗️ Architecture & Commands

  • npm run build:corpus: Crawls the configured VCS organization and downloads docs + catalog data into corpus/manifest.json.
  • npm run build:map: Transforms the manifest into an active dependency graph saved to corpus/system-map.yaml.
  • npm run build: Runs the full pipeline and compiles TypeScript.
  • npm run test: Runs unit tests using the native Node.js test runner.

🧩 System Map & catalog-info.yaml

Corpus automatically generates a global dependency graph of your organization's services. To participate in the system map, each repository should contain a catalog-info.yaml file at its root, conforming to the Backstage Descriptor Format.

Because Corpus acts like a Backstage catalog processor, it extracts any entity type (Component, API, System, Group) and automatically wires up bidirectional relationships. If your Component defines owner: group:auth-team and providesApis: [api:auth-api], Corpus automatically generates the ownedBy/ownerOf and providesApi/apiProvidedBy edges so AI agents can natively traverse your organization's entire service graph.

Example catalog-info.yaml:

apiVersion: backstage.io/v1alpha1
kind: Component
metadata:
  name: my-auth-service
  description: Handles user authentication and token generation
spec:
  type: service
  lifecycle: production
  owner: group:auth-team
  providesApis:
    - api:auth-api
  dependsOn:
    - component:user-database
    - component:email-service

💡 Best Practices & Philosophy

To get the absolute most out of Corpus and your AI agents, we recommend the following ecosystem practices:

  1. Keep Docs Close to Code: Documentation should live in the repository next to the code. The best place to document how a system works is directly beside the system itself. Corpus automatically picks up docs/**/*.md and adr/**/*.md across all your repos.
  2. Central Wiki Repository: If you have company-wide architectural decisions, RFCs, or code-quality standards that span multiple systems, keep them in a central "Wiki" repository as markdown files. Corpus will aggregate them perfectly.
  3. Synergy with Spotify Backstage: If you use Backstage, Corpus is the perfect companion.
    • Backstage is an Internal Developer Portal (IDP) built for humans, providing a rich web UI.
    • Corpus is an IDP built for AI Agents, exposing the exact same context over MCP. Because Corpus natively parses standard catalog-info.yaml files, there is zero duplicated work. If your teams are already defining dependsOn, lifecycle, and owner tags for Backstage, Corpus automatically scoops them up and translates them into an active graph that AI agents can traverse.
  4. Frequent Automated Updates: The Corpus is meant to be a living, breathing snapshot of your organization. Running the build scripts (npm run build) re-fetches and rebuilds the corpus locally. Because it's a simple API scraping script, it consumes zero LLM tokens to build. Ideally, Corpus should be deployed centrally within your company, using a cron job (like a GitHub Action) to rebuild the manifest.json every night and distribute it to your developers.

🤝 Contributing

We welcome contributions! Please see our Contributing Guidelines for details on how to get started, set up your development environment, and submit Pull Requests.

This project enforces Conventional Commits. A pre-commit hook automatically formats your code with Prettier and checks it with ESLint.

See the Setup Skill Guide for more details.

📄 License

Corpus is free to use. All intellectual property is owned by Sayam Hussain.

This project is licensed under the MIT License.

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选