kimi-code-mcp

kimi-code-mcp

Connects Kimi Code with Claude Code, enabling Claude to delegate bulk codebase reading to Kimi (256K context) for cost savings, while Claude focuses on reasoning and code edits.

Category
访问服务器

README

kimi-code-mcp

English | 中文說明


MCP server that connects Kimi Code (K2.5, 256K context) with Claude Code — letting Claude orchestrate while Kimi handles the heavy reading.

<div align="center"> <img src="assets/llm-cost-vs-intelligence.png" alt="LLM Cost vs Intelligence — Kimi K2.5 delivers frontier-level intelligence at a fraction of the cost" width="720" /> <br /> <sub>Kimi K2.5 sits on the efficiency frontier — near-Claude intelligence at 10x lower cost. <a href="https://www.kimi.com/code">kimi.com/code</a></sub> </div>

[!TIP] Stop paying Claude to read files. Kimi K2.5 delivers frontier-class code intelligence at a fraction of the cost (see chart above). Delegate bulk codebase scanning to Kimi (256K context, near-zero cost) and let Claude focus on what it does best — reasoning, decisions, and precise code edits. One kimi_analyze call can replace 50+ file reads.

What is Kimi Code?

Kimi Code is an AI code agent by Moonshot AI, powered by the Kimi K2.5 model (1T MoE, 256K context). It works across Terminal, IDE, and CLI — writing, debugging, refactoring, and analyzing code autonomously.

Key specs:

  • 256K token context — reads entire codebases in one pass
  • Parallel agent spawning — handles concurrent tasks
  • Shell, file, and web access — full developer toolchain
  • Install: curl -L code.kimi.com/install.sh | bash

[!WARNING] Kimi Code membership required. This MCP server calls the Kimi CLI under the hood, which requires an active Kimi Code plan. Make sure you have a valid subscription and have run kimi login before use.

Plan Price Notes
Moderato $0 (7-day free trial) Then $19/mo. Good for trying it out
Allegretto $39/mo Recommended — higher weekly quota + concurrency
Allegro $99/mo For daily, heavy-duty development
Vivace $199/mo Max quota for large codebases

Annual billing saves up to $480. All plans include Kimi membership benefits.

Quick Start

# 1. Install Kimi CLI and log in
curl -L code.kimi.com/install.sh | bash
kimi login

# 2. Install via npm
npm install -g kimi-mcp-server

Add to .mcp.json (project-level or ~/.claude/mcp.json for global):

{
  "mcpServers": {
    "kimi-code": {
      "command": "npx",
      "args": ["-y", "kimi-mcp-server"]
    }
  }
}

Or build from source:

git clone https://github.com/howardpen9/kimi-code-mcp.git
cd kimi-code-mcp && npm install && npm run build
{
  "mcpServers": {
    "kimi-code": {
      "command": "node",
      "args": ["/absolute/path/to/kimi-code-mcp/dist/index.js"]
    }
  }
}

Run /mcp in Claude Code to verify — you should see kimi-code with 7 tools.

Kimi Code API Setup

[!NOTE] Kimi Code API and Moonshot API are separate providers — their API keys are not interchangeable.

There are two ways to configure the Kimi Code API for the CLI:

Option 1: OAuth Login (Recommended)

In the Kimi Code CLI shell, run:

kimi

Then use the /login (or /setup) command:

/login
  1. Select Kimi Code as the platform
  2. Your browser opens for OAuth authorization
  3. Config is saved automatically to ~/.kimi/config.toml

Option 2: Manual API Key Configuration

Get your API Key

  1. Visit code.kimi.com
  2. Sign in → Settings → API Keys
  3. Create a new key (starts with sk-, shown only once)

Edit config file

nano ~/.kimi/config.toml

Add:

[providers.kimi-code]
type = "kimi"
base_url = "https://api.kimi.com/coding/v1"
api_key = "sk-your-api-key"

[models.kimi-for-coding]
provider = "kimi-code"
model = "kimi-for-coding"
max_context_size = 262144
capabilities = ["thinking"]

[defaults]
model = "kimi-for-coding"

Using environment variables (recommended for security)

# Add to ~/.zshrc (macOS) or ~/.bashrc (Linux)
export KIMICODE_API_KEY="sk-your-api-key"

Then reference it in config.toml:

[providers.kimi-code]
type = "kimi"
base_url = "https://api.kimi.com/coding/v1"
api_key = "${KIMICODE_API_KEY}"

Multi-provider config example

You can configure both Kimi Code and Moonshot side by side:

[providers.kimi-code]
type = "kimi"
base_url = "https://api.kimi.com/coding/v1"
api_key = "${KIMICODE_API_KEY}"

[providers.moonshot-cn]
type = "kimi"
base_url = "https://api.moonshot.cn/v1"
api_key = "${MOONSHOT_API_KEY}"

[models.kimi-for-coding]
provider = "kimi-code"
model = "kimi-for-coding"
max_context_size = 262144
capabilities = ["thinking"]

[models.kimi-k2]
provider = "moonshot-cn"
model = "kimi-k2-0905-preview"
max_context_size = 256000
capabilities = ["thinking"]

[defaults]
model = "kimi-for-coding"

Switch models at any time with /model or /model kimi-k2 in the CLI.

Kimi Code vs Moonshot

Feature Kimi Code Moonshot
Focus Optimized for coding General-purpose chat
Endpoint api.kimi.com/coding/v1 api.moonshot.cn/v1
API Key Separate — apply at code.kimi.com Separate
SearchWeb / FetchURL Built-in Not available
Context 262K 256K

What You Can Do

Just tell Claude what you need. It will delegate to Kimi automatically:

Prompt What happens
"Analyze this codebase's architecture" Kimi reads all files (256K ctx), Claude acts on the report
"Scan for security vulnerabilities, then review Kimi's findings" Kimi audits, Claude cross-examines — AI pair review
"Map all dependencies of the auth module, then plan the refactoring" Kimi builds the dependency graph, Claude plans the changes
"Review the recent changes for regressions and edge cases" Kimi reviews full context (not just the diff), Claude synthesizes
"Resume the last Kimi session and ask about the API design" Kimi retains 256K tokens of context across sessions

Why This Exists

Claude Code is powerful but expensive. Every file it reads costs tokens. Meanwhile, many tasks — pre-reviewing large codebases, scanning for patterns, generating audit reports — are high-certainty work that doesn't need Claude's full reasoning power.

[!IMPORTANT] The cost equation: Claude reads 50 files to understand your architecture = expensive. Kimi reads 50 files via kimi_analyze = near-zero cost. Claude then acts on Kimi's structured report = minimal tokens. Total savings: 60-80% fewer Claude tokens on analysis-heavy tasks.

How It Saves Tokens

                          ┌─────────────────────────────┐
                          │   You (the developer)       │
                          └──────────┬──────────────────┘
                                     │ prompt
                                     ▼
                          ┌─────────────────────────────┐
                          │   Claude Code (conductor)   │
                          │   - orchestrates workflow    │
                          │   - makes decisions          │
                          │   - writes & edits code      │
                          └──────┬──────────────┬───────┘
                      precise    │              │  delegate
                      edits      │              │  bulk reading
                      (tokens)   │              │  (FREE)
                                 ▼              ▼
                          ┌──────────┐   ┌──────────────┐
                          │ your     │   │  Kimi Code   │
                          │ codebase │   │  (K2.5)      │
                          └──────────┘   │  - 256K ctx  │
                                         │  - reads all │
                                         │  - reports   │
                                         └──────────────┘
  1. Claude receives your task → decides it needs codebase understanding
  2. Claude calls kimi_analyze via MCP → Kimi reads the entire codebase (256K context, near-zero cost)
  3. Kimi returns a structured analysis
  4. Claude acts on the analysis with precise, targeted edits

Result: Claude only spends tokens on decision-making and code writing, not on reading files.

Mutual Code Review with K2.5

Kimi Code is powered by K2.5 — a 1T MoE model designed for deep code comprehension. This enables AI pair review:

  1. Kimi pre-reviews — 256K context means it sees the entire codebase at once: security issues, anti-patterns, dead code, architectural problems
  2. Claude cross-examines — reviews Kimi's findings, challenges questionable items, adds its own insights
  3. Two perspectives — different models catch different things. What one misses, the other finds

Use Kimi as a Code Reviewer

Beyond ad-hoc analysis, you can use Kimi as a dedicated reviewer in your workflow:

PR Review Workflow

┌──────────────┐   diff    ┌──────────────┐  structured  ┌──────────────┐
│   Your PR    │ ────────► │  Kimi Code   │  findings    │  Claude Code │
│  (changes)   │           │  (reviewer)  │ ────────────►│  (decision)  │
└──────────────┘           └──────────────┘              └──────────────┘

Continuous Audit Pattern

When What Why
Before merging Kimi scans diff + affected modules Catch regressions early
Weekly Full codebase sweep Accumulated tech debt
Pre-release Security-focused audit Ship with confidence

Each review session can be resumed (kimi_resume) — Kimi retains up to 256K tokens of context from previous sessions, building understanding over time.

What Kimi Reviews Well

Review Type Why Kimi Excels
Security audit 256K context sees full attack surface, not just isolated files
Dead code detection Can trace imports/exports across entire codebase
API consistency Compares patterns across all endpoints simultaneously
Dependency analysis Maps full dependency graph in one pass
Architecture review Sees the forest and the trees at the same time

Tools

Tool Description Timeout
kimi_analyze Deep codebase analysis (architecture, audit, refactoring) 10 min
kimi_query Quick programming questions, no codebase context 2 min
kimi_list_sessions List existing Kimi sessions with metadata instant
kimi_resume Resume a previous session (up to 256K token context) 10 min
kimi_status Check CLI installation, version, and authentication status instant
kimi_cache_status View session cache statistics and performance metrics instant
kimi_cache_invalidate Manually invalidate cached sessions (by dir or all) instant

Output Control Parameters

kimi_analyze and kimi_resume support these parameters to control output size:

Parameter Values Default Effect
detail_level summary / normal / detailed normal Controls prompt-side verbosity instructions
max_output_tokens number 15000 Hard ceiling — output truncated at clean boundary if exceeded
include_thinking boolean false Include Kimi's internal reasoning chain (10-30K extra tokens)

kimi_query also supports max_output_tokens and include_thinking.

Token Economics

[!NOTE] The savings come from compression ratio, not from free reading. Kimi's subscription cost still applies, but the key benefit is reducing expensive Claude Code token consumption.

                    Without kimi-code-mcp        With kimi-code-mcp (normal)
                    ─────────────────────        ───────────────────────────
Raw source:         50 files × ~4K = 200K        Kimi reads (subscription cost)
Claude reads:       200K tokens                  5-15K token report
Claude token cost:  $$$                          $

Compression ratio by detail_level:

Level Compression Output Size Equivalent Source Best For
summary 40-100x ~2-5K tokens ~8-20K chars / ~200-500 lines of code Quick orientation, file inventory
normal 15-40x ~5-15K tokens ~20-60K chars / ~500-1500 lines of code Architecture review, dependency mapping
detailed 5-15x ~15-40K tokens ~60-160K chars / ~1500-4000 lines of code Security audit with code snippets

When savings happen:

  • Large codebases (50+ files) — architecture understanding, cross-file scanning
  • Security audits, dead code detection, API consistency checks
  • Pre-review before targeted edits (scan first → edit specific files)

When to skip and let Claude read directly:

  • Small codebases (<10 files) — direct reading is faster
  • Single-file modifications — Claude's built-in file reading is sufficient
  • When you need every line of code — detailed output approaches raw reading cost

How It Works

┌──────────────┐  stdio/MCP   ┌──────────────┐  subprocess   ┌──────────────┐
│  Claude Code │ ◄──────────► │ kimi-code-mcp│ ────────────► │ Kimi CLI     │
│  (conductor) │              │ (MCP server) │               │ (K2.5, 256K) │
└──────────────┘              └──────────────┘               └──────────────┘
  1. Claude Code calls an MCP tool (e.g., kimi_analyze)
  2. This server spawns the kimi CLI with the prompt and codebase path
  3. Kimi autonomously reads files, analyzes the code (up to 256K tokens)
  4. The result is parsed from Kimi's JSON output and returned to Claude Code
  5. Claude acts on the structured results — edits, plans, or further analysis

CLI Invocation Reference

The MCP server calls the Kimi CLI in non-interactive (print) mode:

kimi --work-dir <path> --print -p "<prompt>"
Flag Purpose
--print Non-interactive mode — outputs result and exits (required for subprocess use)
-p / --prompt Pass prompt directly (bypasses interactive shell)
--work-dir / -w Set codebase root directory
-S <id> Resume a specific session by ID
--no-thinking Disable thinking mode

[!NOTE] There is no kimi analyze subcommand. The MCP tool is named kimi_analyze, but the underlying CLI uses the flags above. Use this syntax to call Kimi directly for debugging or scripting.

Advanced Setup

For development (auto-recompile on changes):

{
  "mcpServers": {
    "kimi-code": {
      "command": "npx",
      "args": ["tsx", "/absolute/path/to/kimi-code-mcp/src/index.ts"]
    }
  }
}

npm

Published as kimi-mcp-server on npm.

npx kimi-mcp-server          # run directly
npm install -g kimi-mcp-server # install globally

Project Structure

src/
├── index.ts           # MCP server setup, tool definitions
├── kimi-runner.ts     # Spawns kimi CLI, parses output, handles timeouts
└── session-reader.ts  # Reads Kimi session metadata from ~/.kimi/

Contributing

See CONTRIBUTING.md for guidelines.

Changelog

See CHANGELOG.md for version history.

License

MIT

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选