WSO2 Docs MCP Server
Enables AI assistants to semantically search WSO2 documentation across multiple products using retrieval-augmented generation, with support for local or cloud embeddings.
README
WSO2 Docs MCP Server
"This is an unofficial community project. Not affiliated with or endorsed by WSO2."
A production-ready Model Context Protocol (MCP) server that provides AI assistants (Claude Desktop, Claude Code, Cursor, VS Code) with semantic search over WSO2 documentation via Retrieval-Augmented Generation (RAG).
Under the hood, it uses a blazing-fast dual-ingestion engine:
- GitHub Native: Fetches raw Markdown directly from WSO2's public GitHub repositories via the Git Trees API (avoids web-scraping noise and rate limits)
- Web Crawl Fallback: For products without dedicated GitHub docs repos (like the WSO2 Library)
Architecture
Documentation Sources
| Product | ID | URL |
|---|---|---|
| API Manager | apim |
https://apim.docs.wso2.com |
| Micro Integrator | mi |
https://mi.docs.wso2.com/en/4.4.0 |
| Ballerina Integrator | bi |
https://bi.docs.wso2.com |
| Choreo | choreo |
https://wso2.com/choreo/docs |
| Identity Server | is |
https://is.docs.wso2.com/en/latest |
| Ballerina | ballerina |
https://ballerina.io/learn |
| WSO2 Library | library |
https://wso2.com/library |
Prerequisites
- Node.js ≥ 20
- Docker (for pgvector)
- Embeddings - no API key required by default:
- Ollama (recommended) - runs locally, model auto-downloaded on first run
- If Ollama is not running, the server automatically falls back to HuggingFace ONNX (in-process, also downloads automatically)
- Cloud providers are also supported: OpenAI, Google Gemini, Voyage AI
Quick Start
Choose the setup path that fits your use case:
- Install from npm - simplest, no cloning required
- Clone and build - for development or contributions
Install from npm
Install the package globally to get the wso2-docs-mcp-server, wso2-docs-crawl, and wso2-docs-migrate commands available system-wide:
npm install -g wso2-docs-mcp-server
Prefer no global install? You can use
npx wso2-docs-mcp-server,npx wso2-docs-crawl, andnpx wso2-docs-migratein every step below - just replace the bare command with itsnpxequivalent.
1. Start pgvector
Download the docker-compose.yml and start the database:
curl -O https://raw.githubusercontent.com/iamvirul/wso2-docs-mcp-server/main/docker-compose.yml
docker compose up -d
2. Start Ollama (optional but recommended)
Install Ollama and pull the default embedding model:
ollama pull nomic-embed-text
ollama serve
No Ollama? Skip this step. The server automatically falls back to HuggingFace ONNX - model downloads on first use with no extra setup.
3. Run database migration
DATABASE_URL="postgresql://wso2mcp:wso2mcp@localhost:5432/wso2docs" \
wso2-docs-migrate
Run migration again whenever you change
EMBEDDING_DIMENSIONS(i.e. switch embedding provider). The script detects and handles dimension changes automatically.
4. Index WSO2 documentation
# Index all products (first run downloads the embedding model automatically)
DATABASE_URL="postgresql://wso2mcp:wso2mcp@localhost:5432/wso2docs" \
wso2-docs-crawl
# Index a single product (faster, great for testing)
DATABASE_URL="postgresql://wso2mcp:wso2mcp@localhost:5432/wso2docs" \
wso2-docs-crawl --product ballerina --limit 20
# Force re-index even unchanged pages
DATABASE_URL="postgresql://wso2mcp:wso2mcp@localhost:5432/wso2docs" \
wso2-docs-crawl --force
Available product IDs: apim, mi, bi, choreo, is, ballerina, library
5. Configure your AI client
The MCP server is launched on demand by your AI client - no background process needed.
Claude Desktop - edit ~/Library/Application Support/Claude/claude_desktop_config.json:
{
"mcpServers": {
"wso2-docs": {
"command": "wso2-docs-mcp-server",
"env": {
"DATABASE_URL": "postgresql://wso2mcp:wso2mcp@localhost:5432/wso2docs",
"EMBEDDING_PROVIDER": "ollama"
}
}
}
}
Claude Code - run once in your terminal:
claude mcp add wso2-docs \
--transport stdio \
-e DATABASE_URL="postgresql://wso2mcp:wso2mcp@localhost:5432/wso2docs" \
-e EMBEDDING_PROVIDER="ollama" \
-- wso2-docs-mcp-server
# Verify
claude mcp list
Cursor - create .cursor/mcp.json in your project root:
{
"mcpServers": {
"wso2-docs": {
"command": "wso2-docs-mcp-server",
"env": {
"DATABASE_URL": "postgresql://wso2mcp:wso2mcp@localhost:5432/wso2docs",
"EMBEDDING_PROVIDER": "ollama"
}
}
}
}
VS Code - create .vscode/mcp.json:
{
"servers": {
"wso2-docs": {
"type": "stdio",
"command": "wso2-docs-mcp-server",
"env": {
"DATABASE_URL": "postgresql://wso2mcp:wso2mcp@localhost:5432/wso2docs",
"EMBEDDING_PROVIDER": "ollama"
}
}
}
}
Using
npxinstead of global install? Replace"command": "wso2-docs-mcp-server"with"command": "npx"and add"args": ["-y", "wso2-docs-mcp-server"].
Cloud embedding provider? Add the key to
env, e.g."EMBEDDING_PROVIDER": "openai", "OPENAI_API_KEY": "sk-...".
Clone and build
1. Clone and install
git clone https://github.com/iamvirul/wso2-docs-mcp-server.git
cd wso2-docs-mcp-server
npm install
2. Start Ollama (optional but recommended)
Install Ollama and start it:
ollama serve
No Ollama? Skip this step. The server detects Ollama is not running and automatically falls back to HuggingFace ONNX inference - the model downloads on first use with no extra setup.
3. Configure environment
cp .env.example .env
# Defaults work out of the box with Ollama.
# Only edit if using a cloud provider (OpenAI / Gemini / Voyage).
4. Start pgvector
docker compose up -d
# pgAdmin available at http://localhost:5050 (admin@wso2mcp.local / admin)
5. Run database migration
npm run db:migrate
Note: Run migration again whenever you change
EMBEDDING_DIMENSIONS(i.e. switch embedding provider). The script detects and handles dimension changes automatically.
6. Index documentation
# Index all products
# On first run the embedding model is downloaded automatically (Ollama or HuggingFace)
npm run crawl
# Index a single product (faster, great for testing)
npm run crawl -- --product ballerina --limit 20
# Force re-index even unchanged pages
npm run crawl -- --force
7. Build and start the MCP server
npm run build
npm start
For development (no build step):
npm run dev
8. Configure your AI client
Replace
/ABSOLUTE/PATH/TO/wso2-docs-mcp-serverwith your actual clone path.
Claude Desktop - edit ~/Library/Application Support/Claude/claude_desktop_config.json:
{
"mcpServers": {
"wso2-docs": {
"command": "node",
"args": ["/ABSOLUTE/PATH/TO/wso2-docs-mcp-server/dist/src/index.js"],
"env": {
"DATABASE_URL": "postgresql://wso2mcp:wso2mcp@localhost:5432/wso2docs",
"EMBEDDING_PROVIDER": "ollama"
}
}
}
}
Claude Code:
claude mcp add wso2-docs \
--transport stdio \
-e DATABASE_URL="postgresql://wso2mcp:wso2mcp@localhost:5432/wso2docs" \
-e EMBEDDING_PROVIDER="ollama" \
-- node "/ABSOLUTE/PATH/TO/wso2-docs-mcp-server/dist/src/index.js"
# Verify
claude mcp list
See config-examples/claude_code.sh for a convenience script.
Cursor - create .cursor/mcp.json - see config-examples/cursor_mcp.json.
VS Code - create .vscode/mcp.json - see config-examples/vscode_mcp.json.
MCP Tools
| Tool | Description |
|---|---|
search_wso2_docs |
Semantic search across all products. Optional product and limit filters. |
get_wso2_guide |
Search within a specific product (apim, mi, bi, choreo, is, ballerina, library). |
explain_wso2_concept |
Broad concept search across all products, returns 8 top results. |
list_wso2_products |
Returns all supported products with IDs and base URLs. |
Example response
[
{
"title": "Deploying WSO2 API Manager",
"snippet": "WSO2 API Manager can be deployed in various topologies…",
"source_url": "https://apim.docs.wso2.com/en/latest/install-and-setup/...",
"product": "apim",
"section": "Deployment Patterns",
"score": 0.8712
}
]
Local Embeddings
The default EMBEDDING_PROVIDER=ollama runs entirely on your machine with no API key. The startup sequence is:
Is Ollama running?
├── Yes → Is model present?
│ ├── Yes → Ready (instant)
│ └── No → Pull via Ollama (streamed, runs once)
└── No → Download ONNX model from HuggingFace Hub (~250 MB, cached after first run)
and run inference in-process via @huggingface/transformers
Both paths use nomic-embed-text / Xenova/nomic-embed-text-v1 by default and produce identical 768-dim vectors, so you can switch between them without re-indexing.
Hardware acceleration (HuggingFace ONNX fallback)
When Ollama is not available, the server auto-detects the best compute backend:
| Machine | Detection | ONNX dtype | Batch size | Throughput |
|---|---|---|---|---|
| Apple Silicon (M1/M2/M3/M4) | process.arch === 'arm64' |
q8 INT8 |
32 | ~9 ms/chunk |
| NVIDIA GPU | nvidia-smi probe |
fp32 |
64 | GPU-dependent |
| All others | fallback | q8 INT8 |
16 | ~10 ms/chunk |
Why q8 on Apple Silicon instead of CoreML/Metal?
CoreML compiles Metal shaders on first use (~20 min cold-start). For the typical chunk sizes produced by this server (6–20 chunks per page), the CPU↔GPU transfer overhead eliminates any inference gain. INT8 quantized inference on ARM NEON SIMD is consistently ~100× faster than fp32 CPU with zero cold-start cost.
Benchmark (Apple M-chip, Xenova/nomic-embed-text-v1):
fp32 CPU (before): ~1,000 ms/chunk (68 chunks ≈ 68 s of embedding)
q8 ARM NEON: ~9 ms/chunk (68 chunks ≈ 0.6 s of embedding) ← ~100× speedup
Note: For small crawls (≤ 10 pages) total wall-clock time is dominated by network I/O (HTTPS fetches to docs sites), so the end-to-end improvement is modest. The embedding speedup becomes significant at scale - crawling 500+ pages where embedding previously accounted for hours of runtime. For best crawl performance, run Ollama (
ollama serve) which parallelises inference natively and has no per-chunk overhead.
Environment Variables
Core
| Variable | Default | Description |
|---|---|---|
DATABASE_URL |
- | PostgreSQL connection string (required) |
EMBEDDING_PROVIDER |
ollama |
ollama | openai | gemini | voyage |
EMBEDDING_DIMENSIONS |
768 |
Must match model output dimensions |
CRAWL_CONCURRENCY |
5 |
Concurrent HTTP requests during crawl |
CHUNK_SIZE |
800 |
Approximate tokens per chunk |
CHUNK_OVERLAP |
100 |
Overlap tokens between chunks |
CACHE_TTL_SECONDS |
3600 |
In-memory query cache TTL |
TOP_K_RESULTS |
10 |
Default search result count |
Ollama (default)
| Variable | Default | Description |
|---|---|---|
OLLAMA_BASE_URL |
http://localhost:11434 |
Ollama server URL |
OLLAMA_EMBEDDING_MODEL |
nomic-embed-text |
Model pulled and used via Ollama |
HUGGINGFACE_EMBEDDING_MODEL |
Xenova/nomic-embed-text-v1 |
ONNX fallback when Ollama is not running |
Cloud providers
| Variable | Default | Description |
|---|---|---|
OPENAI_API_KEY |
- | Required if EMBEDDING_PROVIDER=openai |
OPENAI_EMBEDDING_MODEL |
text-embedding-3-small |
OpenAI model |
GEMINI_API_KEY |
- | Required if EMBEDDING_PROVIDER=gemini |
GEMINI_EMBEDDING_MODEL |
text-embedding-004 |
Gemini model |
VOYAGE_API_KEY |
- | Required if EMBEDDING_PROVIDER=voyage |
VOYAGE_EMBEDDING_MODEL |
voyage-3 |
Voyage model |
Embedding dimension reference
| Provider | Model | Dimensions |
|---|---|---|
| Ollama / HuggingFace | nomic-embed-text / Xenova/nomic-embed-text-v1 |
768 (default) |
| Ollama / HuggingFace | mxbai-embed-large / Xenova/mxbai-embed-large-v1 |
1024 |
| Ollama / HuggingFace | all-minilm / Xenova/all-MiniLM-L6-v2 |
384 |
| OpenAI | text-embedding-3-small |
1536 |
| OpenAI | text-embedding-3-large |
3072 |
| Gemini | text-embedding-004 |
768 |
| Voyage | voyage-3 |
1024 |
| Voyage | voyage-3-lite |
512 |
Scheduled Re-indexing
# Run a one-off re-index (checks hashes, skips unchanged pages)
npm run reindex
# Or from the project directory using node-cron (runs daily at 2 AM)
DATABASE_URL=... node -e "
const { ReindexJob } = require('./dist/jobs/reindexDocs');
const job = new ReindexJob();
job.initialize().then(() => job.scheduleDaily());
"
Project Structure
src/
config/ env.ts · constants.ts
vectorstore/ pgvector.ts · schema.sql
ingestion/ crawler.ts · parser.ts · githubFetcher.ts · markdownParser.ts · chunker.ts · embedder.ts
server/ mcpServer.ts · toolRegistry.ts
jobs/ reindexDocs.ts
index.ts
scripts/
crawl.ts CLI ingestion pipeline
migrate.ts Dynamic schema migration
config-examples/ claude_desktop.json · claude_code.sh · cursor_mcp.json · vscode_mcp.json
docker-compose.yml
.env.example
Development
# Type-check
npx tsc --noEmit
# Run crawl with tsx (no build needed)
npm run crawl -- --product ballerina --limit 5
# Run server in dev mode
npm run dev
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。