Bio-MCP BLAST

Bio-MCP BLAST

Enables AI assistants to perform NCBI BLAST sequence similarity searches through natural language, supporting nucleotide and protein searches, custom database creation, and multiple output formats.

Category
访问服务器

README

Bio-MCP BLAST

🔍 MCP server for NCBI BLAST sequence similarity search

Enable AI assistants to perform BLAST searches through natural language. Search nucleotide and protein databases, create custom databases, and get formatted results instantly.

🧬 Features

  • blastn - Nucleotide-nucleotide BLAST search
  • blastp - Protein-protein BLAST search
  • makeblastdb - Create custom BLAST databases
  • Multiple output formats - JSON, XML, tabular, pairwise
  • Flexible input - File paths or raw sequences
  • Queue support - Async processing for large searches

🚀 Quick Start

Installation

# Install BLAST+
conda install -c bioconda blast

# Or via package manager
# macOS: brew install blast
# Ubuntu: sudo apt-get install ncbi-blast+

# Install MCP server
git clone https://github.com/bio-mcp/bio-mcp-blast.git
cd bio-mcp-blast
pip install -e .

Basic Usage

# Start the server
python -m src.server

# Or with queue support
python -m src.main --mode queue

Configuration

Add to your MCP client config:

{
  "mcpServers": {
    "bio-blast": {
      "command": "python",
      "args": ["-m", "src.server"],
      "cwd": "/path/to/bio-mcp-blast"
    }
  }
}

💡 Usage Examples

Simple Sequence Search

User: "BLAST this sequence against nr: ATGCGATCGATCG"
AI: [calls blastn] → Returns top hits with E-values and alignments

File-Based Search

User: "Search proteins.fasta against SwissProt database"
AI: [calls blastp] → Processes file and returns similarity results

Database Creation

User: "Create a BLAST database from reference_genomes.fasta"
AI: [calls makeblastdb] → Creates searchable database files

Long-Running Search

User: "BLAST large_dataset.fasta against nt database"
AI: [calls blastn_async] → "Job submitted! ID: abc123, checking progress..."

🛠️ Available Tools

blastn

Nucleotide-nucleotide BLAST search

Parameters:

  • query (required) - Path to FASTA file or sequence string
  • database (required) - Database name (e.g., "nt", "nr") or path
  • evalue - E-value threshold (default: 10)
  • max_hits - Maximum hits to return (default: 50)
  • output_format - Output format: "tabular", "xml", "json", "pairwise"

blastp

Protein-protein BLAST search

Parameters:

  • Same as blastn, but for protein sequences

makeblastdb

Create BLAST database from FASTA file

Parameters:

  • input_file (required) - Path to FASTA file
  • database_name (required) - Name for output database
  • dbtype (required) - "nucl" or "prot"
  • title - Database title (optional)

Async Variants (Queue Mode)

  • blastn_async - Submit nucleotide search to queue
  • blastp_async - Submit protein search to queue
  • get_job_status - Check job progress
  • get_job_result - Retrieve completed results

⚙️ Configuration

Environment Variables

# Basic settings
export BIO_MCP_MAX_FILE_SIZE=100000000    # 100MB max file size
export BIO_MCP_TIMEOUT=300                # 5 minute timeout
export BIO_MCP_BLAST_PATH="blastn"        # BLAST executable path

# Queue mode settings
export BIO_MCP_QUEUE_URL="http://localhost:8000"

Database Setup

# Download common databases
mkdir -p ~/blast-databases
cd ~/blast-databases

# NCBI databases (large downloads!)
update_blastdb.pl --decompress nt
update_blastdb.pl --decompress nr
update_blastdb.pl --decompress swissprot

# Set environment variable
export BLASTDB=~/blast-databases

🐳 Docker Deployment

Local Docker

# Build image
docker build -t bio-mcp-blast .

# Run container
docker run -p 5000:5000 \
  -v ~/blast-databases:/data/blast-db:ro \
  -e BLASTDB=/data/blast-db \
  bio-mcp-blast

Docker Compose

services:
  blast-server:
    build: .
    ports:
      - "5000:5000"
    volumes:
      - ./databases:/data/blast-db:ro
    environment:
      - BLASTDB=/data/blast-db
      - BIO_MCP_TIMEOUT=600

🔄 Queue System

For long-running BLAST searches, use the queue system:

Setup

# Start queue infrastructure
cd ../bio-mcp-queue
./setup-local.sh

# Start BLAST server with queue support
python -m src.main --mode queue --queue-url http://localhost:8000

Usage

# Submit async job
job_info = await blast_server.submit_job(
    job_type="blastn",
    parameters={
        "query": "large_sequences.fasta",
        "database": "nt",
        "evalue": 0.001
    }
)

# Check status
status = await blast_server.get_job_status(job_info["job_id"])

# Get results when complete
results = await blast_server.get_job_result(job_info["job_id"])

📊 Output Formats

Tabular (Default)

# Fields: query_id, subject_id, percent_identity, alignment_length, ...
Query_1    gi|123456    98.5    500    7    0    1    500    1000    1499    1e-180    633

JSON

{
  "BlastOutput2": [{
    "report": {
      "results": {
        "search": {
          "query_title": "Query_1",
          "hits": [...]
        }
      }
    }
  }]
}

XML

Standard BLAST XML format for programmatic parsing.

🧪 Testing

# Run tests
pytest tests/ -v

# Test with real data
python tests/test_integration.py

# Performance testing
python tests/benchmark.py

📈 Performance Tips

Local Optimization

  • Use SSD storage for databases
  • Increase available RAM
  • Use multiple CPU cores: export BLAST_NUM_THREADS=8

Database Selection

  • Use smaller, specific databases when possible
  • Consider pre-filtering sequences
  • Use appropriate E-value thresholds

Queue Optimization

  • Scale workers based on CPU cores
  • Use separate queues for different database sizes
  • Monitor memory usage with large databases

🔐 Security

Input Validation

  • File size limits prevent resource exhaustion
  • Path validation prevents directory traversal
  • Command injection protection

Sandboxing

  • Containers run as non-root user
  • Temporary files isolated per job
  • Network access restricted in production

🐛 Troubleshooting

Common Issues

BLAST not found

# Check installation
which blastn
blastn -version

# Install via conda
conda install -c bioconda blast

Database not found

# Check BLASTDB environment variable
echo $BLASTDB

# List available databases
blastdbcmd -list /path/to/databases

Out of memory

# Reduce max_target_seqs
blastn -max_target_seqs 100

# Use streaming for large outputs
# Increase system swap space

Timeout errors

# Increase timeout
export BIO_MCP_TIMEOUT=3600  # 1 hour

# Or use queue mode for long searches
python -m src.main --mode queue

📚 Resources

🤝 Contributing

  1. Fork the repository
  2. Create a feature branch
  3. Add tests for new functionality
  4. Ensure all tests pass
  5. Submit a pull request

See CONTRIBUTING.md for detailed guidelines.

📄 License

MIT License - see LICENSE file.

🆘 Support


Happy BLASTing! 🧬🔍

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选