MCP 服务器

LearnMCP Server

Extracts and summarizes learning content from YouTube videos, PDFs, and web articles to provide context for project-based learning. It features automated background processing and integrates with Forest's HTA builder for informed task generation.

README

LearnMCP Server

A standalone MCP server that enhances Forest with learning content extraction and summarization capabilities.

Overview

LearnMCP extracts and summarizes learning content from various sources (YouTube videos, PDFs, web articles) and makes those summaries available to Forest's HTA builder for more informed task generation.

Features

Content Extraction: YouTube videos (with transcripts), PDF documents, web articles
Background Processing: Async content processing with queue management
Smart Summarization: Content chunking and summarization with relevance scoring
Forest Integration: Optional integration with Forest's HTA tree builder
Standalone Operation: Can be enabled/disabled independently of Forest

Architecture

User → LearnMCP Tools → LearnService → BackgroundProcessor ⇄ Extractors ⇄ Summarizer → DataPersistence
                                                                                              ↓
                                                                                    <DATA_DIR>/learn-content/
                                                                                              ↓
                                                                              Forest HTA Builder (optional)

Installation

Install Dependencies:
```
cd learn-mcp-server
npm install
```

Configure MCP: Add to your mcp-config.json:

{
  "mcpServers": {
    "learn-mcp": {
      "command": "node",
      "args": ["server.js"],
      "cwd": "learn-mcp-server",
      "env": {
        "FOREST_DATA_DIR": "<same as Forest>"
      }
    }
  }
}

Start Server: The server starts automatically when Claude Desktop loads the MCP config.

Available Tools

`add_learning_sources`

Add learning sources (URLs) to a project for content extraction.

Parameters:

project_id (string): Project ID to add sources to
urls (array): Array of URLs (YouTube, PDF, articles)

Example:

{
  "project_id": "my_project",
  "urls": [
    "https://youtube.com/watch?v=example",
    "https://example.com/document.pdf",
    "https://blog.example.com/article"
  ]
}

`process_learning_sources`

Start background processing of pending learning sources.

Parameters:

project_id (string): Project ID to process sources for

`list_learning_sources`

List learning sources for a project, optionally filtered by status.

Parameters:

project_id (string): Project ID
status (string, optional): Filter by status (pending, processing, completed, failed)

`get_learning_summary`

Get learning content summary for a project or specific source.

Parameters:

project_id (string): Project ID
source_id (string, optional): Specific source ID (if not provided, returns aggregated summary)
token_limit (number, optional): Maximum tokens for aggregated summary (default: 2000)

`delete_learning_sources`

Delete learning sources and their summaries.

Parameters:

project_id (string): Project ID
source_ids (array): Array of source IDs to delete

`get_processing_status`

Get current processing status for learning sources.

Parameters:

project_id (string): Project ID

Supported Content Types

YouTube Videos

Extracts video metadata (title, author, duration, etc.)
Downloads transcripts when available
Falls back to description if no transcript

PDF Documents

Extracts text content from remote PDF URLs
Preserves document metadata
Handles various PDF formats

Web Articles

Uses Mozilla Readability for clean content extraction
Extracts metadata (title, author, publish date, etc.)
Estimates reading time

Data Storage

LearnMCP stores data in <FOREST_DATA_DIR>/learn-content/:

learn-content/
├── <project_id>/
│   ├── sources.json          # Source registry
│   └── summaries/
│       ├── <source_id>.json  # Individual summaries
│       └── ...

Forest Integration

When both LearnMCP and Forest are active, Forest's HTA builder can optionally include learning content summaries in its task generation prompts. This happens automatically when:

LearnMCP has processed learning sources for a project
Forest builds an HTA tree for the same project
Learning content summaries are injected into the HTA generation prompt

Workflow Examples

Basic Learning Content Workflow

Add Sources:

add_learning_sources(project_id="learn_python", urls=["https://youtube.com/watch?v=python_tutorial"])

Process Content:

process_learning_sources(project_id="learn_python")

Check Status:

get_processing_status(project_id="learn_python")

Get Summary:

get_learning_summary(project_id="learn_python")

Integrated with Forest

Add and process learning sources in LearnMCP
Build HTA tree in Forest - it will automatically include learning content context
Generated tasks will be informed by the processed learning materials