hermes-knowledge-ingestion

hermes-knowledge-ingestion

Enables saving URLs, local files, plain text, and platform content (WeChat, X, YouTube, PDF, images, audio/video, Telegram) as structured Markdown knowledge cards in an Obsidian vault, with optional AI-powered classification and tagging.

Category
访问服务器

README

Hermes Knowledge Ingestion

Local-first, plugin-based knowledge ingestion for Hermes, Obsidian, and other local retrieval tools. It turns links, text, videos, screenshots, PDFs, and local files into structured Markdown knowledge cards and saves them to an Obsidian Vault.

Current release: 0.2.0. Web and file ingestion run cross-platform. WeChat Channels downloading is an optional, experimental macOS integration that requires the desktop WeChat client and a local TLS proxy.

How it works

Hermes Skill / CLI / Web / Telegram
                  |
              MCP / FastAPI
                  |
             Source Adapter
                  |
             ContentItem
                  |
       Cleaner / OCR / Whisper / AI
                  |
      Classifier / Tags / Knowledge Linker
                  |
          Obsidian Markdown + SQLite

Every source is normalized into a ContentItem. To add a platform, implement SourceAdapter.detect() and SourceAdapter.fetch(), then register the adapter in backend/adapters/registry.py.

Supported sources

Source Input Capabilities
Web pages, blogs, and news URL Readability extraction, Markdown conversion, and image download
WeChat Official Accounts URL Article body, author, and images; can also be synced by another tool
X / Twitter Post URL Current post, visible parent context, quoted content, and media when available
YouTube URL Captions first; Whisper fallback when captions are unavailable
Podcast RSS and Apple Podcasts Feed or episode URL Episode metadata, Podcasting 2.0 transcript, audio download, and Whisper fallback
Vimeo URL oEmbed metadata, captions when available, and Whisper fallback
Direct audio, video, and HLS Media URL Streaming download for common media files; yt-dlp resolution for .m3u8
PDF File Text extraction; OCR for scanned pages with the media extra
Images File OCR plus visual and chart descriptions when a vision model is configured
Audio and video File Whisper transcription or vision-model understanding
WeChat Channels Share URL Experimental macOS integration, or upload the original video directly
Telegram Webhook Text, captions, or the first URL found in a message

Quick start

Python 3.11 or newer is required. Media processing requires ffmpeg; OCR requires Tesseract.

python3 -m venv .venv
source .venv/bin/activate
pip install -e ".[media,browser,hermes,dev]"
playwright install chromium
cp config.example.yaml config.yaml
uvicorn backend.main:app --host 127.0.0.1 --port 8787

Open http://127.0.0.1:8787, or use the CLI:

.venv/bin/python scripts/ingest.py 'https://example.com/article'
.venv/bin/python scripts/ingest.py 'https://feeds.example.com/show.rss'
.venv/bin/python scripts/ingest.py 'https://vimeo.com/123456'
.venv/bin/python scripts/ingest.py '/absolute/path/file.pdf'
.venv/bin/python scripts/ingest.py 'A note to keep' --title 'Quick note'

The same pipeline is available through the API:

curl -X POST http://127.0.0.1:8787/api/ingest \
  -H 'content-type: application/json' \
  -d '{"url":"https://example.com/article"}'

Configure AI and Obsidian

The AI layer uses an OpenAI-compatible Chat Completions endpoint. AI is disabled by default; without a model the system still creates a local fallback summary. Enable AI for classification, visual understanding, and richer tags:

export OBSIDIAN_VAULT_DIR=/absolute/path/to/ObsidianVault
export AI_ENABLED=true
export OPENAI_BASE_URL=http://127.0.0.1:11434/v1
export OPENAI_API_KEY=''
export OPENAI_MODEL=qwen2.5:7b
export OPENAI_VISION_MODEL=your-vision-model

You can set the same values in config.yaml. The config file, .env, database, browser login state, and downloaded media are ignored by Git.

Knowledge linking uses qmd when available. If qmd is not installed, it falls back to lexical matching over the latest 1,000 Markdown notes in the Vault. Cards are written to a temporary file and atomically replaced so an indexer never sees a partial note.

Hermes MCP tool

scripts/knowledge_mcp.py exposes two tools:

  • knowledge_ingest: ingest a URL, local file, or text and wait for the knowledge card to finish.
  • knowledge_wechat_prepare: refresh the local WeChat Channels window only when the client connection needs recovery.

In Hermes, configure the MCP command to use this repository's virtual-environment Python and the absolute path to scripts/knowledge_mcp.py. Copy or symlink hermes-skill/personal-knowledge-ingestion into the Hermes skills directory. Hermes can then route intents such as “save”, “archive”, and “ingest” to this tool.

Docker

cp .env.example .env
docker compose up --build

Docker is suitable for web pages, files, OCR, transcription, and the AI pipeline. When the workflow needs the macOS WeChat client, system proxy, or a GUI browser login, run the backend directly on the host. Compose binds the service to 127.0.0.1:8787.

WeChat Channels security boundary

The Channels integration uses the separately maintained ltaoo/wx_channels_download project. Its license and security boundary are separate from this repository. This project does not distribute its binary, root certificate, cookies, or WeChat login data. See integrations/wechat-channels/README.md for installation, licensing, proxy, and macOS permission details.

The downloader creates a local TLS proxy. Use only a trusted, checksum-verified build and never expose the downloader or this service to a LAN. The MCP tool temporarily switches the HTTP/HTTPS proxy for the task and restores the previous settings afterward. The original video is deleted only after both the Obsidian note and SQLite record have been written successfully.

Verification

pytest -q
ruff check backend scripts tests

The test suite covers Markdown formatting, task recovery, text end-to-end ingestion, X context, the WeChat Channels adapter, video transcoding, post-write cleanup, and input classification. Real platform pages and login sessions change over time, so production deployments should still perform a separate end-to-end check for each platform they use.

Contributing and license

Read CONTRIBUTING.md and SECURITY.md before submitting a change. Original project code is licensed under Apache-2.0. Optional third-party components remain under their own licenses; see THIRD_PARTY_NOTICES.md.

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选