kw-notice-mcp

kw-notice-mcp

Local read-only MCP server for Kwangwoon University notices, providing metadata-only crawling and cached notice access through a local STDIO server.

Category
访问服务器

README

kw-notice-mcp

Local, read-only Kwangwoon University notice tooling. The collector is bounded, metadata-minimizing, and exposes only cached results through a local MCP STDIO server.

Install and run

This project uses Python 3.13+ and uv:

uv sync
uv run kw-notice-mcp --help
uv run kw-notice-mcp init-db --db-path data/kw-notice.sqlite3
uv run kw-notice-mcp status --db-path data/kw-notice.sqlite3

Configuration is read from KW_NOTICE_* environment variables. Copy .env.example to .env only if you need local overrides; it contains no credentials or secret-like values. The crawl command always uses the metadata-only operational mode: one direct request to the first 전체 list page. It stores only DUID, canonical category, redacted/capped title, posted/updated dates, department, constructed source URL, collection time, and source status. It never requests detail pages, body text, attachments, images, email addresses, or phone numbers, and it never uses a generic robots bypass.

The command exit codes are stable: 0 success, 10 blocked or budget exhausted, 11 busy because another crawl owns the SQLite lease, 12 invalid configuration, and 13 infrastructure failure.

Commands

uv run kw-notice-mcp init-db --db-path data/kw-notice.sqlite3
uv run kw-notice-mcp crawl --metadata-only --db-path data/kw-notice.sqlite3
uv run kw-notice-mcp status --db-path data/kw-notice.sqlite3
uv run kw-notice-mcp serve --db-path data/kw-notice.sqlite3

crawl --metadata-only makes no /robots.txt request. Its first and only source request on success is https://www.kw.ac.kr/ko/life/notice.jsp?srCategoryId=&mode=list&searchKey=1&searchVal=&tpage=1. This direct-page behavior is the operator-directed policy for the CLI and the scheduled refresh workflow; it is not a claim of permission or authorization. The bounded collector makes one request at a time, with SQLite BEGIN IMMEDIATE locking and stale-run recovery. A blocked run may update only crawl_runs; notices, FTS rows, and revisions remain unchanged. Page 403, 429, 5xx, CAPTCHA/WAF, malformed markup, invalid redirects or targets, timeouts, oversized responses, and budget failures remain blocked. Metadata runs never request details, body text, attachments, images, email addresses, or phone numbers. Logs are JSON records on stderr with run ID, page/detail counters, status, and a safe block reason. Response bodies and personal data are never logged. serve keeps stdout reserved for MCP JSON-RPC and exits when stdin closes.

Current source-policy state

The scheduled workflow uses the same direct page-one operator policy and makes no robots request. The internal FULL collector path retains its strict robots-policy parser for callers that select it explicitly. Tests use local fixtures and fake responses; verification never performs a live crawl.

The GitHub Actions refresh schedule is weekdays, 09:00–18:00 KST, at 15-minute offsets 07,22,37,52 (7,22,37,52 0-8 * * 1-5 UTC). GitHub schedules are best-effort and may start late, so the workflow remains one-concurrent and lease-protected. Actions is the collector runtime; the MCP remains local STDIO. The workflow downloads the prior manifest, database, and checksum, verifies the complete generation before reuse, and initializes only when the notice-protocol assets are truly absent. After a successful metadata-only crawl it uploads immutable generation assets and updates the stable data-latest manifest last. Blocked, busy, invalid, incomplete, checksum, and infrastructure outcomes publish nothing. On 403, 429, or CAPTCHA, the run stops immediately and the operator should cool down before the next selected slot. The older kw-service repositories are architectural precedent only, not an upstream dependency.

Local database operations

The SQLite file contains bounded, redacted metadata only on the CLI path; metadata-only runs keep body NULL and add no body tokens to FTS. Raw HTML, attachments, email addresses, and phone numbers are not stored. Every notice retains a constructed source link for the original page. Keep the database local and restrict it to the operator, for example:

chmod 600 data/kw-notice.sqlite3
sqlite3 data/kw-notice.sqlite3 '.backup data/kw-notice.backup.sqlite3'
sqlite3 data/kw-notice.sqlite3 'PRAGMA integrity_check;'

Restore only while the server and collector are stopped, after checking the backup path and permissions:

cp data/kw-notice.backup.sqlite3 data/kw-notice.sqlite3
chmod 600 data/kw-notice.sqlite3

Scheduling is operator-owned and intentionally out of scope. This repository does not implement a second cron/systemd/Docker scheduler, a remote HTTP server, OAuth, or a public deployment.

Release DB consumption

The durable handoff is the stable GitHub Release tag data-latest, not an Actions artifact. Its single authoritative pointer is the Release body, a strict JSON value naming one immutable, SHA-256-addressed manifest asset. That manifest names the immutable database asset and its checksum asset. The three immutable assets are uploaded and verified before the Release body is edited; there is no mutable pointer asset or pointer-asset --clobber window. GitHub Release does not provide a global multi-asset transaction: a failed or partial asset upload remains unreachable, while a failed later body edit preserves the prior pointer (or is resolved by read-back). An initial body-edit failure leaves no authoritative generation. Consumers keep their current local DB until the strict pointer and complete generation verify.

Resolve the stable Release body first, parse its strict pointer, then download the exact immutable manifest and the two assets it declares into a temporary directory. The repository helper requires all files, verifies the checksum, then revalidates the copied SQLite candidate for integrity, the exact expected schema, and zero retained body content before atomically replacing the local SQLite file:

release_dir=$(mktemp -d)
gh release view data-latest --repo kyowon1108/kw-notice-mcp --json body,assets \
  >"$release_dir/release.json"
jq -r '.body // empty' "$release_dir/release.json" \
  >"$release_dir/release-pointer.json"
uv run python -m kw_notice_mcp.release verify-pointer \
  --pointer "$release_dir/release-pointer.json" \
  >"$release_dir/validated-pointer.json"
manifest_asset=$(jq -r '.manifest_asset' "$release_dir/validated-pointer.json")
gh release download data-latest --repo kyowon1108/kw-notice-mcp \
  --pattern "$manifest_asset" --dir "$release_dir"
uv run python -m kw_notice_mcp.release verify-manifest \
  --manifest "$release_dir/$manifest_asset" >"$release_dir/validated.json"
database_asset=$(jq -r '.database_asset' "$release_dir/validated.json")
checksum_asset=$(jq -r '.checksum_asset' "$release_dir/validated.json")
gh release download data-latest --repo kyowon1108/kw-notice-mcp \
  --pattern "$database_asset" --dir "$release_dir"
gh release download data-latest --repo kyowon1108/kw-notice-mcp \
  --pattern "$checksum_asset" --dir "$release_dir"
uv run python -m kw_notice_mcp.release restore \
  --manifest "$release_dir/$manifest_asset" \
  --assets-dir "$release_dir" \
  --database ./notices.sqlite3
rm -rf "$release_dir"
uv run kw-notice-mcp serve --db-path ./notices.sqlite3

Local producer staging has a separate, precise atomicity boundary: DB, checksum, and manifest are built and verified in a temporary generation directory, then one same-filesystem directory rename exposes generations/<sha256>. A failure before that rename exposes no new manifest or resolvable generation pair. Consumer installation similarly verifies the complete downloaded pair before one filesystem replacement of the local DB. A missing Release or a Release with no notice-protocol assets may initialize safely; an empty/invalid pointer, incomplete pair, checksum mismatch, API failure, or download failure stops the refresh before crawl or publication and leaves the prior Release untouched.

The schedule can be delayed or coalesced by GitHub Actions. The rule-portal search report found no explicit crawling rule in that limited search, but it did not establish permission or settle legal, terms-of-use, privacy, or redistribution questions. Operators should obtain written confirmation before describing automated access as authorized; the collector therefore remains bounded, metadata-only, direct-page, and fail-closed.

Deployment operators acknowledge that these safeguards and the workflow's bounded GitHub token permissions are controls, not a claim of authorization. They are responsible for confirming access and redistribution approval, protecting the local database and Release, and reviewing failed or blocked runs. This responsibility does not disable the user-approved weekday refresh schedule; it governs its operation.

Fixture-only tests and quality checks

No required test contacts the live site. Permissive robots, HTML pages, and transport failures are injected in memory or read from synthetic fixtures:

uv run pytest tests/integration/test_cli.py -q
uv run pytest -q
uv run basedpyright
uv run ruff check
uv run ruff format --check
uv run python scripts/check_no_excuse_rules.py src tests

Generic Hermes STDIO configuration

Hermes can spawn local MCP servers from its mcp_servers configuration. The following JSON is also valid YAML syntax for a generic ~/.hermes/config.yaml entry; replace the project and database paths with operator-owned paths. It contains no token, credential, secret, or Hermes/Discord runtime dependency:

{
  "mcp_servers": {
    "kw-notice": {
      "command": "uv",
      "args": [
        "run",
        "--project",
        "/path/to/kw-notice-mcp",
        "kw-notice-mcp",
        "serve",
        "--db-path",
        "/path/to/kw-notice-mcp/data/kw-notice.sqlite3"
      ]
    }
  }
}

Hermes and Discord remain external consumers. This repository owns only the notice cache and four read-only MCP tools; it does not receive Discord events, hold platform credentials, or route messages.

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选