personalknowhow

personalknowhow

Provides semantic search and listing over a unified personal knowledge graph aggregating LinkedIn, GitHub, course completions, and more, enabling MCP clients to answer questions about skills and experience with evidence-backed results.

Category
访问服务器

README

PKH — PersonalKnowHow

Turns a scattered personal learning/work history — LinkedIn, GitHub, course platforms, Gmail completion emails, sibling project repos — into a unified knowledge graph, served over two real, deployed MCP servers so any MCP-aware client (Claude Desktop, etc.) can query it with semantic search instead of keyword matching.

Try the live demo

https://personalknowhow-demo.kxtwrdzt6g.workers.dev/mcp is a real, deployed MCP server — but it's not a webpage. Opening that URL in a browser sends a plain GET, and MCP servers only speak POST with JSON-RPC framing, so you'll just see a bare {"error":{"message":"Method not allowed."}}. That's expected, not broken — it means you're looking at it the wrong way.

The actual way to use it is as an MCP connector. In Claude Desktop, edit claude_desktop_config.json (config file location):

{
  "mcpServers": {
    "personalknowhow-demo": {
      "command": "npx",
      "args": ["mcp-remote", "https://personalknowhow-demo.kxtwrdzt6g.workers.dev/mcp"]
    }
  }
}

Restart Claude Desktop, then ask something like "use personalknowhow-demo to check if I have Django experience" — Claude calls the query_knowhow tool over MCP and gets back semantically-matched evidence (courses, projects, certifications) with similarity scores, no auth required.

If you just want to confirm the server is alive without setting up a client:

curl -s https://personalknowhow-demo.kxtwrdzt6g.workers.dev/mcp \
  -X POST -H "Content-Type: application/json" -H "Accept: application/json, text/event-stream" \
  -d '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2024-11-05","capabilities":{},"clientInfo":{"name":"test","version":"1.0"}}}'

A 200 with a JSON-RPC response back confirms it's live — the Accept header above is required; without it the server correctly returns 406 Not Acceptable, which is a different, also-expected error from the browser-GET one above.


Why this exists

Course-completion lists and keyword-matched resumes are a weak signal of what someone actually knows. This project builds a real knowledge graph from primary sources (not self-reported summaries), embeds every entry with a real embedding model, and exposes it as a queryable MCP tool — so "do I have Django experience?" gets answered by walking real evidence (a project's README, a course completion, an endorsement) with a similarity score attached, not a guess.

Architecture

ingest/            Source-specific scripts → common schema
                    {title, type, provider, date, description, domain_tags}
corpus/             Generated markdown, one subfolder per source (not tracked — see Privacy below)
graph.json          Extracted nodes/edges from corpus/ (not tracked)
mcp/                Public MCP server (Cloudflare Worker) — semantic search, no auth
mcp-private/        Private MCP server — same search, bearer-token gated, adds
                    signal-only evidence (job applications, career interests)

Ingestion sources: LinkedIn (via the Member Data Portability API, EU-only — see docs/linkedin-connector-notes.md for notes on the manual-export alternative for other regions), GitHub (via the gh CLI, excluding forks — a fork is evidence of browsing, not building), DataCamp/edX/Skilljar course completions, Gmail (completion emails from other platforms), and sibling project repos (auto-discovered, evidenced via README + tracked filenames + a keyword pass, not self-reported).

ingest/merge.py deduplicates across sources (idempotent — safe to re-run). ingest/build_graph.py extracts nodes/edges from corpus/ frontmatter into graph.json.

The RAG layer

Both mcp/ and mcp-private/ are stateless Cloudflare Workers (createMcpHandler, no Durable Object) that embed every corpus entry with Workers AI (@cf/baai/bge-base-en-v1.5, 768-dim) at export time, and embed the query string at request time, then rank by cosine similarity. Two MCP tools are exposed: query_knowhow(topic) for semantic search, and list_by_type(type) for a plain listing. The private server additionally tags every result with an evidence_tier (demonstrated vs. signal_only), so a job application or career-interest entry can never be mistaken for proof of a skill.

Privacy design

corpus/ and graph.json are never public — no public-facing code reads them directly. The only sanctioned public data source is mcp/public_entries.json, built by ingest/build_public_export.py via a fail-closed allowlist: only explicitly listed corpus categories (courses, projects, certifications, education, endorsements, positions, profile, recommendations, articles) get exported. A new corpus category is excluded by default until someone deliberately adds it to the allowlist — the same discipline that keeps job applications and career-interest data out of the public server entirely; that data only exists in mcp-private/, gated behind a bearer token, and is never committed to this repo either (see .gitignore).

Career-agent tooling

A second layer built on top of the same corpus: ingest/analyze_job_postings.py scores scraped job postings against known skill coverage using the same embeddings (graded known/peripheral similarity, not binary keyword matching), ingest/cv_tailor.py matches a posting's requirements against CV bullets with an explicit two-tier system (exact-term matches vs. semantically-related matches, the latter always labeled "verify before claiming" rather than asserted), and ingest/recommend_courses.py cross-references course catalogs against coverage gaps.

Running locally

pip install -r requirements.txt

# Deduplicate corpus after any ingest run
python ingest/merge.py

# Build graph.json from corpus/
python ingest/build_graph.py

# Run a specific ingest source, e.g.:
python ingest/github_ingest.py
python ingest/linkedin_api_ingest.py --domains PROFILE,POSITIONS,SKILLS

Each ingest/*_ingest.py script is independent — run whichever sources apply to you. All of them write markdown into corpus/<source>/ using the shared schema below.

Corpus schema

Every markdown file in corpus/ uses this YAML frontmatter:

---
title: "Advanced Python Programming"
type: "course"              # course | certification | position | project | education | ...
provider: "LinkedIn Learning"
date: "2024-01-15"
description: "Free-text summary."
domain_tags:
  - python
  - programming
---

Deploying the MCP servers

cd mcp && npm install && npm run deploy        # public server
cd mcp-private && npm install && npm run deploy # private server
cd mcp-private && npm run secret                # set PRIVATE_MCP_TOKEN

Both need a Cloudflare account with Workers AI access ([ai] binding, remote = true in wrangler.toml). Rebuilding the embeddings after a corpus change:

CLOUDFLARE_ACCOUNT_ID=... CLOUDFLARE_AI_TOKEN=... python ingest/build_public_export.py
CLOUDFLARE_ACCOUNT_ID=... CLOUDFLARE_AI_TOKEN=... python ingest/build_private_export.py

Adding a new ingest source

  1. Create ingest/<source>_ingest.py that reads the raw export and writes markdown files to corpus/<source>/ using the schema above.
  2. merge.py and build_graph.py require no changes — they scan corpus/ generically.
  3. If the new category should ever be public, add it deliberately to ALLOWLIST in ingest/build_public_export.py — it's excluded by default otherwise.

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选