Ask Samin
Enables users to search Samin Yasar's video library and receive full-video recommendations grounded in transcript evidence with links to exact matched caption cues.
README
Ask Samin
Ask Samin is a hostable, citation-first knowledge app for Samin Yasar's video library. Members start with their goal, stage, tools, and blocker, then get full-video recommendations grounded in transcript evidence with links to the exact matched caption cue.
The same retrieval layer is exposed as a remote MCP server at /mcp, so members can use it inside ChatGPT or Claude without handing this application their AI account credentials. The standalone site intentionally returns deterministic source discovery only; the signed-in MCP host owns model inference for the conversational experience.
What is included
- Intake-first standalone source navigator with verified labels and server-built YouTube timestamp links.
- Searchable video and Shorts library with honest transcript-coverage states; Shorts are browse-only.
- ChatGPT- and Claude-compatible MCP tools for full-video search and exact matched-source lookup.
- Private-by-default admin ingestion preview for YouTube metadata, community calls, documents, and web resources.
- Supabase Postgres schema using pgvector, Postgres full-text search, HNSW, GIN, and reciprocal-rank fusion.
- Static MiniSearch fallback for local development and pre-database deployments.
- Public
/promptsledger sourced from the exact prompts the app runs. - Resumable creator-export ingestion with a shared lock, manifest, bounded workers, and three-attempt failure ceiling.
- Loopy workflow in
LOOPS.md.
Architecture
flowchart LR
A["Member in web app"] --> C["Next.js chat route"]
B["Member in ChatGPT or Claude"] --> M["Remote MCP /mcp"]
C --> R["Grounded retrieval"]
M --> R
R --> P["Supabase pgvector + FTS/RRF"]
R --> F["Static catalog fallback"]
R --> S["Exact source + timestamp links"]
S --> B
D["Admin text ingestion"] --> Q["Atomic private-by-default write"] --> P
See docs/ARCHITECTURE.md for the design and auth decision, and docs/INGESTION.md for the best place to add more videos, community calls, and documents.
Recommendation boundary and token flow
The first recommendation turn is an intake gate: goal, current stage, tools, and blocker. Recommendation APIs then return only full videos backed by timed transcript cues. Shorts and metadata-only entries remain visible in Library browse but can never enter recommendation results.
The standalone site runs no model and consumes no model-inference tokens. When /mcp is connected, ChatGPT or Claude remains the host: it owns sign-in, model selection, and account usage. Ask Samin receives the tool query and returns read-only transcript evidence; it receives neither the member's password nor a personal API key.
Run locally
Requirements: Node.js 20 or newer and Python 3.10 or newer for local transcript imports.
npm install
cp .env.example .env.local
npm run dev
Open http://localhost:3000. The app is useful without secrets: source search and the MCP retrieval endpoint run against the generated catalog.
Optional environment variables:
NEXT_PUBLIC_SUPABASE_URLandSUPABASE_SERVICE_ROLE_KEY: enable persistent hybrid retrieval and ingestion.ADMIN_INGEST_TOKEN: protects the ingestion write route; use at least 32 random bytes.NEXT_PUBLIC_APP_URL: canonical deployment URL used in MCP connection instructions.
The application intentionally does not accept personal model-provider API keys and does not use an owner-funded model key on the public chat route.
Add creator-owned transcripts
Put .vtt, .srt, .json, or .txt exports in imports/. The filename must contain the 11-character YouTube ID. Then run:
npm run ingest:exports
npm run catalog:build
npm run qa
Timed formats preserve exact cues as compact anchors inside each transcript chunk. Plain text becomes searchable but is labeled as untimed. The importer uses five workers by default, retries a failed file no more than three times, and records keep/discard evidence in .cache/ingestion/manifest.json under a file lock.
The current /admin form accepts metadata and pasted creator-owned text, previews normalization, and can atomically persist to Supabase when configured. New records are private unless publication is explicitly selected. File upload, Storage-backed parsing, embedding backfill, and an asynchronous publish worker are future production work; do not run bulk transcription in a Vercel request.
Verify
npm run qa
npm run build
npm run qa runs linting, strict TypeScript, unit tests, and catalog integrity checks. The measurable retrieval QA plan is in docs/QA.md.
Release evidence is recorded in the channel audit and Loopy release receipt.
Deploy
The project is configured for Vercel. Link only the intended showcase project, add environment values there, and deploy:
npx vercel link
npx vercel deploy --prod
No billing, trial, or subscription setup is performed by the repository.
Why there is no third-party “Login with ChatGPT” button
The referenced opencoredev/login-with-chatgpt project currently reuses Codex CLI's OAuth client identity and an undocumented ChatGPT backend. OpenAI's own Codex login warns users to cancel if a website gives them its device code. That is not an appropriate production trust boundary for member credentials.
This build preserves consumer-controlled model use safely: members connect the public /mcp retrieval server to ChatGPT or Claude, where the signed-in client owns authentication, inference, and account usage. The standalone site is retrieval-only. The evidence and exact decision are recorded in docs/AUTH-DECISION.md.
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。