Digital Persona Memory MCP Server
Builds a searchable memory bank from persona-specific PDF content and exposes it as an MCP tool for querying.
README
Digital Persona Memory App
Builds a searchable memory bank from PDF content and exposes it as an MCP tool.
Add it as an MCP server in your favorite AI client (eg: LM Studio) with the included system prompt and chat with your heroes, historical figures, or any collection of PDF content.
Example personas are included for Karl Marx and Friedrich Engels. The personas are built from every work both figures ever published. You can chat, ask them questions, or even debate them as though they were still alive and they'll respond based on their own publicly available works.
Alternative uses include adding your school textbooks in PDF format and asking them questions. No more finding that one page, just ask it to explain whatever you want to know.
A modified version of this app is used to bring Facebook users to life as Digital Personas using their exported Facebook data.
What This App Does
- extracts and indexes persona content from PDF-derived JSON sections
- builds per-person Chroma memory databases
- serves an MCP search tool with semantic ranking plus recency scoring
- supports multiple personas by adding folders under
pdf/
Core Features
parser/contains PDF parsing tools to turn PDF persona text into JSON section files.2_build_memory.pyembeds documents and builds a Chroma memory DB.3_avatar_mcp.pyexposessearch_my_memory()as an MCP tool.--namelets you create and query separate memory databases for different personas.config.jsoncentralizes paths, embedding server settings, memory DB defaults, and search tuning.
Project Files
1_prep_data.py: Prepare imported JSON data and normalize text for memory import.2_build_memory.py: Build the Chroma memory database with optional persona selection.3_avatar_mcp.py: Run the MCP server and query the selected memory database.config.json: Configuration for PDF sources, embeddings, memory DB paths, and search settings.parser/: PDF downloader and parser utilities for creating JSON section files from PDFs.
Python Requirements
Install dependencies in your virtual environment:
pip install -r requirements.txt
Data Layout
PDF Persona Support
Each persona owns a directory under pdf/.
For example:
pdf/marx/(Karl Marx persona, included as a sample)pdf/engles/(Friedrich Engels persona, included as a sample)
Memory DB Layout
The memory database lives under the base path in config.json.
When using --name, a separate collection directory is created:
avatar_memory_db/marx/avatar_memory_db/engles/
Each persona DB contains Chroma persistence files plus a document_index.json.
LM Studio Setup
Before building or querying memory:
- Start LM Studio's Local Server.
- Load the embedding model configured in
config.json. - Confirm
embeddings.api_basematches the Local Server URL.
End-to-End Workflow
1) Add persona-specific PDF data or formatted JSON sections
Option 1: Add PDF files in pdf/<persona>/ and run the parser tools in parser/ to create JSON section files in pdf/<persona>/json/.
You will need to edit the parser tools to meet your specific PDF structure and content. The included parser/ tools are a starting point for common PDF layouts.
Option 2: Add JSON section files directly in pdf/<persona>/json/.
The JSON files should be structured as:
{
"sections": [
{
"title": "Section Title",
"text": "Section text content...",
"timestamp": "2024-01-01T00:00:00Z" // Timestamp is optional but recommended for recency scoring
},
...
]
}
2) Prepare JSON data for ingestion
python 1_prep_data.py
This will normalize the text and prepare the JSON sections for embedding. It will prompt you to select a persona if multiple are present.
3) Build the memory DB
Option 1: Build default collection (no persona name specified. Will use config.json default collection name):
python 2_build_memory.py
Option 2: Build a named persona collection:
JSON files for Karl Marx and Friedrich Engels are included as examples.
python 2_build_memory.py --name marx
python 2_build_memory.py --name engles
4) Start the MCP server
Option 1: Run the default DB server:
python 3_avatar_mcp.py
Option 2: Run a named persona DB server:
python 3_avatar_mcp.py --name marx
python 3_avatar_mcp.py --name engles
The exposed MCP tool is:
search_my_memory(topic, num_results=...)
Search Behavior
Results are ranked by:
- semantic similarity from embeddings
- an optional recency boost
Tune behavior in config.json:
search.default_num_resultssearch.candidate_multipliersearch.recency_half_life_dayssearch.recency_weight_alpha
Troubleshooting
- No memory results:
- Confirm
pdf/<persona>/json/files exist. - Confirm prepared_documents.json exists in the persona folder (or whatever filename you specified in
config.json). - Rebuild the memory DB.
- Confirm
- Collection or DB errors:
- Verify
memory_db.pathandcollection_nameinconfig.json. - Use
--nameconsistently for build and server.
- Verify
- LM Studio errors:
- Confirm the Local Server is running with embedding model loaded.
- Check
embeddings.api_base, and ensure the model name inconfig.jsonmatch LM Studio's model name.
- Long build times:
- Large persona data can take 5-10+ minutes depending on hardware and embedding throughput.
Screenshot of this in action
<img width="1408" height="751" alt="image" src="https://github.com/user-attachments/assets/c374217b-79eb-4d63-8d56-7f8396264fd8" />
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。