Claude Advanced Memory Engine
Provides structured local memory for Claude Desktop, reducing token usage by over 95% through intelligent retrieval and layered storage.
README
Claude Advanced Memory Engine
A production-grade MCP (Model Context Protocol) server that maintains structured local memory for Claude Desktop, reducing token usage by 95%+ without losing important context.
Problem: Traditional conversation summarization replays massive chat histories, wasting 80%+ of token budget.
Solution: Build structured memory once, retrieve intelligently. Never replay history.
Key Features
✅ 95%+ Token Reduction - Structured facts instead of full conversation replay ✅ Local First - Everything runs locally, no cloud, no telemetry, no data leaving your computer ✅ 4-Layer Memory - Hierarchical storage (L1: current, L2: working, L3: semantic, L4: archive) ✅ Intelligent Retrieval - Multi-factor ranking: keyword, semantic, recency, importance, frequency, relationships ✅ Auto Extraction - Automatically extracts facts, decisions, tasks, preferences from responses ✅ Deduplication - Merges equivalent facts, keeps only newest truth ✅ Smart Cleanup - Automatic garbage collection, archival, promotion to hot layers ✅ Context Budgeting - Configurable token limits (minimal/balanced/comprehensive) ✅ Performance - <20ms lookup, <50ms prompt assembly, scales to millions of records
Architecture
Claude Desktop Instance
↓ (MCP Protocol)
┌─────────────────────────────────────────┐
│ Memory Engine MCP Server │
│ │
│ ┌────────────────────────────────────┐ │
│ │ MCP Tool Layer │ │
│ │ (save, search, retrieve, optimize) │ │
│ └──────────────┬─────────────────────┘ │
│ ↓ │
│ ┌────────────────────────────────────┐ │
│ │ Prompt Builder & Optimizer │ │
│ └──────────────┬─────────────────────┘ │
│ ↓ │
│ ┌────────────────────────────────────┐ │
│ │ Intelligent Retrieval Engine │ │
│ │ (multi-factor ranking) │ │
│ └──────────────┬─────────────────────┘ │
│ ↓ │
│ ┌────────────────────────────────────┐ │
│ │ Memory Layer System (L1-L4) │ │
│ └──────────────┬─────────────────────┘ │
│ ↓ │
│ ┌────────────────────────────────────┐ │
│ │ SQLite Database (~/.cache/claude) │ │
│ └────────────────────────────────────┘ │
└─────────────────────────────────────────┘
Memory Layers
| Layer | Scope | TTL | Retrieval | Use Case |
|---|---|---|---|---|
| L1 | Current conversation | Session | Always | Active request context |
| L2 | Working memory | 24h | High | Recent tasks, active decisions |
| L3 | Semantic memory | 30d | Medium | Long-term facts, patterns |
| L4 | Archive | ∞ | Explicit | Historical, rarely accessed |
Installation
Prerequisites
- Node.js 18+
- Claude Desktop (latest)
- ~50MB disk space
Setup
# Clone repository
git clone https://github.com/yourusername/claude-memory-engine.git
cd claude-memory-engine
# Install dependencies
npm install
# Build TypeScript
npm run build
# Initialize database
npm run init-db
# Start MCP server
npm start
Configure Claude Desktop
Add to claude_desktop_config.json:
{
"mcpServers": {
"claude-memory": {
"command": "node",
"args": ["/path/to/claude-memory-engine/dist/index.js"],
"env": {
"CLAUDE_MEMORY_BUDGET": "2000",
"LOG_LEVEL": "info"
}
}
}
}
Usage
Automatic Memory Capture
After every Claude response, the system automatically:
- Extracts facts, decisions, tasks, preferences
- Identifies entities and relationships
- Deduplicates and merges equivalent records
- Stores in appropriate memory layer
- Promotes/archives based on usage
Manual Tools
/save_memory - Save facts/decisions/tasks explicitly
/search_memory - Search by keyword, entity, type
/retrieve_context - Get optimized context for request
/update_memory - Modify existing record
/delete_memory - Remove record
/optimize_prompt - Compress prompt within budget
/memory_stats - View memory metrics and recommendations
/explain_context - Show why a record was retrieved
/export_memory - Export to JSON/CSV/Markdown
/import_memory - Restore from export
Configuration
Environment Variables
# Database location
CLAUDE_MEMORY_DB=~/.cache/claude-memory/memory.db
# Token budget (1000-∞)
CLAUDE_MEMORY_BUDGET=2000
# Enable compression
CLAUDE_MEMORY_COMPRESSION=true
# Enable deduplication
CLAUDE_MEMORY_DEDUP=true
# Enable embeddings (requires model)
CLAUDE_MEMORY_EMBEDDINGS=false
# Auto cleanup interval (ms)
CLAUDE_MEMORY_CLEANUP_INTERVAL=3600000
# Log level
LOG_LEVEL=info
Presets
// Minimal mode (1000 tokens)
CLAUDE_MEMORY_BUDGET=1000
CLAUDE_MEMORY_COMPRESSION=true
// Balanced mode (2000 tokens) - Default
CLAUDE_MEMORY_BUDGET=2000
// Comprehensive mode (4000 tokens)
CLAUDE_MEMORY_BUDGET=4000
// No limit (not recommended)
CLAUDE_MEMORY_BUDGET=-1
Database Schema
Core Tables
- facts - Key-value pairs linked to entities
- entities - People, projects, technologies, files, code
- relationships - Connections between entities
- decisions - Architectural/technical decisions
- tasks - Work items and todos
- preferences - User/project settings
- documents - Indexed files and metadata
- code_index - Code symbols (functions, classes, etc.)
- conversations - Message tracking for extraction status
- memory_usage - Historical metrics
- embeddings - Optional semantic vectors
See docs/SCHEMA.md for detailed schema reference.
Performance
Benchmarks
| Operation | Target | Typical |
|---|---|---|
| Memory lookup | <20ms | 8ms |
| Prompt assembly | <50ms | 25ms |
| Database query | <100ms | 45ms |
| Deduplication | <200ms | 80ms |
| Cleanup cycle | <500ms | 150ms |
Scalability
- Million records: ~200ms query time with indexes
- Database size: ~500MB per million facts
- Memory overhead: <50MB RAM
Token Reduction Examples
Traditional approach:
User: "What was our database decision?"
Needed: Retrieve last 50 messages (~3000 tokens)
Summarize into context (~1000 tokens overhead)
Answer query (~500 tokens)
Total: ~4500 tokens
Memory Engine approach:
User: "What was our database decision?"
Needed: Search "decision" entity (~50ms)
Find 2-3 relevant records (~100 tokens)
Return with relationships (~200 tokens)
Total: ~300 tokens (93% reduction)
Development
Project Structure
src/
├── index.ts # MCP server entry
├── config.ts # Configuration
├── types.ts # TypeScript types
├── database/
│ ├── connection.ts # SQLite management
│ ├── schema.ts # Database schema
│ └── migrations.ts # Schema versions
├── memory/
│ ├── layers.ts # L1-L4 layer system
│ ├── retrieval.ts # Search & ranking
│ ├── extraction.ts # Fact extraction
│ └── deduplication.ts # Merging logic
├── prompt/
│ ├── builder.ts # Prompt assembly
│ ├── tokenizer.ts # Token counting
│ └── budget.ts # Budget management
├── tools/
│ ├── memory-tools.ts # CRUD operations
│ ├── retrieval-tools.ts # Context retrieval
│ └── optimization-tools.ts
└── utils/
├── logger.ts # Structured logging
├── ranking.ts # Scoring engine
├── text-processing.ts # NLP helpers
└── tokenizer.ts # Token counter
Running Tests
# All tests
npm test
# Watch mode
npm test:watch
# Coverage report
npm test:coverage
# Specific suite
npm test -- memory.test.ts
Benchmarking
# Run performance benchmarks
npm run benchmark
# Profile specific operation
npm run benchmark -- --profile retrieval
Building
# Development
npm run dev
# Production build
npm run build
# Type checking
npx tsc --noEmit
Advanced Usage
Custom Extraction Rules
import { ExtractionEngine } from './memory/extraction';
const engine = new ExtractionEngine({
minFactImportance: 5,
extractCodeReferences: true,
customPatterns: {
'technology_stack': /stack:?\s*([^,\n]+)/gi,
'api_endpoint': /endpoint:\s*([^\s]+)/gi,
}
});
const result = engine.extract(claudeResponse);
Semantic Search with Embeddings
import { EmbeddingModel } from './memory/embeddings';
const embedder = new EmbeddingModel('all-MiniLM-L6-v2');
await embedder.initialize();
// Embeddings automatically generated on save
const results = await retrievalEngine.semanticSearch(
'database architecture',
{ useEmbeddings: true }
);
Export/Import Memory
# Export all decisions to Markdown
curl -X POST http://localhost:3000/export \
-d '{"format": "markdown", "type": "decision"}'
# Export facts as CSV
npm run export -- --type facts --format csv --output facts.csv
# Import from backup
npm run import -- --file backup.json --strategy merge
Extensibility
Add Custom Retrievers
class DomainSpecificRetriever extends BaseRetriever {
async retrieve(query: RetrievalQuery): Promise<RetrievalResult> {
// Custom logic
}
}
Add Custom Extractors
class CustomExtractor extends BaseExtractor {
extractCustomType(text: string): CustomItem[] {
// Domain-specific extraction
}
}
Add Rerankers
class CrossEncoderReranker {
rerank(items: RetrievalResult[]): RetrievalResult[] {
// Use larger model for final ranking
}
}
Troubleshooting
Memory growing too fast
# Analyze memory distribution
npm run analyze
# Adjust retention policy in config.ts
# Increase MEMORY_RETENTION_POLICY archiveAfterDays
# Lower minFactImportance threshold
# Manually cleanup old records
curl -X POST http://localhost:3000/cleanup --data '{"olderThanDays": 60}'
Slow retrieval
# Check indexes are present
npm run analyze -- --indexes
# Consider enabling embeddings
CLAUDE_MEMORY_EMBEDDINGS=true npm start
# Reduce context budget
CLAUDE_MEMORY_BUDGET=1000 npm start
High disk usage
# Run VACUUM
npm run analyze -- --optimize
# Export important records, delete others
npm run export -- --type decision --output decisions.json
Performance Tuning
For Limited Resources
# Minimize mode
CLAUDE_MEMORY_BUDGET=1000
CLAUDE_MEMORY_COMPRESSION=true
CLAUDE_MEMORY_CLEANUP_INTERVAL=7200000 # 2 hours
For Maximum Accuracy
# Maximum mode
CLAUDE_MEMORY_BUDGET=4000
CLAUDE_MEMORY_EMBEDDINGS=true
CLAUDE_MEMORY_DEDUP=true
Future Enhancements
- [ ] Semantic embeddings with local models
- [ ] Multi-user support
- [ ] PostgreSQL backend option
- [ ] Obsidian/Roam export plugins
- [ ] Custom extraction templates
- [ ] Graph visualization of relationships
- [ ] Memory import from ChatGPT
- [ ] Audio note support
Contributing
Contributions welcome! See CONTRIBUTING.md.
License
MIT - See LICENSE file
Support
- 📧 Email: support@example.com
- 🐛 Issues: GitHub Issues
- 💬 Discussions: GitHub Discussions
- 📖 Docs: https://claude-memory-engine.dev
Acknowledgments
Built with ❤️ for Claude Desktop users who need smarter memory management.
Inspired by:
- Obsidian's note-taking system
- Roam Research's bidirectional linking
- RAG (Retrieval Augmented Generation) patterns
- Modern database optimization techniques
Ready to 95x your Claude context efficiency? Start building smarter memory today.
git clone https://github.com/yourusername/claude-memory-engine.git
cd claude-memory-engine
npm install && npm run build && npm start
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。