llama-memory

llama-memory

Provides persistent history and semantic memory for llama-server, enabling cross-session recall and memory management through MCP tools.

Category
访问服务器

README

llama-memory

MCP Memory Service for llama-server with persistent history and semantic memory via Postgres + PGVector.

Note: This project is intended for local/demo use only. Do not expose it directly to the internet without additional hardening (HTTPS, proper auth, backups).

What it does

  • Semantic memory: save and retrieve memories by meaning, not just keywords.
  • Conversation bridge: LLM creates conversations automatically; memories are linked.
  • Cross-session recall: ask "what did we talk about before?" and get accurate answers.
  • MCP protocol: works directly with llama-server's built-in MCP support.

Requirements

  • Python 3.11 (Miniconda recommended)
  • PostgreSQL 16+ with PGVector extension
  • llama-server with --jinja flag (required for tool calling)
  • nomic-embed-text running on llama-server (port 8081 by default)

Installation

# Clone the repo
git clone https://github.com/noualit/llama-memory-local.git
cd llama-memory-local

# Create environment
conda create -n llama-memory python=3.11
conda activate llama-memory

# Install dependencies
pip install -e .

Configuration

Copy .env.example to .env and edit:

cp .env.example .env

Example:

# Database
DATABASE_URL="postgresql://postgres:yourpassword@localhost:5432/llamamem"

# Llama-server (LLM)
LLAMA_SERVER_BASE_URL="http://localhost:8080"

# Embedding model (nomic-embed-text via llama-server)
EMBEDDING_MODEL_URL="http://localhost:8081"

# Embedding model name (default: nomic-embed-text)
EMBEDDING_MODEL_NAME="nomic-embed-text"

# Service port
SERVICE_PORT=9001

Setup database

Create the database and run migrations:

psql -U postgres -c "CREATE DATABASE llamamem;"
alembic upgrade head

The application also ensures basic schema on startup for convenience.

Run the service

# Using the script
.\scripts\run_server.ps1

# Or directly
python -m uvicorn app.main:app --host 0.0.0.0 --port 9001

Connect to llama-server

Add to your llama-server MCP configuration:

{
  "mcpServers": {
    "llama-memory": {
      "url": "http://YOUR_SERVER_IP:9001/mcp"
    }
  }
}

The service must be reachable from llama-server. Use the actual IP, not localhost if they run on different machines.

MCP Tools

Tool Description
create_conversation Create a new conversation session
list_conversations List conversations with memory count
get_conversation_history Get all memories in a conversation
search_memories Semantic search across all memories
save_memory Store an important fact or decision

System prompt

You can:

  • Fetch the recommended system prompt from the service:

    • GET /system-prompt → returns plain text.
  • Or paste this minimal version into llama-server:

MEMORY WORKFLOW:
- At the start of each new conversation, call create_conversation with a short title.
- Use the conversation_id from create_conversation when calling save_memory.
- Before answering questions about past topics, call search_memories FIRST.
- When the user shares important information, save it with save_memory.
- If list_conversations has previous chats, check get_conversation_history for context.

Health check

curl http://localhost:9001/health

Returns DB status, embedding service status, and tool count.

Architecture

High-level structure:

  • app/main.py — FastAPI app, lifespan, /system-prompt
  • app/settings.py — Pydantic settings from .env
  • app/clients/embeddings.py — Calls nomic-embed-text for vectors
  • app/db/engine.py — asyncpg connection pool (singleton)
  • app/db/schema.py — Auto-creates tables on startup
  • app/mcp/endpoint.py — MCP protocol handlers, rate limiter
  • app/mcp/tools/ — Individual tool implementations
  • migrations/ — Alembic database migrations

Development

# Run tests
pytest tests/ -v

# Run with auto-reload
python -m uvicorn app.main:app --host 0.0.0.0 --port 9001 --reload

See CONTRIBUTING.md for contribution guidelines.

License

MIT (see LICENSE file).

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选
mcp-server-qdrant

mcp-server-qdrant

这个仓库展示了如何为向量搜索引擎 Qdrant 创建一个 MCP (Managed Control Plane) 服务器的示例。

官方
精选
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选