librechat-mnemonic

librechat-mnemonic

Enables automatic, project-scoped long-term memory for LibreChat conversations by recalling relevant memories before each turn and writing durable facts after each turn. It integrates via a proxy that resolves project context from MongoDB and exposes MCP tools for explicit memory management.

Category
访问服务器

README

librechat-mnemonic

Automatic, project-scoped long-term memory for LibreChat, backed by a local mnemonic MCP server.

Chats inside a LibreChat project recall that project's memories and write new ones back to it. Chats outside a project use the global pool. It is on by default and can be turned off per chat, per user, or entirely.

Nothing is forked or patched. This runs as one container alongside LibreChat.

What it does

  • Recalls before every turn. Relevant memories are retrieved and injected as context before the model is called. This does not depend on the model deciding to call a tool.
  • Writes after every turn. Durable facts are extracted from the exchange and stored, with duplicates detected and skipped.
  • Scopes by LibreChat project. A chat in the "Home Network" project reads and writes memories stamped with that project. Memories live in one global vault, partitioned by project, so nothing is siloed unless you want it to be.
  • Stays out of the way. /memory off in any chat, and that conversation stops recalling and storing.
  • Exposes tools too. An MCP endpoint lets agents search, correct, and forget memories explicitly when the automatic path is not enough.

How it works

browser  →  LibreChat  →  librechat-mnemonic  →  your model provider
                 │               │
                 │               ├── stdio ──→  mnemonic  ──→  vault (markdown + git)
                 │               │
                 └──── mongo ────┘   (read-only: conversation → project)

LibreChat has no server-side plugin API, so the integration hangs off two supported extension points:

  1. Custom endpoints with header placeholders. LibreChat resolves {{LIBRECHAT_USER_ID}} and {{LIBRECHAT_BODY_CONVERSATIONID}} into request headers. That is how the proxy knows who is asking and in which conversation.
  2. MCP servers over streamable HTTP, for the explicit tool surface.

The project is not in the request. LibreChat's ALLOWED_BODY_FIELDS is conversationId, parentMessageId, messageId and nothing else, so the proxy resolves it itself: conversations.chatProjectId → chatprojects.name, read from the same MongoDB LibreChat already uses. Its own collections are never written to.

Why a proxy and not just MCP tools

Because "automatic" and "the model decides" are different things.

LibreChat's own memory feature can be driven externally: with memory.agent.enabled unset, every run loads memories from the MemoryEntry collection and injects them with no tool call involved. But that lookup is keyed by user id alone. There is no conversation or project dimension in the schema, and the load happens inside LibreChat before anything external runs. So that channel gives automatic but user-global.

The only place that can see the project is the request path. Hence a proxy.

The MCP tools are still worth having, they just do a different job: correcting a memory the extractor got wrong, or searching for something the recall query missed.

Requirements

  • LibreChat v0.8.7 or later (projects landed in 0.8.7; header placeholders are older)
  • Access to LibreChat's MongoDB
  • An embedding provider for mnemonic: a local Ollama, or an OpenAI or Gemini key

Quick start

Add the service to your LibreChat compose file. See docker-compose.example.yml for the annotated version.

services:
  librechat-mnemonic:
    image: ghcr.io/claudedowling/librechat-mnemonic:latest
    restart: unless-stopped
    environment:
      LIBRECHAT_MONGO_URI: mongodb://mongo:27017/LibreChat
      UPSTREAMS: >-
        [{"name":"openai","baseUrl":"https://api.openai.com","api":"openai"}]
      OLLAMA_URL: http://ollama:11434
    volumes:
      - mnemonic-vault:/vault
      - mnemonic-projects:/projects
    networks: [librechat-network]

Then point LibreChat at it in librechat.yaml:

endpoints:
  custom:
    - name: 'OpenAI'
      apiKey: '${OPENAI_API_KEY}'
      baseURL: 'http://librechat-mnemonic:8710/openai/v1'
      models:
        default: ['gpt-4o']
      headers:
        x-librechat-user-id: '{{LIBRECHAT_USER_ID}}'
        x-librechat-conversation-id: '{{LIBRECHAT_BODY_CONVERSATIONID}}'

Restart LibreChat. Send a message. Type /memory status to confirm it is wired up.

The full example, including Anthropic and the MCP server, is in examples/librechat.yaml.

In-chat commands

The proxy answers these itself. The model is never called and no tokens are spent.

Command Effect
/memory List the commands
/memory on / /memory off Enable or disable memory for this chat
/memory status Show the current setting, its source, and the project
/memory default on|off Set your personal default for new chats
/memory reset Drop this chat's override and follow your default
/memory save <text> Store a memory now
/memory search <query> Search memory without involving the model
/memory forget <id> Delete a memory by id

Precedence is per-chat, then per-user, then MEMORY_DEFAULT_ENABLED.

How project scoping works

mnemonic derives project identity from a working directory. Its detection order is the git remote of the enclosing repo, then the git root folder name, then the plain basename of the directory. This uses the third branch: each LibreChat project gets a directory under MNEMONIC_PROJECT_ROOT, and its name becomes the mnemonic project.

A LibreChat project called Home Network becomes /projects/Home Network, which mnemonic resolves to { id: "home-network", name: "Home Network", source: "folder" }.

Writes use scope: global with that directory as cwd. mnemonic stores the note in the main vault while stamping it with the detected project. The note's frontmatter carries project: home-network and projectName: Home Network. That is what "one global vault, partitioned by project" means in practice.

What each recall scope actually returns

Verified against mnemonic 0.42, because the tool descriptions are misleading on this point:

MNEMONIC_RECALL_SCOPE A chat in project "Home Network" sees
project Only notes stamped home-network. Hard isolation.
all (default) Everything in the vault, with home-network notes boosted.
global Everything in the main vault, unboosted.

Note that global does not mean "notes with no project". mnemonic's tool description still says it returns only unscoped memories; the implementation returns every note in the main vault regardless of project stamp. If you need memories from one project kept out of another, use project.

Three more things to know:

  • The project directory must exist, and must exist on the filesystem of whichever process runs mnemonic. mnemonic calls simpleGit(cwd) outside its error guard, so a missing path fails the whole call. In the default spawn mode this is handled for you. With MNEMONIC_MODE=remote it is your job.
  • MNEMONIC_PROJECT_ROOT must not be inside a git repository. If it is, mnemonic will attribute every memory to that repo instead of to the project.
  • Project names collide. Two LibreChat users with a project of the same name share one mnemonic project. This service is designed for single-user and small-trusted-team installs; see Limitations.

Configuration

Everything is environment driven. Only LIBRECHAT_MONGO_URI and UPSTREAMS have no useful default.

LibreChat

Variable Default Description
LIBRECHAT_MONGO_URI required Connection string for LibreChat's MongoDB
LIBRECHAT_MONGO_DB from the URI Override the database name
LIBRECHAT_USER_HEADER x-librechat-user-id Header carrying {{LIBRECHAT_USER_ID}}
LIBRECHAT_CONVERSATION_HEADER x-librechat-conversation-id Header carrying {{LIBRECHAT_BODY_CONVERSATIONID}}

Upstreams

UPSTREAMS is a JSON array. Each entry mounts a provider at /<name>/..., and everything after the name is forwarded verbatim.

Field Required Description
name yes Path segment, e.g. openai → /openai/v1/chat/completions
baseUrl yes Provider root, such that <baseUrl>/v1/chat/completions is valid
api no openai (default) or anthropic
apiKey no Static credential replacing whatever LibreChat sends

mnemonic

Variable Default Description
MNEMONIC_MODE spawn spawn runs the bundled mnemonic over stdio; remote connects to a streamable-http instance
MNEMONIC_COMMAND bundled Executable used in spawn mode
MNEMONIC_URL none Required when MNEMONIC_MODE=remote
MNEMONIC_HEADERS {} JSON headers for the remote instance, e.g. auth
MNEMONIC_VAULT_PATH /vault Vault directory, passed through as VAULT_PATH
MNEMONIC_PROJECT_ROOT /projects Where per-project directories live
MNEMONIC_WRITE_SCOPE global global stores in the main vault stamped with the project; project writes a project vault
MNEMONIC_RECALL_SCOPE all project isolates, all boosts the current project, global returns the whole main vault. See the table above.
MNEMONIC_RECALL_LIMIT 6 Memories retrieved per turn
MNEMONIC_MIN_SIMILARITY 0.3 Similarity floor passed to recall
MNEMONIC_TIMEOUT_MS 20000 Per-call timeout
MNEMONIC_TAG librechat Tag added to everything this service writes

mnemonic's own variables (EMBED_PROVIDER, OLLAMA_URL, EMBED_MODEL, OPENAI_API_KEY, GEMINI_API_KEY, DISABLE_GIT, …) are passed through to the spawned process. See mnemonic's configuration.

Behaviour

Variable Default Description
MEMORY_DEFAULT_ENABLED true Whether memory is on for chats with no explicit setting
MEMORY_RECALL_ENABLED true Set false to write memories without injecting them
MEMORY_WRITE_MODE llm llm extracts automatically, explicit only on "remember that …", off disables writing
MEMORY_MAX_CONTEXT_CHARS 4000 Budget for the injected block
MEMORY_QUERY_MESSAGE_COUNT 3 User turns used to build the recall query
MEMORY_MAX_PER_TURN 3 Cap on memories written per exchange
MEMORY_DEDUPE_THRESHOLD 0.82 Recall score above which a candidate is treated as already known
MEMORY_COMMAND_PREFIX /memory Change if it clashes with something
MEMORY_PROJECTLESS global off disables memory entirely in chats not assigned to a project

Extraction model

Leave unset to reuse the chat's own model and credentials. Setting a small dedicated model is cheaper.

Variable Default Description
EXTRACT_BASE_URL none OpenAI-compatible base URL, including /v1
EXTRACT_MODEL none Model name
EXTRACT_API_KEY none Bearer token
EXTRACT_TIMEOUT_MS 30000 Extraction is detached; a timeout drops the write, never the reply

Service

Variable Default Description
PORT 8710 Listen port
HOST 0.0.0.0 Listen address
LOG_LEVEL info pino level
MCP_ENABLED true Serve the MCP endpoint
MCP_PATH /mcp Where to serve it

MCP tools

Available at /mcp for LibreChat agents.

Tool Purpose
search_memory Semantic search, project-scoped
save_memory Store a note deliberately
update_memory Correct an existing note
forget_memory Delete a note
memory_status Report the setting and project for this chat
set_memory_enabled Toggle automatic memory for this conversation

The server ships serverInstructions telling the agent that recall is already automatic, so it should reach for these only when the automatic path falls short.

Limitations

Worth knowing before you rely on it.

  • The first turn of a brand new chat may miss its project. The conversation document may not be written when the first request arrives. The proxy retries once, and the post-turn write re-resolves the project, so writes are correct from turn one. The very first recall can fall back to global.
  • Only traffic routed through the proxy is augmented. Endpoints configured to talk to a provider directly get no memory. That is deliberate: the proxy cannot see what it does not carry.
  • Project names are the identity. Renaming a LibreChat project starts a new mnemonic project; the old memories stay under the old name. Directories under MNEMONIC_PROJECT_ROOT can be renamed to match, but nothing does it for you.
  • Multi-user installs share memory by project name. There is no per-user partition in the vault. Fine for a personal or small-team instance, wrong for a multi-tenant one.
  • Automatic extraction is a judgement call made by a model. It will sometimes store something you would not have, and miss something you would. MEMORY_WRITE_MODE=explicit trades recall for precision.
  • Tool-calling turns are passed through untouched. Memory is injected on the request and extracted from the final text, so intermediate tool rounds are not analysed separately.

Images and releases

Published to the GitHub Container Registry:

ghcr.io/claudedowling/librechat-mnemonic:latest    # newest release
ghcr.io/claudedowling/librechat-mnemonic:0.1       # newest 0.1.x
ghcr.io/claudedowling/librechat-mnemonic:0.1.0     # exact version
ghcr.io/claudedowling/librechat-mnemonic:main      # tip of main, unreleased

Built for linux/amd64 and linux/arm64, with SBOM and signed build provenance. Verify a pull with:

gh attestation verify oci://ghcr.io/claudedowling/librechat-mnemonic:latest \
  --repo claudedowling/librechat-mnemonic

Pin to a minor tag such as :0.1 in production. latest moves across breaking changes while the project is pre-1.0.

To cut a release, bump version in package.json, then tag:

git tag v0.1.0 && git push origin v0.1.0

Development

npm install
npm test          # unit tests
npm run typecheck
npm run dev       # watch mode

Requires Node 22 or later.

The pieces worth understanding first: src/memory/service.ts holds the scoping rules that both entrypoints share, src/proxy/handler.ts is the request path, and src/mnemonic/projects.ts explains the directory trick and the constraints that come with it.

Licence

MIT

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选