AstraMemory Local

AstraMemory Local

Local-first memory daemon for AI coding agents that captures session transcripts, distills typed memories (decisions, facts, lessons, commands, todos), and serves them via hybrid search through MCP tools.

Category
访问服务器

README

AstraMemory Local

Local-first memory daemon for AI coding agents — wire-compatible with memory-plugin.

Why it exists

Claude Code sessions compact and terminate, taking context with them. AstraMemory Local captures every session transcript, distills typed memories (decisions, facts, lessons, commands, todos), and serves them back via hybrid search (BM25 + vector + importance + freshness). It runs entirely on your workstation — no cloud account, no data leaves your machine. The plugin's hooks post to the local daemon instead of the SaaS endpoint through a single environment variable swap.


Quick start (5 commands)

npm install -g @astragenie/astramemory-local
astra-memory init
# follow the wizard — picks Ollama or Azure, writes config.yaml + secrets.env
astra-memory service install
export MEMORY_API_URL=http://127.0.0.1:7777
export MEMORY_BEARER=$(astra-memory token print)

Restart Claude Code. All plugin hooks (PreCompact, SessionEnd, SubagentStop) now post to the local daemon. No other plugin changes needed.


Architecture

 memory-plugin hooks (unchanged)
    |
    |  POST /ingest/transcript
    |  Authorization: Bearer <token>
    v
+------------------+      SQLite (memory.sqlite)
|  HTTP daemon     | ---> +-------------------+
|  Fastify         |      | sessions          |
|  127.0.0.1:7777  |      | messages          |
+------------------+      | transcripts       |
                          | jobs (queue)      |
                          | memories          |
                          | memories_fts (FTS5)|
                          | memories_vec (vec0)|
                          | budget_spend      |
                          +-------------------+
                                   |
                          in-process worker loop
                                   |
                          8-stage distillation
                          (cleanup -> normalize ->
                           chunk -> compact ->
                           extract -> reduce ->
                           memory-normalize ->
                           embed + index)
                                   |
                    +--------------+--------------+
                    |              |              |
              memories        FTS5 index    sqlite-vec
                (rows)       (BM25 search)  (cosine ANN)
                    |              |              |
                    +--------------+--------------+
                                   |
                          hybrid score fusion
                          a*BM25 + b*cosine +
                          c*importance + d*freshness
                                   |
                          GET /search  POST /recall
                                   |
                          /recall in plugin slash commands

Single Node process. Workers run in-process on a polling loop. SQLite is the source of truth. Everything derived (vectors, FTS rows, compactions) can be rebuilt by replaying the jobs table.


Memory types

Type Description Example
decision Architectural or design choice made during a session "Use sqlite-vec for v1 vector storage"
fact Objective project fact, configuration detail "Port 7777 is the default daemon port"
lesson Something that went wrong and how it was resolved "sqlite-vec rowid must match memories rowid"
command CLI command or script worth remembering "npm run build && npm test -- migrate"
todo Outstanding work item surfaced in conversation "Add reembed job when provider changes"

Provider matrix

Concern Ollama (local, free) Azure OpenAI (cloud)
LLM compaction qwen2.5-coder:7b (default) gpt-4.1 or any deployment
LLM extraction qwen2.5-coder:7b (default) gpt-4.1 or any deployment
Embedding nomic-embed-text-v2-moe (1024-dim) text-embedding-3-small (1024 via dimensions)
Cost $0 (local inference) ~$0.02/1K tokens + $0.0001/1K embed tokens
Setup ollama serve + ollama pull <model> Azure portal + endpoint + deployment name

Providers are configurable independently per stage. Embedding provider is system-wide — switching requires astra-memory rebuild --reembed to re-index all memories in the new model's vector space.

See docs/providers.md for full setup instructions.


MCP tools (Claude Code auto-discovery)

The daemon exposes a Model Context Protocol (Streamable HTTP) endpoint at POST /mcp. Claude Code discovers and calls the 4 tools below automatically when configured in .mcp.json.

Tool Description Maps to
search_memory Hybrid FTS + vector search with optional type/repo/project/since filters GET /search
recall_memory Top-K semantic recall (default k=5) POST /recall
remember Direct memory insert, bypasses distillation POST /remember
get_health Daemon health probe: { ok, version } GET /health

Plugin .mcp.json wiring:

{
  "mcpServers": {
    "astramem": {
      "type": "http",
      "url": "${MEMORY_API_URL}/mcp",
      "headers": { "Authorization": "Bearer ${MEMORY_BEARER}" }
    }
  }
}

Set MEMORY_API_URL=http://127.0.0.1:7777 and MEMORY_BEARER to your token (printed by astra-memory token print).


Budget cap

The daily LLM spend cap (default: $10 USD) is enforced before each LLM call.

  • Ollama always reports $0 cost — the cap only applies to Azure usage.
  • When the cap is reached, pending distillation jobs move to paused state. Ingest continues to accept transcripts (no data loss). Distillation resumes the next UTC day automatically.
  • Override: astra-memory budget --reset (logged).
  • Check current spend: astra-memory budget.

Commands reference

Command What it does
astra-memory init Interactive wizard — writes config + secrets, runs migrations, installs service
astra-memory serve [--port N] Start daemon in foreground (dev/debug)
astra-memory service install Register daemon as a user-scope OS service
astra-memory service status Show service state
astra-memory service start Start the service
astra-memory service stop Stop the service
astra-memory service uninstall Remove the service unit
astra-memory doctor Run all health checks, print table
astra-memory doctor --json Machine-readable health check output
astra-memory search "<query>" Hybrid search, print results table
astra-memory search "<query>" --type decision Filter by memory type
astra-memory recall "<question>" Top-5 semantic recall (alias for search k=5)
astra-memory remember "<text>" [--type] Direct insert, bypasses distillation pipeline
astra-memory queue Show pending/failed jobs
astra-memory queue --state failed Show only failed jobs
astra-memory rebuild [--reembed] Rebuild derived indexes; --reembed re-vectors all
astra-memory providers list List configured providers and their health
astra-memory providers test [name] Ping provider, print latency + dim
astra-memory budget Show today and month spend vs cap
astra-memory budget --reset Clear today's spend counter (override, logged)
astra-memory token print Print the current Bearer token
astra-memory token rotate Generate new token, invalidate the old one

Further reading


Status

v0.1.0 — Waves 1-4 of the implementation plan completed.

  • Wave 1: SQLite schema, migration runner, FTS5, sqlite-vec, ingest endpoint, Fastify server, CLI skeleton.
  • Wave 2: Job worker loop, hybrid search, service install adapters, Ollama + Azure providers.
  • Wave 3: 8-stage distillation pipeline, budget tracker, Zod-validated extraction.
  • Wave 4: Install wizard, cross-OS CI matrix, E2E plugin integration test, this documentation.

Spec: astramemory-plugin/docs/superpowers/specs/2026-06-27-astramemory-local-v1-design.md

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选