tokensense

tokensense

Model-agnostic context management middleware for LLM-powered workflows, providing persistent selective memory across sub-conversations using RAG over embedded summaries.

Category
访问服务器

README

TokenSense

Model-agnostic context management middleware for LLM-powered workflows. Gives projects persistent, selective memory across sub-conversations using RAG over embedded summaries, instead of re-explaining prior work or stuffing full history into the context window.

See tokensense_project.md for the full design doc.

Quickstart (using TokenSense in your own project)

Three steps, no local Postgres or pip install required:

1. Get a free hosted Postgres with pgvector. Neon works out of the box — sign up, create a project, and copy the connection string from the dashboard (looks like postgresql://user:pass@ep-xxxx.region.aws.neon.tech/neondb?sslmode=require). TokenSense creates its own extension/tables/indexes on first connect, so an empty database is all it needs. (Self-hosting Postgres instead? See Database.)

2. Install uv if you don't have it (brew install uv, or curl -LsSf https://astral.sh/uv/install.sh | sh). uvx runs TokenSense straight from GitHub — no separate install step, no PATH setup.

3. Add .mcp.json to your project (Claude Code, Cursor, or any MCP-compatible tool):

{
  "mcpServers": {
    "tokensense": {
      "command": "uvx",
      "args": ["--from", "tokensense[server] @ git+https://github.com/AayushKumbhare/tokensense", "tokensense", "serve", "--mcp"],
      "env": {
        "TOKENSENSE_DB_URL": "<your Neon connection string>",
        "OPENAI_API_KEY": "<your OpenAI key>"
      }
    }
  }
}

That's it — restart your MCP client and get_project_context / save_context / stats are available. Project memory is scoped by TOKENSENSE_PROJECT (or inferred from the git repo/folder name), so multiple projects can safely share one database. For automatic session capture instead of relying on the agent to call end_session, also add the hooks in Session capture.

Summarization/embeddings default to OpenAI (gpt-4o-mini + text-embedding-3-small, see docs/decisions.md #8) so the API key above is the only thing you need — no separate download or local model server. Prefer to run everything locally instead (no conversation content leaves your machine, no API key, no per-call cost)? Set both in the same env block:

ollama pull qwen2.5:3b        # summarizer (see docs/decisions.md #5)
ollama pull nomic-embed-text  # embeddings, 768-dim (see docs/decisions.md #4)
"env": {
  "TOKENSENSE_DB_URL": "<your Neon connection string>",
  "TOKENSENSE_SUMMARIZATION_MODEL": "qwen2.5:3b",
  "TOKENSENSE_EMBEDDING_MODEL": "ollama/nomic-embed-text"
}

Any LiteLLM model string works for either variable. Changing the embedding model invalidates stored vectors — see the re-migration note in docs/decisions.md.

Database

TokenSense just needs a TOKENSENSE_DB_URL pointing at any Postgres with the pgvector extension available — it creates the extension, tables, and HNSW indexes itself on first connect.

  • Hosted (recommended for most users): Neon or Supabase both support pgvector on their free tiers and give you a connection string with no local setup. Neon's free tier doesn't auto-pause on inactivity the way Supabase's does, so it's the better fit for an MCP server that connects sporadically between coding sessions.

  • Self-hosted (Docker): the included docker-compose.yml runs Postgres 17 + pgvector locally:

    cp .env.example .env   # then edit TOKENSENSE_DB_PASSWORD to a real value
    docker compose up -d --wait
    

    Exposes localhost:5432 with a persistent volume; URL is postgresql://tokensense:$TOKENSENSE_DB_PASSWORD@localhost:5432/tokensense (user/db name default to tokensense, password is required — compose refuses to start without it). Override TOKENSENSE_DB_USER / TOKENSENSE_DB_NAME / TOKENSENSE_DB_PORT in .env too if you don't want those defaults.

Sharing your database with a collaborator

If you want a friend/teammate to use your instance without them signing up for their own hosted Postgres, don't just hand them your TOKENSENSE_DB_URL — TokenSense's project isolation is enforced by the app, not the database, so whoever holds that connection string can read and write every project's memory, not just their own.

Instead, give each outside user their own database, provisioned from your existing Postgres server (one Neon project can hold many databases):

python scripts/create_user_database.py <your-owner-db-url> <friend-name>

This creates a new database + a Postgres role that owns it (and only it), and prints a connection string scoped to that database alone. Their role has zero grants on your main database or anyone else's — verified with a live connection attempt (permission denied), not just assumed. Give them that connection string; they drop it straight into their own .mcp.json's TOKENSENSE_DB_URL (see Quickstart) and everything else works exactly like a normal install — since their role owns their database, schema bootstrap (CREATE EXTENSION, tables, indexes) runs for them automatically on first connect, same as it did for you.

Install (editable, for development on TokenSense itself)

Only needed if you're changing TokenSense's own code — for using it as a memory server in another project, see the Quickstart above instead.

python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev,server]"

Needs a Postgres instance with pgvector (see Database) and, by default, an OPENAI_API_KEY in the environment — or the two ollama pull commands from the Quickstart above if you'd rather point TOKENSENSE_SUMMARIZATION_MODEL / TOKENSENSE_EMBEDDING_MODEL at local models instead.

Quickstart

from tokensense import TokenSenseClient

client = TokenSenseClient(
    provider="anthropic",
    api_key="your-key-here",
    db_url="postgresql://localhost/tokensense",
)

project = client.project("react-dashboard-build")
project.add_document("api_spec.md")  # optional: pin a file into the session
response = project.chat(messages=[{"role": "user", "content": "Let's work on the auth flow today"}])
print(client.stats())

project.end_session()

Documents added to a session ride along verbatim using the provider's prompt cache while its TTL is live (cheap cached reads); once it expires — or on providers without prompt caching — they're served as retrieved chunks instead (docs/decisions.md #6).

Zero-code server (proxy + MCP)

The SDK above requires writing code against TokenSenseClient. The server exposes the same core engine through two zero-code transports. Run it via uvx (see Quickstart), a local pip install -e ".[server]", or as a container:

docker build -t tokensense .
docker run --rm -p 8317:8317 -e TOKENSENSE_DB_URL=<your-db-url> tokensense serve
# or MCP over stdio: docker run -i --rm -e TOKENSENSE_DB_URL=<your-db-url> tokensense serve --mcp

docker compose --profile server up -d --wait builds and runs the server alongside the self-hosted Postgres from Database in one command.

Proxy transport

Any tool with a base_url override gets memory with no code changes:

export TOKENSENSE_DB_URL=<your-db-url>
tokensense serve            # listens on localhost:8317

Point your existing client at it — your API key is forwarded per-request, never stored:

curl localhost:8317/v1/messages \
  -H "x-api-key: $ANTHROPIC_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "x-tokensense-project: my-project" \
  -d '{"model": "claude-sonnet-5", "max_tokens": 1024,
       "messages": [{"role": "user", "content": "What did we decide last session?"}]}'

/v1/chat/completions (OpenAI-compatible) works the same way with an Authorization header. The response body is a pure passthrough; retrieval transparency comes back in X-TokenSense-* response headers. Project resolution: X-TokenSense-Project header → existing session binding → TOKENSENSE_PROJECT env var → git repo / folder name → default. Streaming is not supported yet.

MCP transport

For Claude Code, Cursor, and similar tools — memory is exposed as callable tools: get_project_context (retrieval), save_context (persist a decision mid-session, without waiting for session end), end_session, switch_project, and stats (token/CO₂ savings: retrieved-summary tokens vs. the raw session tokens they replaced, per docs/decisions.md #7; scoped to the current server process, i.e. one Claude Code session). Add to .mcp.json in your project directory (Claude Code) — see the Quickstart for the uvx-based config. If you installed TokenSense yourself instead (pip install/pipx install, so tokensense is already on PATH), the config simplifies to:

{
  "mcpServers": {
    "tokensense": {
      "command": "tokensense",
      "args": ["serve", "--mcp"],
      "env": { "TOKENSENSE_DB_URL": "<your-db-url>" }
    }
  }
}

The project binds once at session start (from cwd/git inference or TOKENSENSE_PROJECT) and never silently changes; switching mid-session requires the explicit switch_project tool.

Session capture (Claude Code hooks)

The host tool owns the conversation on the MCP transport, so relying on the agent to call end_session with notes makes memory capture best-effort. For deterministic capture, add hooks to .claude/settings.json — Claude Code pipes the session's transcript to TokenSense, which summarizes it (local qwen by default), embeds it, and stores it as project memory. SessionEnd is the primary capture point; PreCompact fires the same command right before Claude Code compacts its context — context exhaustion is exactly when memory matters most, and this snapshots the session before compaction rewrites it:

{
  "hooks": {
    "SessionEnd": [
      {
        "hooks": [
          {
            "type": "command",
            "command": "TOKENSENSE_DB_URL=<your-db-url> uvx --from 'tokensense[server] @ git+https://github.com/AayushKumbhare/tokensense' tokensense ingest-transcript"
          }
        ]
      }
    ],
    "PreCompact": [
      {
        "hooks": [
          {
            "type": "command",
            "command": "TOKENSENSE_DB_URL=<your-db-url> uvx --from 'tokensense[server] @ git+https://github.com/AayushKumbhare/tokensense' tokensense ingest-transcript"
          }
        ]
      }
    ]
  }
}

(If tokensense is already on PATH via pip/pipx install, drop the uvx --from ... wrapper and call tokensense ingest-transcript directly — same as the CLI everywhere else in this doc.)

Ingestion is idempotent per Claude Code session (re-firing either hook, or re-ingesting a resumed session's grown transcript, updates that session's memory in place — so a PreCompact snapshot is simply superseded by the SessionEnd capture), keeps only conversation text (thinking blocks, tool calls/outputs, and subagent sidechains are dropped), and always exits 0 so a capture failure never breaks the host tool. The same command works standalone: tokensense ingest-transcript <path-to-transcript.jsonl>.

See docs/decisions.md for the resolved design decisions (passthrough vs. envelope, agent-invoked retrieval, per-project document duplication).

Benchmarks

python benchmarks/run.py --offline                                  # plumbing + transport parity, no services
python benchmarks/run.py --db-url postgresql://localhost/ts_bench   # real embeddings/summarizer

Prints a transport-comparison table (direct engine path vs. proxy) with token savings and probe recall, plus a tracker-parity check confirming both transports share the same core.

Tests

pytest

When the Docker database is up, tests/test_live_integration.py additionally runs an end-to-end smoke test against it (session memory surviving into a second session, and project isolation on real HNSW-ranked retrieval); it skips itself automatically when the database is unreachable. Set TOKENSENSE_TEST_DB_URL to point it elsewhere.

With TOKENSENSE_LIVE_OLLAMA=1 and a local Ollama serving the default models (qwen2.5:3b + nomic-embed-text), tests/test_live_ollama.py also runs the fully local stack — real summarization and real embeddings, live semantic ranking, no API keys. benchmarks/summarizer_models.py compares candidate local summarizer models on latency and fact retention.

The remaining tests cover the logic that doesn't require a live Postgres connection: sliding window, summarizer chaining, cache-vs-RAG decision, tracker math, document chunking, project isolation guards (compiled query scoping, cascade FKs), the resolution chain, session lifecycle under concurrency, both server transports (with mocked upstreams), and SDK-vs-proxy transport parity. Migrating a pre-multi-project install: python scripts/migrate_add_project_id.py <db-url>.

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选