mcp-data-pipeline-connector

mcp-data-pipeline-connector

Unified MCP server for querying CSV, Postgres, and REST API data sources via embedded DuckDB, enabling cross-source SQL joins with no external query service.

Category
访问服务器

README

MCP Data Pipeline Connector

npm mcp-data-pipeline-connector package

One MCP server for all your data sources — with cross-source SQL joins and no external query service. DuckDB runs embedded in-process, so you can join a CSV file against a Postgres table against a REST API response in a single query, entirely on your machine. Agents work with your data without needing source-specific knowledge or multiple MCP server configs.

Tool reference | Configuration | Contributing | Troubleshooting

Key features

  • Unified query interface: SQL across all connected sources via DuckDB — including cross-source joins.
  • Multiple source types: CSV/JSON files, PostgreSQL databases, and REST API endpoints in a single server.
  • Auto schema detection: Infers column names and types from CSV headers and Postgres metadata.
  • REST caching: REST API responses are cached with a configurable TTL to avoid redundant calls.
  • Schema normalization: Maps source-specific types to a standard set (string, number, date, boolean, json).
  • In-process query engine: DuckDB runs embedded — no separate query service to install or manage.

Why this over separate per-source MCP servers?

The common alternative is running one MCP server per data source — a postgres MCP server, a CSV MCP server, a REST MCP server. Each works fine in isolation, but they can't talk to each other.

mcp-data-pipeline-connector Separate per-source servers
Cross-source joins Native SQL via embedded DuckDB Not possible — agent must fetch and join manually
Config complexity One server entry in your MCP config One entry per source type
Query engine DuckDB in-process — no install, no service Depends on each source's query capabilities
Schema unification Normalizes all types to string/integer/number/datetime/boolean/json/unknown Each source uses its own type system
Data residency All queries run locally Depends on each connector's implementation

If you're asking questions that span multiple data sources — "join my sales CSV with the users table" — this is the right tool. If you only ever query one source type, a dedicated single-source server is simpler.

Disclaimers

mcp-data-pipeline-connector connects to data sources you configure and executes queries against them on behalf of your agent. Ensure agents only have the database permissions they need. Connection strings are never logged or transmitted; keep them out of version-controlled config files. Use environment variables for credentials.

Requirements

  • Node.js v20.19 or newer.
  • npm.
  • Optional: A running PostgreSQL instance for the Postgres connector.

Getting started

Add the following config to your MCP client:

{
  "mcpServers": {
    "data-connector": {
      "command": "npx",
      "args": ["-y", "mcp-data-pipeline-connector@latest"]
    }
  }
}

Define your data sources in ~/.mcp/data-sources.yaml:

sources:
  - name: sales
    type: csv
    path: ~/data/sales-2025.csv
  - name: users
    type: postgres
    connection_string: "${POSTGRES_URL}"
    tables: [users, subscriptions]

Store connection strings in environment variables, not directly in the YAML file.

MCP Client configuration

Amp · Claude Code · Cline · Cursor · VS Code · Windsurf · Zed

Your first prompt

Place a CSV file at ~/data/sample.csv, add it as a source in your config, then enter:

What columns are in the sample table? Show me the first 5 rows.

Your client should return the schema and a preview of the data.

Tools

Sources (2 tools)

  • connect_source
  • list_sources

Schema (2 tools)

  • list_tables
  • get_schema

Data (2 tools)

  • query
  • transform

Health (1 tool)

  • check_health

Configuration

--config / --sources-config

Path to the YAML file defining data sources.

Type: string Default: ~/.mcp/data-sources.yaml

--rest-cache-ttl

Time-to-live in seconds for cached REST API responses. Set to 0 to disable caching.

Type: number Default: 300

--max-rows

Maximum number of rows returned by a single query call. Prevents accidental large result sets.

Type: number Default: 1000

--read-only

Reject any SQL statements that are not SELECT queries. Enforces read-only access across all sources.

Type: boolean Default: true

Pass flags via the args property in your JSON config:

{
  "mcpServers": {
    "data-connector": {
      "command": "npx",
      "args": ["-y", "mcp-data-pipeline-connector@latest", "--max-rows=5000", "--rest-cache-ttl=60"]
    }
  }
}

Verification

Before publishing a new version, verify the server with MCP Inspector to confirm all tools are exposed correctly and the protocol handshake succeeds.

Interactive UI (opens browser):

npm run build && npm run inspect

CLI mode (scripted / CI-friendly):

# List all tools
npx @modelcontextprotocol/inspector --cli node dist/index.js --method tools/list

# List resources and prompts
npx @modelcontextprotocol/inspector --cli node dist/index.js --method resources/list
npx @modelcontextprotocol/inspector --cli node dist/index.js --method prompts/list

# Call a tool (example — replace with a relevant read-only tool for this plugin)
npx @modelcontextprotocol/inspector --cli node dist/index.js \
  --method tools/call --tool-name list_sources

# Call a tool with arguments
npx @modelcontextprotocol/inspector --cli node dist/index.js \
  --method tools/call --tool-name list_sources --tool-arg key=value

Run before publishing to catch regressions in tool registration and runtime startup.

Contributing

Each connector lives in src/connectors/ and must implement the DataConnector interface. Add fixture data files under tests/fixtures/ for integration tests. Never log connection strings or credentials — sanitize before any output or error message.

npm install && npm test

Listings

mcp-data-pipeline-connector is listed on MCP Registry and MCP Market.

Troubleshooting

  • REST source fails to connect: Confirm the URL is reachable and any auth env var is set. Use check_health to retest after startup.
  • Cross-source join returns no results: Ensure both sources are CSV type and registered before using source='_all'.
  • Query returns truncated: true: Increase --max-rows or add a LIMIT clause to your SQL.

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选