tcp-tuner

tcp-tuner

Enables autonomous TCP congestion-control tuning through MCP tools for benchmarking, sysctl parameter adjustments, fault injection, and runbook generation within Kubernetes clusters.

Category
访问服务器

README

TCP Congestion Tuner

An agentic TCP congestion-control tuning system where Bob (IBM watsonx Code Assistant) autonomously benchmarks, diagnoses, tunes Linux kernel TCP parameters, injects network faults, queries live Prometheus metrics, and commits runbooks to git — converging toward a user-defined SLO with zero human involvement per iteration.

Architecture

┌─────────────────────────────────────────────────────────────┐
│                        Bob (TCP Tuner Mode)                  │
│  "cwnd collapsing — switch to BBR, increase rmem_max"       │
└──────────────────────────┬──────────────────────────────────┘
                           │ MCP tools (stdio)
┌──────────────────────────▼──────────────────────────────────┐
│                    MCP Server (Node.js)                      │
│  run_benchmark  │  get_sysctl_params  │  apply_sysctl       │
│                 │  get_benchmark_history                     │
└──────────────────────────┬──────────────────────────────────┘
                           │ kubectl
┌──────────────────────────▼──────────────────────────────────┐
│              kind Kubernetes Cluster (3 nodes)               │
│                                                              │
│  ┌─────────────────┐    ┌──────────────────────────────┐   │
│  │  iperf3-server  │◄───│  iperf3-client (Job)         │   │
│  │  (Deployment)   │    │  measures throughput/RTT      │   │
│  └─────────────────┘    └──────────────────────────────┘   │
│                                                              │
│  ┌──────────────────────────────────────────────────────┐   │
│  │  sysctl-tuner DaemonSet (privileged, hostNetwork)    │   │
│  │  worker-node-1 pod │ worker-node-2 pod               │   │
│  │  reads/writes kernel TCP sysctl params               │   │
│  └──────────────────────────────────────────────────────┘   │
└─────────────────────────────────────────────────────────────┘

Prerequisites

Tool Version Install
Docker Desktop 29+ docker.com
kubectl v1.34+ bundled with Docker Desktop
kind v0.29+ curl -Lo kind.exe https://kind.sigs.k8s.io/dl/v0.29.0/kind-windows-amd64
Node.js v22 LTS winget install OpenJS.NodeJS.LTS
Python 3.12+ python.org

Quick Start

1. Spin up the cluster

cd cluster
.\start.ps1

This creates a 3-node kind cluster (1 control-plane + 2 workers), deploys the iperf3 server, and starts the sysctl-tuner DaemonSet on both worker nodes.

2. Install and build the MCP server

cd mcp-server
npm install
npm run build

3. Register the MCP server with Bob

Add to your Bob MCP config (~/.bob/mcp-settings.json or via Bob UI → Settings → MCP):

{
  "mcpServers": {
    "tcp-tuner": {
      "command": "node",
      "args": ["<absolute-path-to-repo>/tcp-congestion-tuner/mcp-server/src/index.js"],
      "cwd": "<absolute-path-to-repo>/tcp-congestion-tuner"
    }
  }
}

4. Load the custom Bob mode

Copy .bob/custom_modes.yaml to your Bob workspace config, or merge it into your existing custom modes file.

5. Start tuning

Open Bob, switch to TCP Tuner mode, and say:

"Run a baseline benchmark, inspect the current sysctl configuration, and autonomously tune TCP parameters to maximise throughput."

Bob will run the full detect → diagnose → tune → validate loop.


Manual CLI Usage

# Run a benchmark
python scripts/benchmark.py run

# Show last 10 benchmark runs
python scripts/benchmark.py history

# Compare the last two runs
python scripts/benchmark.py compare

# Read current sysctl params
python scripts/benchmark.py sysctl get

# Apply a sysctl change manually
python scripts/benchmark.py sysctl set net.ipv4.tcp_congestion_control bbr

MCP Tools Reference

# Tool Description
1 run_benchmark Launches iperf3 Job, returns throughput/RTT/retransmits + sysctl snapshot
2 get_sysctl_params Reads all TCP sysctl values from a worker node
3 apply_sysctl Writes a sysctl value on one or all worker nodes via privileged DaemonSet pod
4 get_benchmark_history Returns past N runs for trend comparison
5 inject_fault Injects packet loss + delay via tc netem on worker node interfaces
6 clear_fault Removes netem qdiscs, restores clean network
7 save_runbook Writes Markdown runbook to runbooks/, git commit, git push, returns SHA
8 get_metrics Runs PromQL query against in-cluster Prometheus, returns live node metrics
9 check_slo Evaluates last benchmark against SLO targets; returns pass/fail + next-action recommendation
10 export_report Generates a self-contained HTML report with SVG trend charts from benchmark history
11 benchmark_regression Compares latest run vs a golden baseline; returns pass/fail for use in CI

Observability

Deploy Prometheus + Grafana into the cluster:

kubectl apply -f manifests/monitoring.yaml

Open the live dashboard (auto port-forwards and opens browser):

.\cluster\port-forward.ps1
# Grafana: http://localhost:3000/d/tcp-tuner  (admin / tcptuner)
# Prometheus: http://localhost:9090

The TCP Congestion Tuner dashboard shows:

  • Network transmit/receive Mbps (all nodes)
  • TCP retransmits/sec (kernel counter via node-exporter)
  • CPU usage % (iperf3 load visibility)

Known Limitations (kind cluster)

These are expected artifacts of running Kubernetes inside Docker — not bugs:

Limitation Explanation
Baseline retransmits ~1,000–2,000 on clean network veth/bridge interfaces inside kind containers exhibit higher retransmit rates than bare-metal under high-throughput iperf3. The reduction from tuning is real; the absolute baseline is inflated.
net.core.rmem_max / wmem_max unavailable These params are not namespaced and cannot be written from inside a container network namespace. net.ipv4.tcp_rmem/wmem are used instead.
tc netem fault affects control-plane pod too kind's control-plane runs as a container on the same host network. Injecting fault on control-plane's eth0 can cause kubectl timeouts. Use node_selector=worker in production demos.
Throughput limited by loopback BDP 50–75 Gbps is a loopback ceiling, not a real network limit. Buffer sizing changes show their full impact on real network paths with RTT > 1ms.

Key Kubernetes Concepts Demonstrated

  • DaemonSet — sysctl-tuner runs on every worker node automatically
  • Privileged pods with hostNetwork — required to read/write host kernel parameters
  • Job — iperf3 client runs once and terminates cleanly
  • Headless Service — iperf3 server addressable by DNS name within the cluster
  • Node-level sysctl tuning — per-node kernel parameter management in K8s

Tuning Playbook (What Bob Does)

Change Reason Expected Impact
tcp_congestion_control=bbr BBR tracks bottleneck bandwidth directly; better in high-BDP paths +10–40% throughput, fewer retransmits
rmem_max=134217728 Larger receive buffers allow higher in-flight data Higher throughput on high-latency links
tcp_slow_start_after_idle=0 Prevents cwnd reset after idle bursts Important for bursty market-data feeds
tcp_notsent_lowat=16384 Reduces bufferbloat in the socket send queue Lower RTT under load

Teardown

cd cluster
.\teardown.ps1

Bob Integration

This project is designed to be driven entirely by Bob in TCP Tuner mode. Three key prompts cover the full lifecycle:

Baseline + autonomous tuning

Run a baseline benchmark, read the current sysctl configuration, then autonomously tune TCP
parameters to maximise throughput. Show a before/after comparison table after each change.
Save a runbook to the runbooks/ folder when done.

SLO-driven convergence loop

Tune until retransmits < 500 and throughput > 50000 Mbps. Run autonomously — apply one
sysctl change per iteration, benchmark, check SLO, repeat until SLO passes or playbook
exhausted. Save a runbook to git when done.

Fault-injection resilience

Inject 3% packet loss. Run the fault-resilience tuning loop: benchmark under fault, tune
to compensate, clear the fault, confirm recovery, save runbook.

Bob uses the tcp-tuner skill (.bob/skills/tcp-tuner/SKILL.md) for structured agentic best practices: 4 phases, 6 invariants, a full tuning playbook, and runbook templates.


Project Structure

tcp-congestion-tuner/
├── cluster/
│   ├── kind-config.yaml          # 1 control-plane + 2 worker nodes
│   ├── start.ps1                 # cluster bootstrap script
│   ├── port-forward.ps1          # Grafana + Prometheus port-forward for demos
│   └── teardown.ps1              # cluster cleanup
├── manifests/
│   ├── iperf-server.yaml         # iperf3 server Deployment + headless Service
│   ├── iperf-client-job.yaml     # iperf3 benchmark Job
│   ├── tuner-daemonset.yaml      # privileged sysctl-tuner DaemonSet (alpine + iproute2)
│   ├── monitoring.yaml           # Prometheus + node-exporter DaemonSet + Grafana
│   └── alerts.yaml               # TcpHighRetransmits + TcpThroughputDrop alert rules
├── mcp-server/
│   ├── package.json
│   ├── tsconfig.json
│   └── src/
│       └── index.ts              # 9 MCP tools (TypeScript source)
├── scripts/
│   └── benchmark.py              # CLI benchmark runner, comparator, rollback
├── tests/
│   └── test_benchmark.py         # 119 pytest tests, 98% coverage
├── runbooks/                     # Bob-generated runbooks (auto-committed by save_runbook)
├── .bob/
│   ├── custom_modes.yaml         # TCP Tuner Bob mode with SLO convergence loop
│   └── skills/tcp-tuner/
│       └── SKILL.md              # Reusable agentic best-practices skill
└── README.md

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选