tcp-tuner
Enables autonomous TCP congestion-control tuning through MCP tools for benchmarking, sysctl parameter adjustments, fault injection, and runbook generation within Kubernetes clusters.
README
TCP Congestion Tuner
An agentic TCP congestion-control tuning system where Bob (IBM watsonx Code Assistant) autonomously benchmarks, diagnoses, tunes Linux kernel TCP parameters, injects network faults, queries live Prometheus metrics, and commits runbooks to git — converging toward a user-defined SLO with zero human involvement per iteration.
Architecture
┌─────────────────────────────────────────────────────────────┐
│ Bob (TCP Tuner Mode) │
│ "cwnd collapsing — switch to BBR, increase rmem_max" │
└──────────────────────────┬──────────────────────────────────┘
│ MCP tools (stdio)
┌──────────────────────────▼──────────────────────────────────┐
│ MCP Server (Node.js) │
│ run_benchmark │ get_sysctl_params │ apply_sysctl │
│ │ get_benchmark_history │
└──────────────────────────┬──────────────────────────────────┘
│ kubectl
┌──────────────────────────▼──────────────────────────────────┐
│ kind Kubernetes Cluster (3 nodes) │
│ │
│ ┌─────────────────┐ ┌──────────────────────────────┐ │
│ │ iperf3-server │◄───│ iperf3-client (Job) │ │
│ │ (Deployment) │ │ measures throughput/RTT │ │
│ └─────────────────┘ └──────────────────────────────┘ │
│ │
│ ┌──────────────────────────────────────────────────────┐ │
│ │ sysctl-tuner DaemonSet (privileged, hostNetwork) │ │
│ │ worker-node-1 pod │ worker-node-2 pod │ │
│ │ reads/writes kernel TCP sysctl params │ │
│ └──────────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────┘
Prerequisites
| Tool | Version | Install |
|---|---|---|
| Docker Desktop | 29+ | docker.com |
| kubectl | v1.34+ | bundled with Docker Desktop |
| kind | v0.29+ | curl -Lo kind.exe https://kind.sigs.k8s.io/dl/v0.29.0/kind-windows-amd64 |
| Node.js | v22 LTS | winget install OpenJS.NodeJS.LTS |
| Python | 3.12+ | python.org |
Quick Start
1. Spin up the cluster
cd cluster
.\start.ps1
This creates a 3-node kind cluster (1 control-plane + 2 workers), deploys the iperf3 server, and starts the sysctl-tuner DaemonSet on both worker nodes.
2. Install and build the MCP server
cd mcp-server
npm install
npm run build
3. Register the MCP server with Bob
Add to your Bob MCP config (~/.bob/mcp-settings.json or via Bob UI → Settings → MCP):
{
"mcpServers": {
"tcp-tuner": {
"command": "node",
"args": ["<absolute-path-to-repo>/tcp-congestion-tuner/mcp-server/src/index.js"],
"cwd": "<absolute-path-to-repo>/tcp-congestion-tuner"
}
}
}
4. Load the custom Bob mode
Copy .bob/custom_modes.yaml to your Bob workspace config, or merge it into your existing custom modes file.
5. Start tuning
Open Bob, switch to TCP Tuner mode, and say:
"Run a baseline benchmark, inspect the current sysctl configuration, and autonomously tune TCP parameters to maximise throughput."
Bob will run the full detect → diagnose → tune → validate loop.
Manual CLI Usage
# Run a benchmark
python scripts/benchmark.py run
# Show last 10 benchmark runs
python scripts/benchmark.py history
# Compare the last two runs
python scripts/benchmark.py compare
# Read current sysctl params
python scripts/benchmark.py sysctl get
# Apply a sysctl change manually
python scripts/benchmark.py sysctl set net.ipv4.tcp_congestion_control bbr
MCP Tools Reference
| # | Tool | Description |
|---|---|---|
| 1 | run_benchmark |
Launches iperf3 Job, returns throughput/RTT/retransmits + sysctl snapshot |
| 2 | get_sysctl_params |
Reads all TCP sysctl values from a worker node |
| 3 | apply_sysctl |
Writes a sysctl value on one or all worker nodes via privileged DaemonSet pod |
| 4 | get_benchmark_history |
Returns past N runs for trend comparison |
| 5 | inject_fault |
Injects packet loss + delay via tc netem on worker node interfaces |
| 6 | clear_fault |
Removes netem qdiscs, restores clean network |
| 7 | save_runbook |
Writes Markdown runbook to runbooks/, git commit, git push, returns SHA |
| 8 | get_metrics |
Runs PromQL query against in-cluster Prometheus, returns live node metrics |
| 9 | check_slo |
Evaluates last benchmark against SLO targets; returns pass/fail + next-action recommendation |
| 10 | export_report |
Generates a self-contained HTML report with SVG trend charts from benchmark history |
| 11 | benchmark_regression |
Compares latest run vs a golden baseline; returns pass/fail for use in CI |
Observability
Deploy Prometheus + Grafana into the cluster:
kubectl apply -f manifests/monitoring.yaml
Open the live dashboard (auto port-forwards and opens browser):
.\cluster\port-forward.ps1
# Grafana: http://localhost:3000/d/tcp-tuner (admin / tcptuner)
# Prometheus: http://localhost:9090
The TCP Congestion Tuner dashboard shows:
- Network transmit/receive Mbps (all nodes)
- TCP retransmits/sec (kernel counter via node-exporter)
- CPU usage % (iperf3 load visibility)
Known Limitations (kind cluster)
These are expected artifacts of running Kubernetes inside Docker — not bugs:
| Limitation | Explanation |
|---|---|
| Baseline retransmits ~1,000–2,000 on clean network | veth/bridge interfaces inside kind containers exhibit higher retransmit rates than bare-metal under high-throughput iperf3. The reduction from tuning is real; the absolute baseline is inflated. |
net.core.rmem_max / wmem_max unavailable |
These params are not namespaced and cannot be written from inside a container network namespace. net.ipv4.tcp_rmem/wmem are used instead. |
tc netem fault affects control-plane pod too |
kind's control-plane runs as a container on the same host network. Injecting fault on control-plane's eth0 can cause kubectl timeouts. Use node_selector=worker in production demos. |
| Throughput limited by loopback BDP | 50–75 Gbps is a loopback ceiling, not a real network limit. Buffer sizing changes show their full impact on real network paths with RTT > 1ms. |
Key Kubernetes Concepts Demonstrated
- DaemonSet — sysctl-tuner runs on every worker node automatically
- Privileged pods with hostNetwork — required to read/write host kernel parameters
- Job — iperf3 client runs once and terminates cleanly
- Headless Service — iperf3 server addressable by DNS name within the cluster
- Node-level sysctl tuning — per-node kernel parameter management in K8s
Tuning Playbook (What Bob Does)
| Change | Reason | Expected Impact |
|---|---|---|
tcp_congestion_control=bbr |
BBR tracks bottleneck bandwidth directly; better in high-BDP paths | +10–40% throughput, fewer retransmits |
rmem_max=134217728 |
Larger receive buffers allow higher in-flight data | Higher throughput on high-latency links |
tcp_slow_start_after_idle=0 |
Prevents cwnd reset after idle bursts | Important for bursty market-data feeds |
tcp_notsent_lowat=16384 |
Reduces bufferbloat in the socket send queue | Lower RTT under load |
Teardown
cd cluster
.\teardown.ps1
Bob Integration
This project is designed to be driven entirely by Bob in TCP Tuner mode. Three key prompts cover the full lifecycle:
Baseline + autonomous tuning
Run a baseline benchmark, read the current sysctl configuration, then autonomously tune TCP
parameters to maximise throughput. Show a before/after comparison table after each change.
Save a runbook to the runbooks/ folder when done.
SLO-driven convergence loop
Tune until retransmits < 500 and throughput > 50000 Mbps. Run autonomously — apply one
sysctl change per iteration, benchmark, check SLO, repeat until SLO passes or playbook
exhausted. Save a runbook to git when done.
Fault-injection resilience
Inject 3% packet loss. Run the fault-resilience tuning loop: benchmark under fault, tune
to compensate, clear the fault, confirm recovery, save runbook.
Bob uses the tcp-tuner skill (.bob/skills/tcp-tuner/SKILL.md) for structured agentic
best practices: 4 phases, 6 invariants, a full tuning playbook, and runbook templates.
Project Structure
tcp-congestion-tuner/
├── cluster/
│ ├── kind-config.yaml # 1 control-plane + 2 worker nodes
│ ├── start.ps1 # cluster bootstrap script
│ ├── port-forward.ps1 # Grafana + Prometheus port-forward for demos
│ └── teardown.ps1 # cluster cleanup
├── manifests/
│ ├── iperf-server.yaml # iperf3 server Deployment + headless Service
│ ├── iperf-client-job.yaml # iperf3 benchmark Job
│ ├── tuner-daemonset.yaml # privileged sysctl-tuner DaemonSet (alpine + iproute2)
│ ├── monitoring.yaml # Prometheus + node-exporter DaemonSet + Grafana
│ └── alerts.yaml # TcpHighRetransmits + TcpThroughputDrop alert rules
├── mcp-server/
│ ├── package.json
│ ├── tsconfig.json
│ └── src/
│ └── index.ts # 9 MCP tools (TypeScript source)
├── scripts/
│ └── benchmark.py # CLI benchmark runner, comparator, rollback
├── tests/
│ └── test_benchmark.py # 119 pytest tests, 98% coverage
├── runbooks/ # Bob-generated runbooks (auto-committed by save_runbook)
├── .bob/
│ ├── custom_modes.yaml # TCP Tuner Bob mode with SLO convergence loop
│ └── skills/tcp-tuner/
│ └── SKILL.md # Reusable agentic best-practices skill
└── README.md
推荐服务器
Baidu Map
百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。
Playwright MCP Server
一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。
Magic Component Platform (MCP)
一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。
Audiense Insights MCP Server
通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。
VeyraX
一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。
graphlit-mcp-server
模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。
Kagi MCP Server
一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。
e2b-mcp-server
使用 MCP 通过 e2b 运行代码。
Neon MCP Server
用于与 Neon 管理 API 和数据库交互的 MCP 服务器
Exa MCP Server
模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。