Unitree Go2 MCP Server

Unitree Go2 MCP Server

Enables LLM agents to control and monitor a Unitree Go2 robot through MCP tools, including live telemetry, navigation, camera feeds, and waypoint missions.

Category
访问服务器

README

🤖 MCP Web Portal — Unitree Go2 Robot Control Interface

A browser-based control and monitoring portal for the Unitree Go2 robot dog, built with Gradio and ROS2 keeping it 100% pythonic. The portal streams live camera feeds, displays sensor telemetry, provides remote navigation controls, supports autonomous waypoint missions, and integrates LLM-powered scene description. Portal Screenshot

Table of Contents


Feature Overview

Feature Description
🎥 Live Camera Front-facing video stream + Intel RealSense RGB/Depth feed, with YOLO object detection overlaid on either camera
🗺️ Map & Navigation 2D occupancy-grid map rendering with Nav2 goal-setting from the browser
🕹️ Remote Control Virtual joystick / directional controller for driving the robot over WirelessController messages
📡 Telemetry Live battery %, pose (x/y/yaw), sport-mode state, and IMU roll/pitch/yaw
🧠 AI Scene Description LLM-based image analysis of the camera feed (Azure OpenAI or local Ollama models via LiteLLM), with a configurable system prompt
🔊 Audio / Sounds Upload audio files and play them through the robot's onboard speaker
💡 LED Controller Adjust headlight color and brightness
📊 ROS Graph View Auto-generated visual graph of active ROS 2 topics, publishers, and subscribers
🛠️ Development Tab Diagnostic tools and dev utilities for debugging the ROS bridge/connections
🌐 MCP Server Exposes all portal-managed topics/services as MCP tools so an LLM agent can inspect and control the robot
📍 Waypoint Missions Save and replay autonomous navigation waypoints (data/waypoints/waypoints.json)

Architecture

Architecture Diagram

The robot connects to the host machine via rosbridge (default 127.0.0.1:9090), using either the DDS or WebRTC transport mode (MODE in config.py). All ROS 2 nodes are registered onto a shared executor in main.py, and their live data is surfaced to both the Gradio UI and the MCP server from the same DataStream object in web_backend/data_stream.py.


Project Structure

.
├── main.py                    # App entry point — ROS2 init, node registration, Gradio launch
├── config.py                  # Central settings: topic names, LLM config, rosbridge connection, UI flags
├── server.py                  # MCP server exposing ROS2 topics/services as agent tools
├── test.py                    # Test/scratch script
├── pyproject.toml / uv.lock   # uv-managed dependency lockfile
├── requirements.txt           # Full pinned dependency list (ROS2, Gradio, ML, MCP stack)
│
├── web_backend/
│   ├── action_sub.py          # SportMode action interface (stand, sit, hello, dance, etc.)
│   ├── audio_sub.py           # Access to the Go2's onboard speaker
│   ├── bm_status.py           # Battery / motor / IMU (roll, pitch, yaw) status
│   ├── camera.py              # Main front camera access
│   ├── camera_rs.py           # RealSense RGB/Depth camera access (requires RealSense ROS2 pkg on the Go2)
│   └── data_stream.py         # Core hub: ROS2 subscribers/publishers, YOLO inference, LLM calls, map builder
│
├── web_frontend/
│   ├── index.py                # Main tab UI — camera, map, telemetry
│   ├── action.py                # Actions tab UI — waypoints, missions, sport commands
│   ├── dev.py                   # Development tab UI — diagnostics
│   └── style.css                # Custom CSS
│
├── backends/                  # Additional backend service modules
├── utils/                     # Shared helper utilities
│
├── data/
│   ├── yolo/best2.pt           # YOLO model weights used for object detection
│   ├── waypoints/waypoints.json# Saved navigation waypoints
│   └── sounds/                 # Uploaded audio files for robot playback
│
└── github_media/               # Screenshots / media used in repo documentation

Tech Stack

Based on the project's pinned requirements.txt, the portal is built on:

  • Robotics / middleware: ROS 2 (rclpy, ros2cli tooling), rosbridge-suite for the WebSocket bridge to the robot, unitree_go/unitree_api/unitree_hg message packages, and the unitree_sdk2_python SDK (does not need seperate install, requirements.txt alread has it as an editable Git dependency)
  • Navigation: nav2-msgs, nav2-simple-commander, slam-toolbox, cartographer-ros-msgs for occupancy-grid mapping and goal navigation. You may choose any code you like. I used following repo [go2_slam_nav2] (https://github.com/andy-zhuo-02/go2_ros2_toolbox)
  • Web UI: gradio (v6.x) and gradio_client for the browser interface; fastapi / starlette / uvicorn underneath
  • Computer vision: opencv-python, ultralytics (YOLO) for object detection, torch / torchvision
  • LLM / agent layer: litellm (unified model API), openai, ollama (Python client), and mcp (the official Model Context Protocol SDK) for the agent-facing tool server
  • Audio: gTTS, pydub for text-to-speech / audio handling
  • Real-time transport: aiortc/aioice/av for optional WebRTC-based video/data channels
  • Misc: pandas, matplotlib/networkx (for the ROS topic/service graph visualization), redis, python-dotenv

Prerequisites

  • Ubuntu 22.04 (recommended)
  • ROS 2 Humble or later
  • Python 3.10+
  • Unitree ROS 2 SDK — unitreerobotics/unitree_ros2, installed and sourced
  • rosbridge_suite (bash ros2 launch rosbridge_server rosbridge_websocket_launch.xml ) (if you run it on laptop you can find it on 127.0.0.1:9090)
  • Nav2 (optional — required only for autonomous waypoint navigation)
  • Intel RealSense ROS 2 package installed on the Go2 (optional — required only for the RealSense RGB/Depth tab)
  • Ollama (optional — for local LLM inference instead of Azure OpenAI)
  • An NVIDIA GPU is not required, but the pinned requirements include CUDA-enabled torch/nvidia-* wheels for faster YOLO inference if one is available

Installation

1. Clone the repository

git clone https://github.com/sallu-786/Unitree_Go2_Web_Portal.git
cd Unitree_Go2_Web_Portal

2. Install Python dependencies

install from the pinned requirements.txt (note: this file includes ROS 2 Python packages, so it assumes a ROS 2 environment is already sourced/available):

pip install -r requirements.txt

3. Source ROS 2 and the Unitree setup script

source /opt/ros/humble/setup.bash
source /home/<your-user>/unitree_ros2/setup.sh

Update UNITREE_ROS2_SETUP_SH_PATH in config.py to match the actual path on your machine.

4. Configure config.py

At minimum, review and set:

  • ROSBRIDGE_IP / ROSBRIDGE_PORT — where rosbridge is running
  • MODE — "DDS" or "WEBRTC"
  • INTERFACE — your network interface for ROS 2 (ip a to find it)
  • ROBOT — a friendly name for your robot
  • Topic names (camera, cmd_vel, LIDAR, pose, odom, map, etc.) if they differ from your setup
  • UNITREE_ROS2_SETUP_SH_PATH and ROS_JS_LIB_PATH

5. (Optional) Set up .env for API keys

Rather than hard-coding credentials in config.py, create a .env file:

# .env
AZURE_API_KEY=your_key_here

config.py already includes a warning that hard-coded keys are unsafe — load them via python-dotenv instead:

from dotenv import load_dotenv
import os
load_dotenv()
AZURE_API_KEY = os.getenv("AZURE_API_KEY")

6. Run the portal

python main.py

The Gradio app launches at http://0.0.0.0:7860 by default.


Configuration Reference

All settings live in config.py. Key groups:

Connection

Setting Purpose
ROSBRIDGE_IP / ROSBRIDGE_PORT Address of the rosbridge WebSocket server (default 127.0.0.1:9090)
MODE Transport mode — "DDS" or "WEBRTC"
INTERFACE Network interface used for ROS 2 DDS traffic
ROBOT Display name for the connected robot

LLM / Scene Description

Setting Purpose
LLM_MODE "azure" or "ollama"
MODELS Dict mapping mode → friendly name → LiteLLM model string
DEFAULT_MODEL Default model per mode
AZURE_API_BASE / AZURE_OPENAI_DEPLOYMENT / AZURE_API_KEY / AZURE_API_VERSION Azure OpenAI credentials (use .env, not literals)
OLLAMA_API_BASE / OLLAMA_API_KEY Local Ollama endpoint (default http://localhost:11434)
SYSTEM_PROMPT / LLM_PROMPT Prompts used for periodic scene description; the shipped example is tuned for factory-floor PPE/hazard detection
MCP_AGENT_PROMPT System prompt for the MCP-connected agent, instructing it to use tools for robot state/control and never claim success without a confirmed tool result

UI Feature Flags

Setting Purpose
SHOW_CAMERA / SHOW_TOPICS / SHOW_SERVICES / SHOW_CONTROLLER / SHOW_DESCRIPTION / SHOW_LIDAR Toggle individual UI panels on/off
TTS_LANGUAGE Language code for text-to-speech ("ja" by default in the sample config)
UPDATE_INTERVAL Seconds between periodic scene-description calls
IMAGE_HEIGHT / IMAGE_WIDTH Camera stream display dimensions
YOLO_MODE "main" for the front camera or "rs" for RealSense as the YOLO detection source

Paths

Setting Purpose
YOLO_MODEL Path to YOLO weights (data/yolo/best2.pt)
SOUNDS_DIR Directory for uploaded playback audio
WAYPOINT_FILE JSON file storing saved navigation waypoints
UNITREE_ROS2_SETUP_SH_PATH Path to the Unitree ROS 2 setup.sh
ROS_JS_LIB_PATH Path to the JS library used for browser-side map/nav rendering

Topic Names — all remappable to match your robot's actual topic names: CAMERA_TOPIC_NAME, REALSENSE_CAMERA_COLOR, REALSENSE_CAMERA_DEPTH, CMD_VEL_PUB_TOPIC_NAME (+ _TYPE), LIDAR (+ LIDAR_MAX_POINTS), POSE (+ POSE_HEADER_FRAME_ID), ODOM, MAP, SPORTS, LFLOWCMD.

ROS Graph Styling — TOPIC_COLOR, PUBLISHER_COLOR, SUBSCRIBER_COLOR, NODE_SIZE, TOPIC_SIZE, PLOT_WIDTH, PLOT_HEIGHT control the appearance of the topic/service graph shown in the UI.


Running the MCP Portal

python main.py

This will:

  1. Initialize rclpy, instantiate all ROS 2 subscriber/publisher nodes defined in web_backend/, and register them on a shared executor.
  2. Launch the Gradio app with tabs for the main dashboard, actions/waypoints, and development diagnostics.
  3. Start the background loop that periodically grabs a camera frame, runs it through the configured LLM, and updates the on-screen scene description.

MCP Server & LLM Agent Integration

server.py starts an MCP server that mirrors the robot's ROS 2 surface as callable tools — battery/pose/telemetry reads, topic/service introspection, and movement/action commands. Any MCP-compatible client (a custom agent script, an IDE assistant, or a chat UI wired up with an MCP connector) can attach to it and:

  • List and inspect active ROS 2 topics and services
  • Read live telemetry (battery, pose, sport-mode state, sensors)
  • Issue movement or action commands through the exposed tools
  • Get grounded, tool-verified answers rather than the model guessing at robot state

The MCP_AGENT_PROMPT in config.py explicitly instructs the connected agent to rely on tool calls for anything robot-related and to never report success unless a tool call actually confirms it — useful guardrails when letting an LLM drive a physical robot.

To customize which model powers the natural-language side of the agent, add entries to MODELS in config.py using LiteLLM's model string format, e.g.:

MODELS = {
    "ollama": {
        "Gemma3": "ollama/gemma3:latest",
        "Llama3": "ollama/llama3:latest",   # ← new entry
    }
}

Accessing the Portal

Access Type URL
Local (same machine) http://localhost:7860
LAN (other devices) http://<robot-host-ip>:7860

Extending the Project

Add a new ROS 2 subscriber

  1. Create a new subscriber class in web_backend/, following the pattern of an existing one (e.g. the camera subscribers).
  2. Instantiate it inside DataStream.__init__() in web_backend/data_stream.py.
  3. Register the node with the executor in main.py:
    executor.add_node(launcher.your_new_subscriber)
    
  4. Expose the data via a property or method on DataStream so the frontend can read it.

Add a new UI tab

  1. Create web_frontend/my_tab.py and define a get_my_tab_page(demo, launcher) function using Gradio components.
  2. Wire it into main.py inside the gr.Tabs() block:
    with gr.Tab("My Tab"):
        get_my_tab_page(demo, launcher)
    

Change the LLM scene-description prompt

Edit SYSTEM_PROMPT and LLM_PROMPT in config.py:

SYSTEM_PROMPT = "You are a robot assistant."
LLM_PROMPT = "Describe the scene and highlight any hazards."

Add a new LLM model — add an entry to the MODELS dict as shown above in the MCP section.


License & Acknowledgements

This project is intended for internal/research use. Please respect the licenses of its third-party dependencies, including Gradio, ROS 2, the Unitree SDK, and the MCP SDK. See the repository's LICENSE file for details.

Acknowledgements:

推荐服务器

Baidu Map

Baidu Map

百度地图核心API现已全面兼容MCP协议,是国内首家兼容MCP协议的地图服务商。

官方
精选
JavaScript
Playwright MCP Server

Playwright MCP Server

一个模型上下文协议服务器,它使大型语言模型能够通过结构化的可访问性快照与网页进行交互,而无需视觉模型或屏幕截图。

官方
精选
TypeScript
Magic Component Platform (MCP)

Magic Component Platform (MCP)

一个由人工智能驱动的工具,可以从自然语言描述生成现代化的用户界面组件,并与流行的集成开发环境(IDE)集成,从而简化用户界面开发流程。

官方
精选
本地
TypeScript
Audiense Insights MCP Server

Audiense Insights MCP Server

通过模型上下文协议启用与 Audiense Insights 账户的交互,从而促进营销洞察和受众数据的提取和分析,包括人口统计信息、行为和影响者互动。

官方
精选
本地
TypeScript
VeyraX

VeyraX

一个单一的 MCP 工具,连接你所有喜爱的工具:Gmail、日历以及其他 40 多个工具。

官方
精选
本地
graphlit-mcp-server

graphlit-mcp-server

模型上下文协议 (MCP) 服务器实现了 MCP 客户端与 Graphlit 服务之间的集成。 除了网络爬取之外,还可以将任何内容(从 Slack 到 Gmail 再到播客订阅源)导入到 Graphlit 项目中,然后从 MCP 客户端检索相关内容。

官方
精选
TypeScript
Kagi MCP Server

Kagi MCP Server

一个 MCP 服务器,集成了 Kagi 搜索功能和 Claude AI,使 Claude 能够在回答需要最新信息的问题时执行实时网络搜索。

官方
精选
Python
e2b-mcp-server

e2b-mcp-server

使用 MCP 通过 e2b 运行代码。

官方
精选
Neon MCP Server

Neon MCP Server

用于与 Neon 管理 API 和数据库交互的 MCP 服务器

官方
精选
Exa MCP Server

Exa MCP Server

模型上下文协议(MCP)服务器允许像 Claude 这样的 AI 助手使用 Exa AI 搜索 API 进行网络搜索。这种设置允许 AI 模型以安全和受控的方式获取实时的网络信息。

官方
精选