A UI/HUD parser for the da Vinci Xi.
hudini extracts system state from the da Vinci Xi user interface embedded in surgical video and turns it into a timestamped event log. It recovers mounted instruments, surgeon-controlled arms, energy pedal presses, off-screen indicators, tool association indicators and UI text - from video alone.
With hudini you can:
- label video with the instrument on each arm and its time under surgeon control, without manual annotation
- search a video archive for every stapler firing, coagulation, vessel sealing, and any other pedal-triggered action
- mask the HUD before training to avoid shortcut learning
- screen for popups that leak the surgeon's account name before sharing a recording
Note
hudini requires the HUD to be visible in the recording. The HUD can lag the device by its rendering latency. The instrument catalogs cover the English and German system locales.
The instrument timeline of one SurgVU video, recovered by hudini. Each row is one arm. The color shows the instrument class. Saturated color marks the time under surgeon control. Ticks mark pedal presses.
uv tool install "hudini[rfdetr]"
hudini fetch
hudini parse video.mp4
hudini timeline video.hudini.jsonl.gz hudini parse writes one compressed observation log. hudini timeline
turns the log into a self-contained HTML page. Open the page in a
browser. If the video is in the same folder, the page plays it at the
selected time.
No Xi video at hand? Try hudini on a public sample from Open-H-Embodiment:
curl -LO https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-Open-H-Embodiment/resolve/main/Surgical/utenn/surgical_video_datasets/videos/chunk-000/observation.images.color/episode_000009.mp4
hudini parse episode_000009.mp4
hudini timeline episode_000009.hudini.jsonl.gz| Signal | Output | Method |
|---|---|---|
status |
arm status: under surgeon control, inactive, or warning | color rule, CNN for the endoscope pod |
arm |
the arm digit, 1 to 4 | CNN |
instrument |
the mounted instrument, matched to the catalog | OCR, catalog match |
pedals |
yellow and blue pedal presses | color rule |
pedal_label |
the action of a press (CUT, COAG, ...) | catalog, OCR when necessary |
laser |
the Firefly laser readout | color rule |
popups, banner |
the message text of each column, the system banner | OCR |
offscreen |
off-screen indicator bars, with state and arm | RF-DETR, CNN |
tool_association |
tool-association badges, with arm | RF-DETR |
Every observation carries a timestamp and a confidence score. The layout, the popups, and the detected indicators also carry their bounding boxes in frame coordinates.
A parse writes one file, the observation log <video>.hudini.jsonl.gz.
Everything else is a view of this log, computed when you read it:
- timeline: a self-contained HTML page with the video and the
intervals of each arm (
hudini timeline) - frame view: the state at each sampled frame, for comparison with
frame-level labels (
hudini timeline --frames). The rate of the layout signal decides which frames are sampled. - interval view: the runs of each lane, for example instrument presence and time under surgeon control (the data that the timeline shows)
The views apply the correction rules of the run as patches on top of the log. The log itself never changes. To print the state at one moment:
hudini query video.hudini.jsonl.gz --at 98:23The log stores observations, not one complete record for each frame. Each line records one reading of one signal:
{
"t": 98.4,
"f": 2952,
"s": "pedals",
"k": [2, "blue"],
"v": {"type": "press", "pressed": true},
"c": 0.99
}| Field | Meaning |
|---|---|
t |
video time in seconds |
f |
frame index |
s |
signal |
k |
key of the reading, here arm 2 and the blue pedal |
v |
value, null clears the key |
c |
confidence |
The first line is the header with the run configuration, and the last line is the footer. The file is gzip JSONL, one object per line, so any JSON lines tool can read it:
zcat video.hudini.jsonl.gz | head -1 | jq .signals
zcat video.hudini.jsonl.gz | grep '"s":"pedals"' | head -3When a user applies an energy preset, the Xi shows a popup with that
user's account name. If this is the real name of the surgeon, the
recording identifies the surgeon. hudini screen finds these popups in
a video or in a stored log. It reports each episode with the bounding
box of the popup, so you can redact the popup and keep the rest of the
frame.
hudini screen video.mp4
hudini screen video.hudini.jsonl.gz --jsonhudini is on PyPI. It needs Python 3.12 or newer. A GPU makes parsing faster.
uv tool install "hudini[rfdetr]"or, with pip, in a virtual environment:
pip install "hudini[rfdetr]"To use the newest unreleased code instead, install from GitHub:
uv tool install "hudini[rfdetr] @ git+https://github.com/claasdeboer/hudini"The [rfdetr] extra installs the two indicator detectors. Without it,
hudini reads every signal except offscreen and tool_association.
Skip the extra if you do not need these two signals.
The model weights are not in the package. hudini downloads them from nct-tso/hudini at a pinned revision on first use. To download them in advance:
hudini fetchOpenCV dependency note
paddleocr pins opencv-contrib-python==4.10. This pin installs a
second cv2 next to opencv-python-headless. This is an upstream
issue in paddlex. The parser runs with both installed.
| Command | Description |
|---|---|
hudini parse video.mp4 |
parse a video into its observation log |
hudini timeline log |
build a self-contained HTML timeline (--frames also writes the per-frame export) |
hudini query log --at 98:23 |
print the state at one moment, as JSON |
hudini serve folder/ |
serve an overview of the logs in a folder |
hudini screen video.mp4 |
find the popups that show an account name |
hudini frame image.png |
parse one image, JSON to stdout |
hudini catalog --locale de |
list the known instruments and pedal actions |
hudini fetch |
download the model checkpoints |
Options of hudini parse:
hudini parse video.mp4 --signals pedals # one signal, its dependencies enable themselves
hudini parse video.mp4 --rate pedals=30 # sample one signal at 30 fps
hudini parse video.mp4 --fast # lower sampling rates for every signalParse from Python:
from hudini.parser import Parser
parser = Parser() # loads every model once
log = parser.parse_video("video.mp4") # writes the log and returns itA stored log opens without the models. The views take the log, the correction patches of the run, and the catalog:
from hudini.catalog import Catalog
from hudini.corrections import correct, rules_from_settings
from hudini.storage import load
from hudini.views import frame_view, interval_view
log = load("video.hudini.jsonl.gz")
catalog = Catalog.load()
patches = correct(log, rules_from_settings(log.header.corrections), catalog)
records = frame_view(log, patches, catalog) # one dict for each sampled frame
intervals = interval_view(log, patches, catalog) # the temporal runs of each arm and lanegit clone https://github.com/claasdeboer/hudini && cd hudini
uv sync --extra dev --extra rfdetr
uv run pytest
uv run ruff check src/ tests/
uv run ty check src/hudini/hudini is accepted at AE-CAI @ MICCAI 2026. The BibTeX entry will follow when it is available.
The interface annotations for DSAD, hSDB-instrument, and SurgVU are in nct-tso/hudini-annotations.
Apache-2.0.
