Skip to content
This repository was archived by the owner on Aug 31, 2026. It is now read-only.

Latest commit

 

History

History
143 lines (100 loc) · 4.65 KB

File metadata and controls

143 lines (100 loc) · 4.65 KB

RoboCache

Archived August 2026. Last GPU validation 2025-11-08 on H100 PCIe at commit 0db3726; numbers below are from that run and have not been re-validated.

GPU-Accelerated Data Preprocessing for Robot Learning

License CUDA Python PyTorch


Overview

RoboCache is a GPU-accelerated data preprocessing library for robot foundation models: trajectory resampling, multimodal sensor fusion, and point cloud voxelization as CUDA kernels with a PyTorch fallback.

Measured results (H100 PCIe, November 2025):

Benchmark inputs were synthetic dataset-shaped tensors (torch.randn) - shaped like Isaac Gym, TartanAir, nuScenes, and KITTI samples, but no real dataset files were loaded (see benchmarks/real_world_datasets.py).


Installation

Not published to PyPI. Install from source:

git clone https://github.com/GOATnote-Inc/robogoat.git
cd robogoat/robocache
pip install -e .

Prerequisites: Python 3.10+, PyTorch 2.0+, CUDA 12.1+ toolkit for the CUDA extensions (a CPU-only install works but uses the fallback - see Known Issues in the repository README).


Quick Start

import torch
import robocache

# GPU-accelerated trajectory resampling
source_data = torch.randn(32, 500, 256, device='cuda', dtype=torch.bfloat16)
source_times = torch.linspace(0, 5, 500, device='cuda').unsqueeze(0).expand(32, -1)
target_times = torch.linspace(0, 5, 250, device='cuda').unsqueeze(0).expand(32, -1)

resampled = robocache.resample_trajectories(source_data, source_times, target_times)
# H100 measured: 2.605 ms P50 for this config (32x500->256, dim 256, bf16)

Performance

All numbers from the November 2025 H100 PCIe run; see the CSV for raw data.

Trajectory resampling (CUDA kernel vs. PyTorch CPU)

Source: bench/results/benchmark_h100_20251106_172811.csv

Config (B x S -> T, D) CUDA P50 PyTorch CPU P50 Speedup
8 x 250, 128 0.184 ms 20.14 ms ~110x
32 x 500, 256 2.605 ms 38.39 ms ~15x
64 x 1000, 512 20.05 ms 75.69 ms ~3.8x

End-to-end training (H100)

Pipeline ms/step Speedup
Baseline (PyTorch preprocessing) 18.28 1.00x
RoboCache preprocessing 14.04 1.30x

Where it loses

Config (B x S -> T, D) RoboCache PyTorch GPU Result
64 x 4096 -> 1024, 32 0.190 ms 0.140 ms 0.74x (slower)

Long sequences fall out of L1 cache; see ../KNOWN_LIMITATIONS.md for the crossover analysis.


Architecture

RoboCache implements three memory patterns:

  1. L1-resident (trajectory, fusion): binary search + linear interpolation; effective while timestamp arrays fit in L1.
  2. Bandwidth-bound (voxelization): atomic scatter operations.
  3. BF16 storage with FP32 interpolation internally.

See: ../docs/ARCHITECTURE.md


Documentation


Citation

@software{robocache2025,
  author = {Dent, Brandon},
  title = {RoboCache: GPU-Accelerated Data Preprocessing for Robot Learning},
  year = {2025},
  publisher = {GitHub},
  howpublished = {\url{https://github.com/GOATnote-Inc/robogoat}},
  version = {1.0.0}
}

License

Apache 2.0 - See LICENSE for details.