This document covers all steps needed to reproduce training and evaluation results for Uni-World VLA.
After completing all setup steps, your repository should look like:
UniWorldVLA/
├── DA3/ # Depth Anything 3 (git submodule)
├── checkpoints/
│ ├── show-o-w-clip-vit/ # Show-O backbone (HuggingFace)
│ ├── phi-1_5/ # Phi-1.5 LLM (HuggingFace)
│ ├── magvitv2/ # MagViT-v2 tokenizer config
│ └── i3d/
│ └── i3d_torchscript.pt # I3D model for FVD evaluation
├── pretrained_models/
│ ├── tokenizer/
│ │ └── diffusion_pytorch_model.safetensors # fine-tuned VQ tokenizer (released)
│ ├── pretrain_ckpt/
│ │ └── unwrapped_model/
│ │ └── pytorch_model.bin # pre-trained PWM checkpoint (released)
│ ├── ckpt_sft_navsim/
│ │ └── unwrapped_model/
│ │ └── pytorch_model.bin # SFT checkpoint (released)
│ └── DA3-GIANT-LARGE/ # DA3 model weights
├── dataset/
│ └── navsim/
│ ├── nuplan_scene_blobs/
│ ├── navsim/
│ │ └── nuplan_img_logs/
│ ├── navsim_logs/
│ ├── maps/
│ └── ...
├── depth_cache_8_futrue_frame_flash/ # DA3 depth cache (self-generated, not released)
│ └── {scene_token}.pt
├── configs/
├── models/
├── training/
└── ...
DA3 is included as a git submodule pointing to the official Depth Anything 3 repository. Initialize it after cloning:
git submodule update --init --recursiveThen install the DA3 package:
pip install -e DA3/After installing, download the DA3-GIANT-LARGE model weights from the Depth Anything 3 releases and place them at:
pretrained_models/DA3-GIANT-LARGE/
Update the path in configs/sft_navsim/navsim.yaml:
model:
da3:
pretrained_model_path: "${experiment.base_root}/pretrained_models/DA3-GIANT-LARGE"Download from HuggingFace showlab/show-o and place at checkpoints/show-o-w-clip-vit/:
huggingface-cli download showlab/show-o --local-dir checkpoints/show-o-w-clip-vitDownload from HuggingFace microsoft/phi-1_5 and place at checkpoints/phi-1_5/:
huggingface-cli download microsoft/phi-1_5 --local-dir checkpoints/phi-1_5The MagViT-v2 tokenizer architecture config is already included in checkpoints/magvitv2/ within the repository. The fine-tuned tokenizer weights are released separately — see Section 3 below.
Download i3d_torchscript.pt from flateon/FVD-I3D-torchscript and place it at:
mkdir -p checkpoints/i3d
huggingface-cli download flateon/FVD-I3D-torchscript i3d_torchscript.pt \
--local-dir checkpoints/i3dAll released weights are hosted at SII-Rigby/UniWorldVLA on HuggingFace.
| Checkpoint | HF filename | Local path |
|---|---|---|
| VQ Tokenizer | tokenizer/diffusion_pytorch_model.safetensors |
pretrained_models/tokenizer/ |
| Pre-trained PWM | pretrain_ckpt/unwrapped_model/pytorch_model.bin |
pretrained_models/pretrain_ckpt/unwrapped_model/ |
| SFT NavSim | ckpt_sft_navsim/unwrapped_model/pytorch_model.bin |
pretrained_models/ckpt_sft_navsim/unwrapped_model/ |
Download with:
huggingface-cli download SII-Rigby/UniWorldVLA --local-dir pretrained_modelsOnce downloaded, the layout under pretrained_models/ should match the directory structure shown at the top of this document.
Follow the official NAVSIM instructions to download the dataset.
The expected structure under dataset/navsim/ is:
dataset/navsim/
├── nuplan_scene_blobs/ # raw scene blob files
├── navsim/
│ └── nuplan_img_logs/ # image logs
├── navsim_logs/ # log files
└── maps/ # HD map data
During training and evaluation, DepthEncoder runs DA3 inference on every NavSim scene's camera frames and caches the results to disk, avoiding redundant computation across epochs.
Cache directory (relative to base_root):
depth_cache_8_futrue_frame_flash/ # note: "futrue" is a typo in the original name — keep it as-is
└── {scene_token}.pt # one file per NavSim scene token
Note: The pre-computed depth cache is not released due to its large size. You need to generate it yourself using DA3 on the raw NavSim data (see below).
Use the following command to do a dry run over the training set — no model training or inference is performed:
EVAL_ONLY=1 RUN_FLASH_DATA_LOADER=1 \
bash scripts/finetune/navsim/run_sft_navsim_baseline8.shHow it works:
EVAL_ONLY=1: the val dataloader automatically loads the train split instead of the test splitRUN_FLASH_DATA_LOADER=1: tokens that are already cached are skipped; for uncached tokens, DA3 is called and the result is saved to disk — no model forward pass is executed- The run is safely resumable: already-generated files are never recomputed
To generate cache for the test split only (e.g. for evaluation), omit
EVAL_ONLY=1. The val dataloader defaults to the test split.
Open configs/sft_navsim/navsim.yaml and set:
experiment:
base_root: '/absolute/path/to/UniWorldVLA'Verify by running a quick import check:
LOCAL_RUN_PWM=1 python -c "
from omegaconf import OmegaConf
cfg = OmegaConf.load('configs/sft_navsim/navsim.yaml')
print('base_root:', cfg.experiment.base_root)
print('showo path:', cfg.model.showo.pretrained_model_path)
print('Config OK')
"