CLP (CKA-Guided Layer Pruning) identifies and removes redundant transformer layers in Vision-Language-Action (VLA) backbones prior to finetuning. By utilizing Centered Kernel Alignment (CKA) to isolate structurally repetitive states, CLP drastically reduces computational footprint and accelerates inference speeds while fully preserving or even improving manipulation success rates across downstream imitation learning tasks.
- TODO List
- 📦 Installation
- 🚀 Part 1: Quickstart — Finetune & Evaluate CLP on Your Own Dataset
- 🔁 Part 2: Reproducing Paper Experiments
- Citation
- Acknowledgement
- <Upload remaining pruned / finetuned checkpoints to Hugging Face and update links below>
- <Add CLP server/client evaluation scripts for SimplerEnv (youliangtan/SimplerEnv fork), wiring
Gr00tPolicyinto Google Robot / WidowX+Bridge tasks>
Clone the repository:
git clone https://github.com/jibby2803/CLP_VLA.git
cd CLP_VLACreate an environment and install the base dependencies (this repo builds on the NVIDIA Isaac GR00T codebase):
conda create -n clp_vla python=3.10
conda activate clp_vla
pip install uv
uv pip install --upgrade setuptools
uv pip install -e ".[base]"
# install flash attention
wget "https://github.com/Dao-AILab/flash-attention/releases/download/v2.7.1.post4/flash_attn-2.7.1.post4+cu12torch2.4cxx11abiFALSE-cp310-cp310-linux_x86_64.whl"
uv pip install flash_attn-2.7.1.post4+cu12torch2.4cxx11abiFALSE-cp310-cp310-linux_x86_64.whlTo run LIBERO / RoboCasa simulation for evaluation:
pip install uv
git clone https://github.com/ARISE-Initiative/robosuite
cd robosuite && uv pip install -e . && cd ..
git clone <robocasa_repo_url> ./robocasa
cd robocasa && uv pip install -e . && cd ..
python robocasa/scripts/download_kitchen_assets.py
python robocasa/scripts/setup_macros.pyThis section is a detailed, step-by-step guide for running one specific model on one specific dataset — i.e. how to bring your own robot demonstrations, apply CKA-Guided Layer Pruning (CLP), finetune, and evaluate. If you instead want to reproduce the exact experiments from the paper (LIBERO, RoboCasa, SimplerEnv, with all model variants), skip ahead to Part 2.
Convert your demonstrations into the LeRobot v2.1 dataset format. E.g: https://huggingface.co/datasets/binhng/groceries_to_basket
Run a forward pass over your prepared dataset and cache the per-layer hidden states. This is a debug/analysis-only run — it loads the full pretrained backbone, runs a single training step, and dumps each transformer layer's hidden states to disk (save_hidden_state=True, capped at 1 step by default):
bash my_scripts/debug.shKey flags in my_scripts/debug.sh you may want to edit:
MODEL_DIR— path to the base pretrained checkpoint (e.g. GR00T-N1.5-3B).DATA_PATH— path to your LeRobot-format dataset from Step 1.--save-hidden-state True— enables the hidden-state dump (off by default everywhere else).--save-hidden-state-dir/--save-hidden-state-max-steps— where to write the.npyfiles and how many steps to capture (default: 1 step, then auto-disables).
Then open the notebook below to visualize the layer-wise CKA similarity matrix and identify the high-CKA (redundant) layers to prune:
jupyter notebook notebook/cka.ipynbThe notebook outputs a list of layer indices to keep for both the backbone (VLM) and the action head (DiT / VL self-attention). Update kept_layer_idx_list in gr00t/model/backbone/eagle_backbone.py and kept_layer_idx_list_dit / kept_layer_idx_list_vl_self_attn in gr00t/model/action_head/flow_matching_action_head.py with the layers you've identified.
Once the redundant layers are identified, finetune the pruned backbone on your dataset:
bash my_scripts/train_libero.sh(or copy this script and point it at your own DATA_PATH / data_config — see my_scripts/train_libero.sh and my_scripts/AKHOA_closepot_cka.sh for examples on different datasets).
Key flags exposed by scripts/gr00t_finetune.py:
| Flag | Default | Description |
|---|---|---|
--base-model-path |
— | Path to the full pretrained checkpoint. |
--dataset-path |
— | Path to your LeRobot-format dataset. |
--data_config |
— | module:ClassName for your dataset's modality config / transforms. |
--prune-model |
True |
Whether to prune the architecture (using the kept_layer_idx_list* you set in Step 2) after loading the full pretrained weights. Set to False to train a non-pruned baseline. |
--tune-llm / --tune-visual |
False |
Whether to finetune the backbone's LLM / vision tower. |
--tune-projector / --tune-diffusion-model |
True |
Whether to finetune the action head's projector / DiT. |
--batch-size, --max-steps, --save-steps |
— | Standard training hyperparameters. |
--save-hidden-state |
False |
Debug-only; leave off for real training runs (see Step 2). |
--output-dir |
/tmp/gr00t |
Where checkpoints get saved. |
Why pruning order matters: training always loads the full pretrained checkpoint first (
pruned=False), then prunes the in-memory architecture afterward via--prune-model. This avoids any shape/key mismatch that would happen if you tried to load full pretrained weights directly into an already-pruned architecture.
Evaluate your finetuned, pruned checkpoint in LIBERO:
bash evaluation/libero.shKey flags in evaluation/eval_libero.py (set via evaluation/libero.sh):
| Flag | Default | Description |
|---|---|---|
--args.pretrained_model_path |
— | Path to your finetuned checkpoint directory. |
--args.load_pruned_model |
True |
Build the pruned architecture before loading weights — use this for a finetuned checkpoint saved after pruning (Step 3). Set to False only when evaluating a base/unpruned checkpoint. |
--args.model_type |
pi0 |
Set to gr00tn15 for this codebase. |
--args.task_suite_name |
libero_goal |
Which LIBERO task suite to run. |
--args.num_trials_per_task |
10 |
Number of rollouts per task. |
Why this flag matters: evaluation builds the architecture pruned first, then loads your finetuned weights — the opposite order from training, but necessary because the checkpoint you're loading is already pruned.
--args.load_pruned_model Trueensures the architecture matches the savedstate_dictexactly.
For RoboCasa or your own custom simulation/real-robot eval harness, follow the same pattern: build a Gr00tPolicy(..., load_pruned_model=<True/False to match your checkpoint>) and call policy.get_action(obs).
This section provides scripts to reproduce the simulation experiments reported in the paper, organized by benchmark. Each benchmark section is split into Evaluation (run our released checkpoints) and Finetuning (train the CLP model yourself from scratch).
Click to expand: Evaluation & Finetuning
We release the following checkpoints reported in our paper:
| Model | Description | 🤗 Download Link |
|---|---|---|
<checkpoint_name> |
Base (no pruning), LIBERO | Coming soon |
<checkpoint_name> |
CLP-pruned (CKA), LIBERO | Coming soon |
Step 1 — Set up the LIBERO simulator.
pip install uv
git clone https://github.com/ARISE-Initiative/robosuite
cd robosuite && uv pip install -e . && cd ..(See Installation above for the full LIBERO/RoboCasa simulator setup.)
Step 2 — Run the evaluation.
bash evaluation/libero.shwhich under the hood runs:
python evaluation/eval_libero.py \
--args.model_type=gr00tn15 \
--args.pretrained_model_path="<path_to_checkpoint>" \
--args.load_pruned_model=True \
--args.task_suite_name="libero_goal" \
--args.num_trials_per_task=50 \
--args.exp_name="<exp_name>"Key evaluation arguments:
| Argument | Description |
|---|---|
--args.pretrained_model_path |
Path to a checkpoint dir (contains model.safetensors / *.safetensors shards and config.json). |
--args.load_pruned_model |
True to build the pruned architecture before loading weights (use for the CLP checkpoint above). Set to False when evaluating the Base (no pruning) checkpoint. |
--args.task_suite_name |
LIBERO suite to evaluate (e.g. libero_spatial, libero_object, libero_goal, libero_10). |
--args.num_trials_per_task |
Number of rollout episodes per task. |
--args.exp_name |
Name used for the results/video/log output directory, written to eval_results/<exp_name>/. |
| Dataset | Description | 🤗 Download Link |
|---|---|---|
merged_libero_mask_noops_lerobot_v4 |
LIBERO (LeRobot v2.1 format) | binhng/merged_libero_mask_noops_lerobot_v4 |
Step 1 — Download the base GR00T-N1.5 pretrained checkpoint and the LIBERO dataset from the table above (or your own, converted to LeRobot v2.1 format — see Part 1, Step 1).
Step 2 — Identify the redundant layers to prune, following Part 1, Step 2 (bash my_scripts/debug.sh then notebook/cka.ipynb), and set kept_layer_idx_list* accordingly in gr00t/model/backbone/eagle_backbone.py / gr00t/model/action_head/flow_matching_action_head.py.
Step 3 — Launch training.
bash my_scripts/my_libero2.shwhich under the hood runs:
CUDA_VISIBLE_DEVICES=0 python scripts/gr00t_finetune.py \
--base-model-path <path_to_GR00T-N1.5_checkpoint> \
--dataset-path <path_to_libero_dataset> \
--data_config examples.My_libero.data_config:LiberoDataConfig \
--video-backend torchvision_av \
--tune-diffusion-model \
--tune-projector \
--prune-model True \
--batch-size 32 \
--save-steps 10000 \
--max-steps 100000 \
--num-gpus 1 \
--output-dir <output_dir> \
--report-to wandbKey training arguments:
| Argument | Description |
|---|---|
--base-model-path |
Path to the full pretrained GR00T-N1.5 checkpoint. |
--dataset-path |
Path to the LIBERO dataset from Step 1. |
--prune-model |
True to prune the architecture (using the kept_layer_idx_list* from Step 2) after loading the full pretrained weights — this is what makes the run "CLP" rather than "Base". |
--tune-projector / --tune-diffusion-model |
Finetune the action head's projector / DiT. |
--batch-size, --max-steps, --save-steps |
Standard training hyperparameters. |
--num-gpus |
Set >1 to automatically launch via torchrun. |
--output-dir |
Where checkpoints and logs are written. |
--report-to |
wandb to log to Weights & Biases, tensorboard for local logging. |
Click to expand: Evaluation & Finetuning
We release the following checkpoints reported in our paper:
| Model | Description | 🤗 Download Link |
|---|---|---|
<checkpoint_name> |
Base (no pruning), RoboCasa | Coming soon |
<checkpoint_name> |
CLP-pruned (CKA), RoboCasa | Coming soon |
Step 1 — Install the RoboCasa simulator.
pip install uv
git clone https://github.com/ARISE-Initiative/robosuite
cd robosuite && uv pip install -e . && cd ..
git clone <robocasa_repo_url> ./robocasa
cd robocasa && uv pip install -e . && cd ..
python robocasa/scripts/download_kitchen_assets.py
python robocasa/scripts/setup_macros.pyStep 2 — Run the evaluation.
bash evaluation/robocasa0.sh
bash evaluation/robocasa1.shPoint each script at your checkpoint and set --args.load_pruned_model=True for the CLP checkpoint, or False for the Base (no pruning) checkpoint — same convention as LIBERO evaluation above.
Key evaluation arguments (same flags as LIBERO evaluation, plus):
| Argument | Description |
|---|---|
--args.pretrained_model_path |
Path to a checkpoint dir. |
--args.load_pruned_model |
True for the CLP checkpoint, False for Base. |
--args.num_trials_per_task |
Number of rollout episodes per task. |
--args.exp_name |
Results/log output directory name. |
| Dataset | Description | 🤗 Download Link |
|---|---|---|
robocasa_30_demos_lerobot_5_chosen_tasks_v3 |
RoboCasa, 5 chosen tasks, 30 demos each (LeRobot v2.1 format) | binhng/robocasa_30_demos_lerobot_5_chosen_tasks_v3 |
The procedure is identical to LIBERO finetuning — the same Steps 1-2 apply (download the base GR00T-N1.5 checkpoint, identify redundant layers via CKA). The only difference is the dataset and data config:
bash my_scripts/robocasa_cka.shwhich under the hood runs:
CUDA_VISIBLE_DEVICES=0 python scripts/gr00t_finetune.py \
--base-model-path <path_to_GR00T-N1.5_checkpoint> \
--dataset-path <path_to_robocasa_dataset> \
--data_config examples.My_robocasa.data_config:RobocasaDataConfig \
--video-backend torchvision_av \
--tune-diffusion-model \
--tune-projector \
--prune-model True \
--batch-size 32 \
--save-steps 10000 \
--max-steps 100000 \
--num-gpus 1 \
--output-dir <output_dir> \
--report-to wandbClick to expand: Evaluation
We follow youliangtan/SimplerEnv (a GR00T-compatible fork of SimplerEnv) to set up the simulation environment.
Step 1 — Install system dependencies.
sudo apt update
sudo apt install libegl1-mesa-dev libglu1-mesaStep 2 — Clone and install the SimplerEnv fork.
git clone https://github.com/youliangtan/SimplerEnv.git
cd SimplerEnv
# follow the installation instructions in this repo's README to set up
# the SAPIEN/ManiSkill2 simulator, Google Robot, and WidowX+Bridge environments🚧 The CLP-specific server/client evaluation scripts that wire
Gr00tPolicy(fromgr00t/model/policy.py) into this SimplerEnv fork are not yet included in this repo — see the TODO list. Once added, evaluation will follow the same pattern as LIBERO/RoboCasa above: load a checkpoint with--load_pruned_model=Truefor the CLP model orFalsefor the Base model, then roll out episodes against the Google Robot / WidowX+Bridge tasks provided by the SimplerEnv fork.
Dataset. Bridge (WidowX) data used for SimplerEnv finetuning, in LeRobot format:
| Dataset | Description | 🤗 Download Link |
|---|---|---|
bridge_orig_lerobot |
Bridge (WidowX), original demos (LeRobot format) | IPEC-COMMUNITY/bridge_orig_lerobot |
| Model | Description | 🤗 Download Link |
|---|---|---|
gr15_base_bridge_v1 |
Base (no pruning), Bridge / SimplerEnv | binhng/gr15_base_bridge_v1 |
gr15_cka_bridge_v1 |
CLP-pruned (CKA), Bridge / SimplerEnv | binhng/gr15_cka_bridge_v1 |
If you find this work useful, please cite our paper:
@article{nguyen2026finetuning,
title={Finetuning Vision-Language-Action Models Requires Fewer Layers Than You Think},
author={Nguyen, Gia-Binh and Ho, Trong-Bao and Ha, Thien-Loc and Vo, Khoa and M{\o}ller, Philip Lund and Nguyen, Quang T and Dinh, Long and Dam, Tuan and Duong, Vu and Luu, Tung M and others},
journal={arXiv preprint arXiv:2606.20246},
year={2026}
}This repository builds on NVIDIA Isaac GR00T and lerobot, with evaluation environments from LIBERO, SimplerEnv, and RoboCasa. Thanks for their excellent open-source work!
