JQ Huang2β Rakesh Ranjan2β Aviral Chharia1,β β Fernando De la Torre1
Achieve 93.5% of full-token performance with just 8% of the tokens! We present CoVeR - a deterministic, training-free token pruner for multi-view inputs to 2D VLMs that selects a subset of visual tokens collectively covering the scene.
- β Coming Soon: Full CoVeR Codebase. Stay Tuned!
- β Sep. 8, 2026: We released the CoVeR on arXiv. Check the preprint!
Representing a 3D scene as multi-view images allows 2D VLMs to reason in 3D by reusing priors from pre-training, sidestepping the scarcity of annotated 3D data. However, it produces thousands of redundant visual tokens whose cost grows with every view. Existing visual token pruners fall into two families, each limited in the 3D multi-view setting. Learned importance methods rank tokens by attention or encoder features; because redundancy here is fundamentally spatial, they keep near-duplicate tokens from a few prominent regions and leave most of the scene unrepresented. Voxelization methods improve spatial coverage but cannot enforce an exact token budget and saturate as multi-view observations overlap in 3D, capping retention well below the target. We show that spatial coverage is associated with 3D reasoning performance and introduce CoVeR, a deterministic, training-free selector that uses only token coordinates, with no learned signals. CoVeR selects tokens that collectively cover every region of the scene, and solves the limitations of both families: it enforces an exact per-scene budget, breaks the voxelization saturation plateau, and avoids the near-duplicate selections of learned importance. Extensive experiments show CoVeR outperforms prior SOTAs on all three 3D reasoning benchmarks and generalizes as a plug-and-play module tested across four VLMs. Notably, with only β8% of visual tokens, it preserves 93.5% of full-token performance, surpassing SOTA by 3.9 percentage points on average across benchmarks.
If you find our work useful for your project, please consider adding a star to this repo and citing our paper:
@misc{bui2026covercoveragebasedtokenpruning,
title={CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs},
author={Nhat-Tan Bui and Varshini Elangovan and Arun Reddy Anugu and Sreyas Mohan and Wei Ye and Dilin Wang and JQ Huang and Rakesh Ranjan and Aviral Chharia and Fernando De la Torre},
year={2026},
eprint={2609.08345},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2609.08345},
}