Skip to content

Latest commit

 

History

93 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

[CVPR 2026] Selective, Regularized, and Calibrated: Harnessing Vision Foundation Models for Cross-Domain Few-Shot Semantic Segmentation

Official implementation of our CVPR 2026 paper Selective, Regularized, and Calibrated: Harnessing Vision Foundation Models for Cross-Domain Few-Shot Semantic Segmentation

Paper | Project Page | Code

The proposed HERA harnesses vision foundation models for cross-domain few-shot semantic segmentation through selective feature extraction, regularized adaptation, and calibrated attention refinement. It dynamically identifies task-relevant representations from intermediate layers, integrates complementary multi-level features, and improves prediction reliability under severe domain shifts while requiring only a few annotated support exemplars.

Data Preparation

We evaluate HERA on the standard cross-domain few-shot semantic segmentation benchmarks.

The source-domain dataset follows the conventional CD-FSS setting, while HERA is evaluated directly on the target domains without source-data retraining.

Source Domain

▸ PASCAL VOC 2012

 PASCAL VOC 2012 is commonly used as the source-domain dataset in CD-FSS.

wget http://host.robots.ox.ac.uk/pascal/VOC/voc2012/VOCtrainval_11-May-2012.tar

Target Domains

▸ DeepGlobe

 DeepGlobe is a satellite-image segmentation dataset with substantial variations in texture, scale, and spatial layout.

▸ ISIC 2018

 ISIC 2018 contains dermoscopic skin-lesion images with irregular boundaries and low-contrast foreground regions.

▸ Chest X-ray

 The Chest X-ray contains radiographic images and lung masks with substantial grayscale and structural variations.

▸ FSS-1000

 FSS-1000 is a large-scale few-shot segmentation dataset containing 1,000 object categories from natural images.

Pretrained Models and Benchmark Results

Models

The default implementation uses DINOv3 ViT-L/16 (Download from Google Drive) as the vision foundation model.

After downloading, place the checkpoint in the checkpoints/ directory.

checkpoints/
└── dinov3_vitl16_pretrain_lvd1689m-8aa4cbdd.pth

Performance

The following results are reported using DINOv3 under the standard 1-shot and 5-shot CD-FSS evaluation protocols.

Target Dataset 1-Shot mIoU 5-Shot mIoU
DeepGlobe 44.6% 63.4%
ISIC 2018 61.2% 73.6%
Chest X-ray 85.8% 87.9%
FSS-1000 81.6% 86.7%
Average 68.3% 77.9%

Dataset Organization

After downloading and preprocessing the datasets, organize them using the following structure:

HERA-CDFSS/                                           # project root
|── codes/                                            # source code
├── data/                                             # datasets
│   ├── VOC2012/                                      # source dataset: PASCAL VOC 2012
│   │   ├── JPEGImages/
│   │   └── SegmentationClassAug/
│   │
│   ├── DeepGlobe/                                    # target dataset: DeepGlobe
│   │   ├── 01_train_ori/                             # original data
│   │   ├── ...
│   │   └── 04_train_cat/                             # processed data
│   │       ├── 1/                                    # category
│   │       │   └── test/
│   │       │       ├── origin/                       # images
│   │       │       └── groundtruth/                  # masks
│   │       └── ...
│   │
│   ├── ISIC/                                         # target dataset: ISIC 2018
│   │   ├── ISIC2018_Task1-2_Training_Input/          # images
│   │   │   ├── 1/                                    # category
│   │   │   └── ...
│   │   ├── ISIC2018_Task1_Training_GroundTruth/      # masks
│   │   └── class_id.csv
│   │
│   ├── LungSegmentation/                             # target dataset: Chest X-ray
│   │   ├── CXR_png/                                  # images
│   │   └── masks/                                    # masks
│   │
│   └── FSS-1000/                                     # target dataset: FSS-1000
│       ├── ab_wheel/                                 # category
│       └── ...
│
└── checkpoints/                                      # pretrained model checkpoints
    └── dinov3_vitl16_pretrain_lvd1689m-8aa4cbdd.pth

Environment Setup

To set up your environment, execute the following commands:

conda create -n hera python=3.10 -y
conda activate hera

pip install torch==2.8.0 torchvision==0.23.0 torchaudio==2.8.0
pip install scipy pandas matplotlib seaborn
pip install opencv-python scikit-image safetensors timm tensorflow tensorboardX

Run the Code

HERA follows a source-free test-time adaptation setting and does not require separate source-domain training.

Please ensure that the target dataset and DINOv3 checkpoint are properly prepared before running the code.

We use DeepGlobe as an example below. More evaluation commands are provided in scripts.sh.

Run the 1-shot evaluation on DeepGlobe:

CUDA_VISIBLE_DEVICES=0 python main_hera.py \
  --test_datapath ./data/deepglobe \
  --backbone DINOv3 \
  --benchmark deepglobe \
  --fold 0 \
  --nshot 1 \
  --refine always \
  --fusion on \
  --feat_id 12 13 14 15 16 17 18 19 20 21 22 23 \
  --attn_strategy dual_attn_gauss \
  --logdir ./logs/deepglobe \
  --logfile Dinov3_deepglobe_shot1.txt

Run the 5-shot evaluation on DeepGlobe:

CUDA_VISIBLE_DEVICES=0 python main_hera.py \
  --test_datapath ./data/deepglobe \
  --backbone DINOv3 \
  --benchmark deepglobe \
  --fold 0 \
  --nshot 5 \
  --refine auto \
  --fusion on \
  --feat_id 12 13 14 15 16 17 18 19 20 21 22 23 \
  --attn_strategy dual_attn_gauss \
  --logdir ./logs/deepglobe \
  --logfile Dinov3_deepglobe_shot5.txt

The same evaluation pipeline can be applied to other target datasets by updating --benchmark, --test_datapath, --logdir, and --logfile.

Evaluation performance may vary slightly across random seeds, GPU devices, software environments, and dataset preprocessing implementations.

HERA is designed as a general framework for vision foundation models (VFMs). The current repository provides the standard DINOv3 implementation as the default reference. To use other VFMs, please follow the code comments and update the model-specific components accordingly.

Citation

If you find HERA useful in your research, please cite our paper:

@inproceedings{ma2026selective,
  title     = {Selective, Regularized, and Calibrated: Harnessing Vision Foundation Models for Cross-Domain Few-Shot Semantic Segmentation},
  author    = {Ma, Junyuan and Xiang, Xunzhi and Li, Wenbin and Fan, Qi and Gao, Yang},
  booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
  pages     = {12385--12395},
  year      = {2026}
}

Acknowledgement

Our codebase is built upon the official implementations of DR-Adapter and SSP. We sincerely thank the authors for releasing their valuable code and providing a solid foundation for the development of this project.

We also thank PATNet and other FSS and CD-FSS works for their valuable contributions to this research community.

About

[CVPR 2026] Official repository for "Selective, Regularized, and Calibrated: Harnessing Vision Foundation Models for Cross-Domain Few-Shot Semantic Segmentation".

Topics

Resources

Stars

7 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages