SenseNova-Vision training combines original images from public datasets with
training JSONL annotations and derived assets released in
SenseNova-Vision-Corpus-50M.
This document describes the repository layout expected by
data/dataset_info.py and organizes data preparation by training task.
Run all commands from the repository root. Review and follow the license, terms of use, and citation requirements of every original dataset.
The training configuration uses the following repository-local roots:
jsonl_generate/train_jsonls/ # training JSONL annotations
datas/ # original public data and converted assets
datas/train_data/ # source images for released training annotations
datas/SenseNova-Vision-Corpus-50M/ # released derived assets
Large datasets may be stored outside the repository. Symbolic links are recommended as long as these paths are present from the repository root.
After preparing and activating the project environment with setup.sh,
download the dataset repository with its pinned huggingface_hub package:
CORPUS_RELEASE_DIR=/absolute/path/SenseNova-Vision-Corpus-50M-release
python -m huggingface_hub.commands.huggingface_cli download \
sensenova/SenseNova-Vision-Corpus-50M \
--repo-type dataset \
--local-dir "$CORPUS_RELEASE_DIR"Copy the four released training annotation directories and the image-generation
key lists into jsonl_generate/train_jsonls/. This preserves any benchmark
annotations already prepared under jsonl_generate/:
mkdir -p jsonl_generate/train_jsonls
cp -r "$CORPUS_RELEASE_DIR/dense_geometric_prediction" \
jsonl_generate/train_jsonls/
cp -r "$CORPUS_RELEASE_DIR/multiview_visual_geometry" \
jsonl_generate/train_jsonls/
cp -r "$CORPUS_RELEASE_DIR/segmentation" \
jsonl_generate/train_jsonls/
cp -r "$CORPUS_RELEASE_DIR/structure_view_understanding" \
jsonl_generate/train_jsonls/
mkdir -p jsonl_generate/train_jsonls/image_generation/keys
cp "$CORPUS_RELEASE_DIR"/image_generation/keys/*.keys.txt.gz \
jsonl_generate/train_jsonls/image_generation/keys/The corpus assets are distributed as ordered
SenseNova-Vision-Corpus-50M.tar.gzNNN shards. Join and extract them:
CORPUS_ASSET_PARENT=/absolute/path/sensenova-corpus-assets
mkdir -p "$CORPUS_ASSET_PARENT"
cat "$CORPUS_RELEASE_DIR"/SenseNova-Vision-Corpus-50M.tar.gz* | \
tar -xz -C "$CORPUS_ASSET_PARENT"
mkdir -p datas
ln -s "$CORPUS_ASSET_PARENT/SenseNova-Vision-Corpus-50M" \
datas/SenseNova-Vision-Corpus-50MThe resulting release layout is:
jsonl_generate/train_jsonls/
├── dense_geometric_prediction/
├── multiview_visual_geometry/
├── segmentation/
├── structure_view_understanding/
└── image_generation/
└── keys/
├── BLIP3o-Pretrain-Long-Caption.keys.txt.gz
├── BLIP3o-Pretrain-Long-Caption-part2.keys.txt.gz
└── BLIP3o-Pretrain-Short-Caption.keys.txt.gz
datas/SenseNova-Vision-Corpus-50M/
├── coco2017/
├── object365/
├── sa_1b/
├── scannetpp/
├── DL3DV/
├── wild_rgbd/
├── Cityscapes/
├── coconut/
└── ...
The release does not duplicate original source images. Its source-image links are listed in the corpus download references. The full SN-VC source inventory is defined by the report sections linked below.
Corpus details: Section 7.3, Tables 12–13.
The released segmentation JSONL files load original images from
datas/train_data/ and datas/gcg_seg_data/, while their target assets are
loaded from datas/SenseNova-Vision-Corpus-50M/. The benchmark-compatible
COCO 2014 and COCO 2017 layouts remain under datas/gen_seg_data/.
| Source dataset | Download | Target path |
|---|---|---|
| COCONut-XL | COCONut | COCO images under datas/train_data/coco2017/ |
| Objects365 | Objects365 | datas/train_data/object365/ |
| Cityscapes | cityscapesScripts | datas/train_data/Cityscapes/ |
| Hypersim | Hypersim | datas/train_data/Hypersim/ |
| EntityV2 | EntitySeg | datas/train_data/Entityv2/images/ |
| Trashcan | TrashCan download | datas/train_data/TrashCan/dataset/instance_version/ |
| Pidray | PIDRay | datas/train_data/PIDRay/ |
| ZeroWaste-f | ZeroWaste | datas/train_data/ZeroWaste-f/ |
| LVIS | LVIS | COCO images under datas/train_data/coco2017/ |
| IDD-1/2 | IDD | datas/train_data/IDD/IDD_Segmentation/ and datas/train_data/IDD/idd20kII/ |
| IDDAv3 | IDDA | datas/train_data/IDDA/IDDAv3/ |
| Mapillary Vistas | Mapillary Vistas | datas/train_data/MapillaryVistas/ |
| NuScenes | nuScenes devkit | datas/train_data/nuImages/samples/ |
| 51WORLD | DataOne sample | datas/train_data/51WORLD/train/ |
| StreetHazards | StreetHazards | datas/train_data/StreetHazards/train/images/ |
| KITTI | KITTI | datas/train_data/KITTI/ |
| TAS500 | TAS500 download | datas/train_data/TAS500/ |
| UDD5/6 | UDD | datas/train_data/UDD/UDD5/ and datas/train_data/UDD/UDD6/ |
| TTPLA | TTPLA | datas/train_data/TTPLA/ |
| LoveDA | LoveDA preparation | datas/train_data/LoveDA/ |
| VIPSeg | VIPSeg | datas/train_data/VIPSeg/imgs/ |
| GranDf | GroundingLMM data | datas/gcg_seg_data/images/GranDf_HA_images/ |
| RefCOCOg | REFER | datas/gcg_seg_data/images/coco2014/ |
| PSG | OpenPSG | datas/gcg_seg_data/images/coco2017/ |
| Flickr30k | Flickr30k Entities | datas/gcg_seg_data/images/flickr30k/ |
The segmentation layout follows the public preparation conventions documented
by X-SAM.
The benchmark preparation downloads COCO 2017 validation images but not the
training images required here. Download and extract train2017 into the same
COCO directory:
mkdir -p datas/gen_seg_data/coco2017
COCO17_DIR=datas/gen_seg_data/coco2017
wget -c http://images.cocodataset.org/zips/train2017.zip \
-O "$COCO17_DIR/train2017.zip"
unzip "$COCO17_DIR/train2017.zip" -d "$COCO17_DIR"
rm "$COCO17_DIR/train2017.zip"
unset COCO17_DIROnly the original training images are needed; the derived training targets are provided by SenseNova-Vision-Corpus-50M. Reuse common image collections with symbolic links where necessary:
mkdir -p datas/train_data datas/gen_seg_data datas/gcg_seg_data/images
ln -s ../gen_seg_data/coco2017 datas/train_data/coco2017
ln -s ../../gen_seg_data/coco2014 datas/gcg_seg_data/images/coco2014
ln -s ../../gen_seg_data/coco2017 datas/gcg_seg_data/images/coco2017Prepare the public binary-mask datasets listed in Section 7.3, Table 12 as follows.
Generate the following JSONL files under
jsonl_generate/train_jsonls/segmentation/.
| Dataset entry | JSONL file | Required media directories |
|---|---|---|
refcoco_train |
seg_refcoco_train_binary.jsonl |
datas/ref_seg_data/images/coco2014/train2014/, datas/ref_seg_data/ref_seg/binary_masks/refcoco_train/ |
refcoc+_train |
seg_refcoco+_train_binary.jsonl |
datas/ref_seg_data/images/coco2014/train2014/, datas/ref_seg_data/ref_seg/binary_masks/refcoco+_train/ |
refcocog_train |
seg_refcocog_train_binary.jsonl |
datas/ref_seg_data/images/coco2014/train2014/, datas/ref_seg_data/ref_seg/binary_masks/refcocog_train/ |
refclef_train |
seg_refclef_train_binary.jsonl |
datas/ref_seg_data/images/saiapr_tc-12/, datas/ref_seg_data/ref_seg/binary_masks/refclef_train/ |
grefcoco_train |
seg_grefcoco_train_binary.jsonl |
datas/ref_seg_data/images/coco2014/train2014/, datas/ref_seg_data/ref_seg/binary_masks/grefcoco_train/ |
rea_train |
seg_reason_train_repeat100.jsonl |
datas/rea_seg_data/train/, datas/rea_seg_data/rea_seg/binary_masks/train/ |
coco_interactive_psalm |
seg_coco_interactive_psalm.jsonl |
datas/gen_seg_data/coco2017/train2017/, datas/inter_seg_data/inter_seg/binary_masks/coco_interactive_psalm/train/ |
Reuse the COCO 2014 images, RefCOCO-family annotations, and COCO-Interactive
annotations prepared in docs/data_prepare.md. Check them
before downloading:
bash tools/data_prepare/segmentation/check_reusable_data.sh \
datas/gen_seg_data/coco2014/train2014 \
datas/gen_seg_data/coco2014/train2014.zip \
'datas/ref_seg_data/refcoco/refs(unc).p' \
datas/ref_seg_data/refcoco.zip \
'datas/ref_seg_data/refcoco+/refs(unc).p' \
datas/ref_seg_data/refcoco+.zip \
'datas/ref_seg_data/refcocog/refs(umd).p' \
datas/ref_seg_data/refcocog.zip \
datas/inter_seg_data/annotations/coco_interactive_train_psalm.json \
datas/inter_seg_data/PSALM_data.zip[READY] means the extracted data can be reused. [ARCHIVE READY] means a
valid ZIP is available for extraction. For missing COCO or RefCOCO-family data,
follow the corresponding instructions in
docs/data_prepare.md.
Download the remaining train-only sources following the public X-SAM preparation. RefCLEF requires both the REFER annotations and the ReferItGame image subset:
REF_DIR=datas/ref_seg_data
mkdir -p "$REF_DIR/images"
if [ ! -f "$REF_DIR/refclef/refs(unc).p" ]; then
wget -c \
https://web.archive.org/web/20220413011631/https://bvisionweb1.cs.unc.edu/licheng/referit/data/refclef.zip \
-O "$REF_DIR/refclef.zip"
unzip "$REF_DIR/refclef.zip" -d "$REF_DIR"
fi
if [ ! -d "$REF_DIR/images/saiapr_tc-12" ]; then
wget -c \
https://web.archive.org/web/20220413011744/http://bvisionweb1.cs.unc.edu/licheng/referit/data/images/saiapr_tc-12.zip \
-O "$REF_DIR/saiapr_tc-12.zip"
unzip "$REF_DIR/saiapr_tc-12.zip" -d "$REF_DIR/images"
fiDownload the official gRefCOCO annotations directly into the directory used by the converter:
if [ ! -f datas/ref_seg_data/grefcoco/instances.json ] || \
[ ! -f 'datas/ref_seg_data/grefcoco/grefs(unc).json' ]; then
python -m huggingface_hub.commands.huggingface_cli download \
FudanCVL/gRefCOCO \
--repo-type dataset \
--local-dir datas/ref_seg_data/grefcoco
fiDownload train.zip and train.json from the
LISA/ReasonSeg release
to datas/rea_seg_data/. Reuse val.zip and test.zip from benchmark
preparation; only the train split is added here:
REA_DIR=datas/rea_seg_data
mkdir -p "$REA_DIR/explanatory"
[ -d "$REA_DIR/train" ] || unzip "$REA_DIR/train.zip" -d "$REA_DIR"
[ -f "$REA_DIR/explanatory/train.json" ] || \
mv "$REA_DIR/train.json" "$REA_DIR/explanatory/train.json"The benchmark download of PSALM_data.zip already contains both train and val
annotations. If the archive exists but the train annotation has not been
extracted, recover only that JSON file:
INTER_DIR=datas/inter_seg_data
mkdir -p "$INTER_DIR/annotations"
if [ ! -f "$INTER_DIR/annotations/coco_interactive_train_psalm.json" ]; then
unzip -j "$INTER_DIR/PSALM_data.zip" \
'*/coco_interactive_train_psalm.json' \
-d "$INTER_DIR/annotations"
fiCreate shared image links without duplicating COCO images:
mkdir -p datas/ref_seg_data/images datas/inter_seg_data
[ -e datas/ref_seg_data/images/coco2014 ] || \
ln -s ../../gen_seg_data/coco2014 datas/ref_seg_data/images/coco2014
[ -e datas/inter_seg_data/coco2017 ] || \
ln -s ../gen_seg_data/coco2017 datas/inter_seg_data/coco2017The resulting layout should include:
datas/
├── gen_seg_data/
│ ├── coco2014/train2014/
│ └── coco2017/train2017/
├── ref_seg_data/
│ ├── grefcoco/
│ ├── images/
│ │ ├── coco2014/
│ │ └── saiapr_tc-12/
│ ├── refclef/
│ ├── refcoco/
│ ├── refcoco+/
│ ├── refcocog/
│ └── ref_seg/binary_masks/
├── rea_seg_data/
│ ├── explanatory/
│ ├── train/
│ └── rea_seg/binary_masks/train/
└── inter_seg_data/
├── annotations/
├── coco2017/
└── inter_seg/binary_masks/coco_interactive_psalm/train/
Run the converters after preparing the source data:
python tools/data_prepare/segmentation/prepare_binary.py refcoco
python tools/data_prepare/segmentation/prepare_binary.py reasonseg
python tools/data_prepare/segmentation/prepare_binary.py coco-interactiveThe commands write training JSONL files to
jsonl_generate/train_jsonls/segmentation/. Benchmark and test JSONL files
remain under jsonl_generate/.
Download each dataset to the target directory listed below.
| Source dataset | Download | Target path |
|---|---|---|
| DOORS | Zenodo | datas/ref_seg_data/DOORS/ |
| NDISPark | Zenodo | datas/ref_seg_data/NDISPark/ |
| MinneApple | Project repository | datas/ref_seg_data/MinneApple/ |
| EYTH | EgoYouTubeHands project | datas/ref_seg_data/EYTH/ |
| PST900 | Project repository | datas/ref_seg_data/PST900/ |
| PSTRGB | Project repository | datas/ref_seg_data/PSTRGB/ |
| SUIM | UMN IRVLab | datas/ref_seg_data/SUIM/ |
| MyFood | AIcrowd Food Recognition Challenge | datas/ref_seg_data/MyFood/ |
| CO-SKEL | Project repository | datas/ref_seg_data/CO-SKEL/ |
YouTube VOS 2022 (VIS2022) |
YouTube-VIS data page | datas/ref_seg_data/VIS2022/ |
MVTec D2S (MVTecD2S) |
MVTec dataset page | datas/ref_seg_data/MVTecD2S/ |
| VizWiz-FewShot | VizWiz download page | datas/ref_seg_data/VizWiz-FewShot/ |
| Trans10K | Trans10K-v1 project | datas/ref_seg_data/Trans10K/ |
| CIHP | LIP Challenge | datas/ref_seg_data/CIHP/ |
| ATR | HumanParsing-Dataset repository | datas/ref_seg_data/ATR/ |
| LIP | SYSU-HCP dataset page | datas/ref_seg_data/LIP/ |
| FAT-single / FAT-mixed | NVIDIA Falling Things | datas/ref_seg_data/FAT/ |
| Fashionpedia | Official download page | datas/ref_seg_data/Fashionpedia/ |
| PartImageNet / PartImageNet-Whole | Project repository | datas/ref_seg_data/PartImageNet/ |
| WaterOVS | Hugging Face dataset | datas/ref_seg_data/WaterOVS/ |
| RaidaR-rainy / RaidaR-sunny | Official download page | datas/ref_seg_data/RaidaR/ |
| FSS-1000 | Project repository | datas/ref_seg_data/FSS-1000/ |
DAVIS 2017 (DAVIS) |
Official download page | datas/ref_seg_data/DAVIS/ |
OCID-VLG (OCID) |
Project repository | datas/ref_seg_data/OCID-VLG/ |
| PIC | IEEE Person in Context Challenge | datas/ref_seg_data/PIC/ |
| LaPa | Official repository | datas/ref_seg_data/LaPa/ |
| DeepFashion2 | Official repository | datas/ref_seg_data/DeepFashion2/ |
| MattingHumanHalf | Matting Human Datasets | datas/ref_seg_data/MattingHumanHalf/ |
Follow each source's access and license requirements.
Download DOORS v1.0 from Zenodo. The archive is CC BY 4.0.
mkdir -p datas/ref_seg_data
if [ ! -f datas/ref_seg_data/DOORS.zip ]; then
wget -c 'https://zenodo.org/records/7107409/files/DOORS.zip?download=1' \
-O datas/ref_seg_data/DOORS.zip
fi
if [ ! -d datas/ref_seg_data/DOORS/Segmentation/DS1/DS ]; then
unzip datas/ref_seg_data/DOORS.zip -d datas/ref_seg_data
fi
python tools/data_prepare/segmentation/prepare_binary.py doorsFor YouTube-VIS 2022, accept the official terms of use, then download the 2022 training images and annotations from the YouTube-VIS data page. The annotations are CC BY 4.0; the dataset is limited to non-commercial research use.
Use this layout:
datas/ref_seg_data/VIS2022/
└── train/
├── instances.json
└── JPEGImages/
└── <video_id>/
└── <frame>.jpg
Run:
python tools/data_prepare/segmentation/prepare_binary.py vis2022 --num-workers 8Corpus details: Section 7.2, Table 10.
Reuse Hypersim, COCO 2017, and Objects365 from Segmentation. Download the remaining sources.
| Source dataset | Download | Target path |
|---|---|---|
| Hypersim | Hypersim | datas/train_data/Hypersim/ |
| Virtual KITTI | Virtual KITTI 1.3.1 | datas/train_data/vkitti_depth/ |
| InteriorVerse | InteriorVerse | datas/train_data/InteriorVerse_85/ |
| IRS | IRS dataset | datas/train_data/IRS/ |
| TartanAir | TartanAir | datas/train_data/tartanair/ |
| SceneNet RGB-D | SceneNet RGB-D | datas/train_data/ScenenetRGBD/ |
| Taskonomy | Taskonomy | datas/train_data/taskonomy/ |
| ScanNet++ | ScanNet++ | datas/train_data/scannetpp/ |
| COCO 2017 | COCO downloads | datas/train_data/coco2017/ |
| SA-1B | Segment Anything | datas/train_data/sa_1b/ |
| Objects365 | Objects365 | datas/train_data/object365/ |
Use the iPhone-captured ScanNet++ data.
The corpus download steps above install the released depth and normal JSONLs
under jsonl_generate/train_jsonls/dense_geometric_prediction/. Download the
corresponding source images listed above; the generated targets are loaded from
datas/SenseNova-Vision-Corpus-50M/.
| Source dataset | Dataset entries |
|---|---|
| COCO 2017 | coco_depth, coco_normal |
| Objects365 | object365_depth, object365_normal |
| ScanNet++ | scannetpp_depth, scannetpp_normal |
| SA-1B | sa_1b_depth, SA_1B_normal |
| Taskonomy | taskonomy_depth, taskonomy_normal |
For public raw-data conversion, see the Dense Geometric Prediction conversion guide.
Corpus details: Section 7.4, Table 16.
Reuse Hypersim, IRS, TartanAir, SceneNet RGB-D, and ScanNet++ from Dense Geometric Prediction. Download the remaining sources.
| Source dataset | Download | Target path |
|---|---|---|
| Hypersim | Hypersim | datas/train_data/Hypersim/ |
| IRS | IRS dataset | datas/train_data/IRS/ |
| TartanAir | TartanAir | datas/train_data/tartanair/ |
| SceneNet RGB-D | SceneNet RGB-D | datas/train_data/ScenenetRGBD/ |
| AriaSyntheticENV | Aria Synthetic Environments | datas/train_data/AriaSyntheticEnvironment/ |
| BlendedMVG | BlendedMVS and BlendedMVG | datas/train_data/BlendedMVG/ |
| MegaSynth | MegaSynth | datas/train_data/MegaSynth/ |
| MvsSynth | MVS-Synth | datas/train_data/MVS-Synth/ |
| OmniObject3D | OmniObject3D | datas/train_data/OmniObject3D/ |
| Objaverse | Objaverse | datas/train_data/objaverse_v1/ |
| CO3Dv2 | CO3Dv2 | datas/train_data/CO3Dv2/ |
| DeMoN-MVE | DeMoN datasets | datas/train_data/demon-mve/ |
| ScanNetV2 | OpenDataLab | datas/train_data/scannetv2/ |
| ScanNet++ | ScanNet++ | datas/train_data/scannetpp/ |
| DL3DV | DL3DV-10K | datas/train_data/DL3DV/ALL-960P/ |
| WildRGB-D | WildRGB-D | datas/train_data/wild_rgbd/ |
The released reconstruction JSONLs use DL3DV, ScanNet++, ScanNetV2, and
WildRGB-D. Use the 960P resolution version for DL3DV, and extract ScanNetV2
scenes from their .sens files with the ScanNet tools.
Camera-pose JSONLs for the other sources can be generated with the
multi-view conversion guide.
Corpus details: Section 7.1, Tables 6–7.
Reuse SA-1B, Objects365, and COCO 2017 from Dense Geometric Prediction, and reuse NuImages and the RefCOCO family from Segmentation. Download the remaining sources.
| Source dataset | Download | Target path |
|---|---|---|
| APTv2 | APTv2 | datas/train_data/APTv2/ |
| BDD100K | BDD100K | datas/train_data/BDD100K/ |
| Blood Cell | Roboflow | datas/train_data/Blood Cell Detection/ |
| CARPK | CARPK project | datas/train_data/CARPK/ |
| CrowdHuman | Official download page | datas/train_data/CrowdHuman/ |
| DOTAv2 | DOTA | datas/train_data/DOTAv2/ |
| DeepFashion | DeepFashion-MultiModal | datas/train_data/DeepFashion/ |
| EgoObjects | EgoObjects | datas/train_data/EgoObjects/ |
| FAIR1M | FAIR1M 2.0 | datas/train_data/FAIR1M/ |
| FSC147 | Learning to Count Everything | datas/train_data/FSC147/ |
| FiftyOne | Hugging Face | datas/train_data/dense_object_detection_FiftyOne/ |
| Fish | Roboflow | datas/train_data/fish-detection-dataset/ |
| Football | Roboflow | datas/train_data/football-object-detection/ |
| GroceryStore | GroceryStoreDataset | datas/train_data/GroceryStore/ |
| HomeObjects-3k | Ultralytics | datas/train_data/homeobjects-3K/ |
| HumanParts | Human-Parts | datas/train_data/HumanParts/ |
| ImageNetPart | PartImageNet | datas/train_data/ImageNetPart/ |
| Industrial Site Safety | Hugging Face | datas/train_data/Industrial-Site-Safety-Detection-v1-DATASET/ |
| LVIS Fruits & Vegetables | Hugging Face | datas/train_data/LVIS_Fruits_And_Vegetables/ |
| Locount | Dataset repository | datas/train_data/Locount/ |
| METU-ALET | Dataset repository | datas/train_data/METU-ALET/ |
| NuImages | nuImages | datas/train_data/nuImages/ |
| OWOD | Open World Dense Object Detection | datas/train_data/owdod/ |
| Objects365 | Objects365 | datas/train_data/object365/ |
| PACO-LVIS | PACO annotations and COCO images | datas/train_data/PACO/ |
| PixMo-Points | PixMo-Points | datas/train_data/pixmo/ |
| S2TLD | Dataset repository | datas/train_data/S2TLD/ |
| SA-1B | Segment Anything | datas/train_data/sa_1b/ |
| SKU110K | Official repository | datas/train_data/SKU110k/ |
| Shoes | Kaggle | datas/train_data/Shoes_data/ |
| TinyPerson | TinyBenchmark | datas/train_data/TinyPerson/ |
| V3Det-OVD | V3Det | datas/train_data/V3Det___V3Det/raw/ |
| VisDrone | VisDrone | datas/train_data/VisDrone/ |
| WiderPerson | Dataset index | datas/train_data/WiderPerson/ |
| Pill | Medical Pills | datas/train_data/Medical-pills/ |
| Sheep | Aerial Sheep | datas/train_data/aerial-sheep-object-detection/ |
| HumanRef | Hugging Face | datas/train_data/humanref_cot_45k_converted/ |
| OpenImages | Open Images V7 | datas/train_data/openimages/ |
| RefCOCO/+/g | REFER and COCO 2014 | datas/ref_seg_data/ |
| RexVerse | RexVerse-2M | datas/train_data/RexVerse-2M/ |
| BLIP3-OCR-200M | Hugging Face | datas/train_data/OCR/blip3-ocr-200m/ |
| HierText | Dataset repository | datas/train_data/OCR/Hiertext/ |
| ICDAR2013 | ICDAR 2013 Robust Reading | datas/train_data/OCR/icdar2013/ |
| ICDAR2015 | ICDAR 2015 Incidental Scene Text | datas/train_data/OCR/icdar2015/ |
| ICDAR2019 | ICDAR 2019 ArT | datas/train_data/OCR/icdar2019/ |
| LSVT2019 | ICDAR 2019 LSVT | datas/train_data/OCR/LSVT2019/ |
| MTWI | Tianchi | datas/train_data/OCR/mtwi/ |
| RCTW | RCTW-17 | datas/train_data/OCR/RCTW/ |
| ReCTS | ICDAR 2019 ReCTS | datas/train_data/OCR/ReCTS/ |
| SROIE | SROIE 2019 | datas/train_data/sroie-datasetv2/ |
| SynthText | SynthText in the Wild | datas/train_data/OCR/SynthText/ |
| TextOCR | Official dataset page | datas/train_data/OCR/TextOCR/ |
| WildReceipt | OpenMMLab archive | datas/train_data/OCR/wildreceipt/ |
| AP-10K | AP-10K | datas/train_data/keypoints/ap-10k/ |
| APT36K | APT-36K | datas/train_data/keypoints/APT36k/ |
| COCO2017 | COCO | datas/train_data/coco2017/ |
| CrowdPose | CrowdPose | datas/train_data/keypoints/crowdpose/ |
| Human-Art | Human-Art | datas/train_data/keypoints/Human-Art/ |
| MPII | MPII Human Pose | datas/train_data/keypoints/mpii/ |
| MacaquePose V1 | MacaquePose | datas/train_data/keypoints/macaquepose_v1/ |
| OCHuman | OCHuman | datas/train_data/keypoints/ochuman/ |
| CDLA | CDLA | datas/train_data/Layout/CDLA_DATASET/ |
| DocLayNet Core | DocLayNet | datas/train_data/Layout/DocLayNet_core/ |
| PubLayNet | Official repository | datas/train_data/Layout/publaynet/ |
| TabRecSet | TabRecSet | datas/train_data/Layout/TabRecSet/ |
| TableBank | TableBank | datas/train_data/Layout/TableBank/ |
| OS-Atlas | OS-Atlas-data | datas/train_data/GUI/OS-Atlas-data/ |
| ShowUI Desktop | ShowUI-desktop | datas/train_data/GUI/ShowUI-desktop/ |
Apply these dataset-specific adjustments after downloading the source data.
Objects365: keep the annotation Parquet shards under
datas/train_data/object365/data/ and the training images under
datas/train_data/object365/patch*/.
PACO-LVIS: place paco_lvis_v1_train.json in datas/train_data/PACO/. From
that directory, link the COCO 2017 images:
ln -s ../coco2017/train2017 train2017EgoObjects: use EgoObjectsV1_images.zip and
EgoObjectsV1_unified_train.json.
APTv2: extract both archives so that datas/train_data/APTv2/ contains
annotations/train_annotations.json, data/easy/, and data/hard/.
OCR datasets: keep the released folder names from the original downloads.
- SROIE:
extract to
datas/train_data/sroie-datasetv2/versions/4/SROIE2019/train/so bothbox/andimg/are present. - ICDAR2013:
place the training images in
datas/train_data/OCR/icdar2013/Challenge2_Training_Task12_Images/and the word annotations indatas/train_data/OCR/icdar2013/Challenge2_Training_Task1_GT/. - ICDAR2015:
place the training images in
datas/train_data/OCR/icdar2015/ch4_training_images/and the word annotations indatas/train_data/OCR/icdar2015/ch4_training_localization_transcription_gt/.
Generated JSONLs are written to
jsonl_generate/train_jsonls/structure_view_understanding/.
Copy the published
structure_view_understanding
JSONLs to jsonl_generate/train_jsonls/structure_view_understanding/. Download
the source images listed above; do not run a converter for these entries.
| Source dataset | Dataset entries |
|---|---|
| SA-1B | grounding_SA1B, SA1B_pointing, SA_1B_visual |
| Objects365 | Objects365_pointing, object365_refbbox_merge, object365_refpoint_merge |
| OpenImages | openimages_refbbox_merge, openimages_refpoint_merge |
| APTv2 | APT_pointing |
| DeepFashion | DeepFashion_pointing |
| EgoObjects | EgoObjects_pointing |
| HumanParts | HumanParts_pointing |
| ImageNetPart | ImageNetPart_pointing |
| PACO-LVIS | PACO_LVIS_pointing |
| V3Det-OVD | V3Det_ovd_pointing |
| BDD100K, DOTAv2, FAIR1M, NuImages, VisDrone | BDD100K_pointing, DOTAv2_pointing, FAIR1M_pointing, NuImages_pointing, VisDrone_pointing |
| PixMo-Points | pixmo_detect |
| GroceryStore | GroceryStore_detect, GroceryStore_visual |
| FSC147 | FSC147_detect, FSC147_visual |
| BLIP3-OCR-200M | blip3_ocr_200m_text_bbox, blip3_ocr_200m_text_poly_OCR, blip3_ocr_200m_word_bbox_OCR, blip3_ocr_200m_word_poly_OCR |
These datasets are converted directly from the downloaded annotations without an intermediate JSONL.
| Source dataset | Dataset entries | Workflow cases |
|---|---|---|
| APTv2 | APT_detect |
aptv2 |
| Blood Cell | blood_cell_detect, blood_cell_visual |
blood-cell-bbox, blood-cell-visual |
| FSC147 | FSC147_pointing |
fsc147 |
| COCO 2017 | coco2017_keypoint |
coco2017-keypoint |
| CrowdPose | crowdpose_keypoint |
crowdpose-keypoint |
The commands below cover every workflow case in this table. Run only the cases for the datasets you prepared.
bash tools/data_prepare/structured_visual_understanding/prepare.sh aptv2
bash tools/data_prepare/structured_visual_understanding/prepare.sh blood-cell-bbox blood-cell-visual
bash tools/data_prepare/structured_visual_understanding/prepare.sh fsc147
bash tools/data_prepare/structured_visual_understanding/prepare.sh coco2017-keypoint
bash tools/data_prepare/structured_visual_understanding/prepare.sh crowdpose-keypointThe workflow first converts the raw annotations to a common bbox, keypoint, or OCR JSONL, then creates the final training JSONL.
| Source dataset | Dataset entries | Workflow cases |
|---|---|---|
| Objects365 | Objects365_detect, Objects365_visual |
objects365-bbox, objects365-visual |
| PACO-LVIS | PACO_LVIS_detect |
paco-lvis-bbox |
| SKU110K | SKU110k_detect, SKU110k_visual |
sku110k-bbox, sku110k-visual |
| EgoObjects | EgoObjects_detect |
egoobjects-bbox |
| V3Det-OVD | V3Det_ovd_detect |
v3det-ovd-bbox |
| OWOD | owdod_detect, owdod_visual |
owdod-bbox, owdod-visual |
| LVIS Fruits & Vegetables | LVIS_Fruits_And_Vegetables_detect, LVIS_Fruits_And_Vegetables_visual |
lvis-fruit-vegetable-bbox, lvis-fruit-vegetable-visual |
| Sheep | sheep_detect, sheep_visual |
sheep-bbox, sheep-visual |
| Football | football_detect, football_visual |
football-bbox, football-visual |
| Industrial Site Safety | Industrial_Site_Safety_detect, Industrial_Site_Safety_visual |
industrial-safety-bbox, industrial-safety-visual |
| AP-10K | ap-10k_keypoint |
ap-10k-keypoint |
| APT36K | APT36k_keypoint |
apt36k-keypoint |
| Human-Art | Human-Art_keypoint |
human-art-keypoint |
| MacaquePose V1 | macaquepose_v1_keypoint |
macaquepose-keypoint |
| MPII | mpii_keypoint |
mpii-keypoint |
| OCHuman | ochuman_keypoint |
ochuman-keypoint |
| SROIE | SROIE_text_bbox_OCR |
sroie |
| ICDAR2013 | icdar2013_word_bbox_OCR |
icdar2013-word-bbox |
| ICDAR2015 | icdar2015_word_bbox_OCR, icdar2015_word_poly_OCR |
icdar2015-word-bbox, icdar2015-word-poly |
| CDLA | CDLA_Layout |
cdla-layout |
| DocLayNet Core | DocLayNet_core_Layout |
doclaynet-core-layout |
| TableBank | TableBank_Layout |
tablebank-layout |
| TabRecSet | TabRecSet_Layout |
tabrecset-layout |
| OS-Atlas | OS-Atlas-data_desktop_domain_GUI, OS-Atlas-data_mobile_domain_GUI, OS-Atlas-data_rico_GUI, OS-Atlas-data_web_domain_GUI |
os-atlas-desktop-gui, os-atlas-mobile-gui, os-atlas-rico-gui, os-atlas-web-gui |
| ShowUI Desktop | ShowUI-desktop_GUI |
showui-desktop-gui |
The commands below are representative bbox/visual, keypoint, layout, and GUI examples. For other datasets, use the workflow case listed in the table.
bash tools/data_prepare/structured_visual_understanding/prepare.sh objects365-bbox objects365-visual
bash tools/data_prepare/structured_visual_understanding/prepare.sh owdod-bbox owdod-visual
bash tools/data_prepare/structured_visual_understanding/prepare.sh ap-10k-keypoint
bash tools/data_prepare/structured_visual_understanding/prepare.sh sroie icdar2013-word-bbox icdar2015-word-bbox icdar2015-word-poly
bash tools/data_prepare/structured_visual_understanding/prepare.sh cdla-layout
bash tools/data_prepare/structured_visual_understanding/prepare.sh os-atlas-rico-guiThese auxiliary multimodal sources belong to the training mixture described in Section 4; they are not part of the SN-VC source tables in Appendix Section 7.
| Task | Dataset entries | Official source | Local source path | Preparation |
|---|---|---|---|---|
| Understanding | llava_v1_5 |
LLaVA-Instruct-150K, LLaVA data guide | datas/train_data/llava_images |
Supported converter |
| Understanding | finevision_image, finevision_multi_image, finevision_text |
FineVision | datas/train_data/finevision_source |
Official source only |
| Understanding | mammoth_image, mammoth_text |
MAmmoTH-VL-Instruct-12M | datas/train_data/mammoth_vl_source |
Official source only |
| Generation | BLIP3o-Pretrain-Long-Caption, BLIP3o-Pretrain-Short-Caption, BLIP3o-Pretrain-Long-Caption-part2 |
BLIP3o datasets | datas/train_data/BLIP3o |
Released key lists and supported converter |
| Generation/editing | ShareGPT-4o text-to-image, ShareGPT_4o_edit |
ShareGPT-4o-Image | datas/train_data/sharegpt_4o |
Editing converter only |
| Editing | Nano-consistent-150k |
Nano-consistent-150k | datas/train_data/nano_consistent_150k |
Official source only |
| Editing | multi_edit |
MultiEdit | datas/train_data/multiedit |
Official source only; gated |
| Editing | GPT_Image_Edit_OmniEdit, GPT_Image_Edit_HQEdit, GPT_Image_Edit_UltraEdit |
GPT-Image-Edit-1.5M, training JSON | datas/train_data/gpt_image_edit/gpt-edit |
Supported converter |
Prepare LLaVA annotations and images. Reuse COCO 2017 from the segmentation preparation with a relative symbolic link:
LLAVA_ANNOTATION_DIR=/absolute/path/LLaVA-Instruct-150K
LLAVA_IMAGES_DIR=datas/train_data/llava_images
mkdir -p \
"$LLAVA_ANNOTATION_DIR" \
"$LLAVA_IMAGES_DIR/coco" \
"$LLAVA_IMAGES_DIR/gqa" \
"$LLAVA_IMAGES_DIR/ocr_vqa/images" \
"$LLAVA_IMAGES_DIR/textvqa" \
"$LLAVA_IMAGES_DIR/vg"
wget -c \
https://huggingface.co/datasets/liuhaotian/LLaVA-Instruct-150K/resolve/main/llava_v1_5_mix665k.json \
-O "$LLAVA_ANNOTATION_DIR/llava_v1_5_mix665k.json"
ln -s ../../coco2017/train2017 "$LLAVA_IMAGES_DIR/coco/train2017"
wget -c https://downloads.cs.stanford.edu/nlp/data/gqa/images.zip \
-O "$LLAVA_IMAGES_DIR/gqa/images.zip"
unzip "$LLAVA_IMAGES_DIR/gqa/images.zip" -d "$LLAVA_IMAGES_DIR/gqa"
wget -c https://dl.fbaipublicfiles.com/textvqa/images/train_val_images.zip \
-O "$LLAVA_IMAGES_DIR/textvqa/train_val_images.zip"
unzip "$LLAVA_IMAGES_DIR/textvqa/train_val_images.zip" \
-d "$LLAVA_IMAGES_DIR/textvqa"
wget -c https://cs.stanford.edu/people/rak248/VG_100K_2/images.zip \
-O "$LLAVA_IMAGES_DIR/vg/images.zip"
wget -c https://cs.stanford.edu/people/rak248/VG_100K_2/images2.zip \
-O "$LLAVA_IMAGES_DIR/vg/images2.zip"
unzip "$LLAVA_IMAGES_DIR/vg/images.zip" -d "$LLAVA_IMAGES_DIR/vg"
unzip "$LLAVA_IMAGES_DIR/vg/images2.zip" -d "$LLAVA_IMAGES_DIR/vg"
wget -c \
'https://drive.usercontent.google.com/download?id=1r0tyZUwGCc4wIG4RkiglCGNL_nFJjR6Q&export=download&confirm=t' \
-O "$LLAVA_IMAGES_DIR/ocr_vqa/dataset.json"
python tools/data_prepare/general_understanding/download_ocr_vqa.pyDownload the other public releases with the project-pinned Hugging Face CLI:
python -m huggingface_hub.commands.huggingface_cli download \
HuggingFaceM4/FineVision --repo-type dataset \
--local-dir datas/train_data/finevision_source
python -m huggingface_hub.commands.huggingface_cli download \
MAmmoTH-VL/MAmmoTH-VL-Instruct-12M --repo-type dataset \
--local-dir datas/train_data/mammoth_vl_source
BLIP3O_ROOT=datas/train_data/BLIP3o
mkdir -p "$BLIP3O_ROOT"
python -m huggingface_hub.commands.huggingface_cli download \
BLIP3o/BLIP3o-Pretrain-Long-Caption --repo-type dataset \
--local-dir "$BLIP3O_ROOT/BLIP3o-Pretrain-Long-Caption"
python -m huggingface_hub.commands.huggingface_cli download \
BLIP3o/BLIP3o-Pretrain-Short-Caption --repo-type dataset \
--local-dir "$BLIP3O_ROOT/BLIP3o-Pretrain-Short-Caption"
# The converter can read tar files directly. Training expects regular image
# files, so extract every archive into a directory named after its tar stem.
for archive in \
"$BLIP3O_ROOT"/BLIP3o-Pretrain-Long-Caption/*.tar \
"$BLIP3O_ROOT"/BLIP3o-Pretrain-Short-Caption/*.tar; do
shard_dir="${archive%.tar}"
mkdir -p "$shard_dir"
tar -xf "$archive" -C "$shard_dir"
done
unset BLIP3O_ROOT
python -m huggingface_hub.commands.huggingface_cli download \
Yejy53/Nano-consistent-150k --repo-type dataset \
--local-dir datas/train_data/nano_consistent_150k
# MultiEdit requires Hugging Face login and acceptance of its access terms.
python -m huggingface_hub.commands.huggingface_cli download \
inclusionAI/MultiEdit --repo-type dataset \
--local-dir datas/train_data/multiedit
python -m huggingface_hub.commands.huggingface_cli download \
FreedomIntelligence/ShareGPT-4o-Image --repo-type dataset \
--include '*.json' '*.tar' \
--local-dir datas/train_data/sharegpt_4o
for archive in datas/train_data/sharegpt_4o/*.tar; do
tar -xf "$archive" -C datas/train_data/sharegpt_4o
doneGPT-Image-Edit provides a parallel downloader for its 4.53 TB image release.
GNU parallel is required:
GPT_EDIT_DIR=datas/train_data/gpt_image_edit
GPT_EDIT_IMAGES="$GPT_EDIT_DIR/gpt-edit"
GPT_EDIT_ANNOTATIONS="$GPT_EDIT_DIR/annotations"
python -m huggingface_hub.commands.huggingface_cli download \
UCSC-VLAA/GPT-Image-Edit-1.5M --repo-type dataset \
--include download.sh --local-dir "$GPT_EDIT_DIR/release"
for family in hqedit omniedit ultraedit; do
bash "$GPT_EDIT_DIR/release/download.sh" \
-d "$family" -o "$GPT_EDIT_IMAGES" -p 8
cat "$GPT_EDIT_IMAGES/$family/$family.tar.gz.part"* | \
tar -xz -C "$GPT_EDIT_IMAGES/$family"
done
python -m huggingface_hub.commands.huggingface_cli download \
UCSC-VLAA/gpt-image-edit-training \
--include 'training_json/*.json' \
--local-dir "$GPT_EDIT_ANNOTATIONS"Create the task JSONL directories and run the deterministic converters:
mkdir -p \
jsonl_generate/train_jsonls/understanding \
jsonl_generate/train_jsonls/editing \
jsonl_generate/train_jsonls/image_generation
python tools/data_prepare/general_understanding/prepare_llava_v1_5.py \
--input-json "$LLAVA_ANNOTATION_DIR/llava_v1_5_mix665k.json"
python tools/data_prepare/general_editing/prepare_sharegpt_4o.py
python tools/data_prepare/general_editing/prepare_gpt_image_edit.pyReconstruct the three BLIP3o image-generation JSONLs from the released key lists. Long Caption and Long Caption part2 both read the official Long Caption source.
python tools/data_prepare/image_generation/convert_blip3o_keys.py \
--blip3o-root datas/train_data/BLIP3o \
--key-list jsonl_generate/train_jsonls/image_generation/keys/BLIP3o-Pretrain-Short-Caption.keys.txt.gz \
--jsonl-out-path jsonl_generate/train_jsonls/image_generation/BLIP3o-Pretrain-Short-Caption.jsonl
python tools/data_prepare/image_generation/convert_blip3o_keys.py \
--blip3o-root datas/train_data/BLIP3o \
--key-list jsonl_generate/train_jsonls/image_generation/keys/BLIP3o-Pretrain-Long-Caption.keys.txt.gz \
--jsonl-out-path jsonl_generate/train_jsonls/image_generation/BLIP3o-Pretrain-Long-Caption.jsonl
python tools/data_prepare/image_generation/convert_blip3o_keys.py \
--blip3o-root datas/train_data/BLIP3o \
--key-list jsonl_generate/train_jsonls/image_generation/keys/BLIP3o-Pretrain-Long-Caption-part2.keys.txt.gz \
--jsonl-out-path jsonl_generate/train_jsonls/image_generation/BLIP3o-Pretrain-Long-Caption-part2.jsonlThe BLIP3o converter rejects duplicate keys, validates each image-caption pair, and preserves key-list order. The understanding and editing converters check their expected counts and decode every image in the first case for each source component. There is no separate validation subcommand.