Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

2 Commits
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Generation

Our code base includes folowing stages:

  1. Finetune Vision Encoder
  • Finetune CTViT model on Vietnamese PET/CT-report dataset. Details can be found in pet-clip/README.md
  • Finetune Cosmos model on Vietnamese PET/CT-report dataset. Details can be found in Cosmos/README.md
  1. Training and Inference VLMs
  • Training and Inference Vision-Language model. Details can be found in VLMs/README.md
  1. Clinical Evaluation
  • Extract structured lesion information from the LLM output and clinically evaluate the predictions. Details can be found in clinical_evaluation/README.md

Acknowledgments

This research was supported by the NVIDIA Academic Grant Program, which provided the NVIDIA A100 GPU resources used for all model training and evaluation in this work. We also thank NVIDIA for publicly releasing the Cosmos Tokenizer, which served as one of the 3D vision encoders in our benchmark.

About

[NeurIPS 2025] Toward a Vision-Language Foundation Model for Medical Data: Multimodal Dataset and Benchmarks for Vietnamese PET/CT Report Generation

Resources

Stars

15 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages