Cross-modal extension: adding speech input to Qwen2-VL with a frozen Whisper encoder and a trained projector. 11.3% WER / 6.7% CER on the 8,086-sample LargeScaleASR test partition.
machine-learning transformers pytorch speech-recognition whisper multimodal qlora qwen2-vl-7b architecture-grafting
-
Updated
Jul 25, 2026 - Jupyter Notebook