Please check whether this paper is about 'Voice Conversion' or not.
article info.
- title: Synthetic Audio Generation Framework for Air Traffic Control Speech Recognition
- summary: Automatic Speech Recognition (ASR) systems, despite achieving remarkable accuracy in general-purpose domains with native speech (L1), struggle in domains like Air Traffic Control (ATC) due to strong channel noise, a presence of non-native (L2) English accents, and data scarcity. We propose a synthetic data generation pipeline with acoustical properties simulations specifically designed to address this lack of real data to improve recognition accuracy in the ATC domain. Our approach leverages a combination of neural generation techniques, including Text-to-Speech, Voice Conversion, L2-to-L1 accent conversion, and a novel controllable L1-to-L2 accent conversion framework built to simulate accented speech. Our experiments with the Whisper model on the ATCO2 corpus demonstrate that fine-tuning with either synthetic data alone, or a mix of real and synthetic data, significantly improves the word error rate over out-of-the-box and real data only baselines respectively.
- id: http://arxiv.org/abs/2606.21340v1
judge
Write [vclab::confirmed] or [vclab::excluded] in comment.
Please check whether this paper is about 'Voice Conversion' or not.
article info.
judge
Write [vclab::confirmed] or [vclab::excluded] in comment.