Hi, thanks for making the models publicly available. My question is about the speech decoder.
It seems that the released decoder was fine-tuned with just a single (female) voice:
Would it be possible to provide the checkpoint before this fine-tuning? Otherwise the speaker similarity for audio reconstruction is extremely poor. Thanks!
Hi, thanks for making the models publicly available. My question is about the speech decoder.
It seems that the released decoder was fine-tuned with just a single (female) voice:
Would it be possible to provide the checkpoint before this fine-tuning? Otherwise the speaker similarity for audio reconstruction is extremely poor. Thanks!