A fast speech-to-speech & speech-to-text translation model that supports simultaneous decoding and offers 28× speedup.
-
Updated
Oct 22, 2024 - Python
A fast speech-to-speech & speech-to-text translation model that supports simultaneous decoding and offers 28× speedup.
Real time audio to audio translation over sockets. With virtual microphones, you can use this in any video conferencing software you'd like!
Code for NeurIPS 2023 paper "DASpeech: Directed Acyclic Transformer for Fast and High-quality Speech-to-Speech Translation".
List of direct speech-to-speech translation papers.
End-to-end speech-to-speech translation pipeline with voice cloning (RVC) and automatic lip-sync (Wav2Lip).
Code for ACL 2024 main conference paper "Can We Achieve High-quality Direct Speech-to-Speech Translation Without Parallel Speech Data?".
Applying deep learning to translate animation and re-generate audio.
cascaded speech-to-speech translation (STST), mapping from source speech in any language to target speech in English
Official repository for our NeurIPS 2024 paper: DiffNorm: Self-Supervised Normalization for Non-autoregressive Speech-to-speech Translation
Tool to generate English AI Dubbing for a YouTube video
HF Space app for End-to-End Speech-to-Speech Translation from Spanish to English using ESPnet
Implementation of Speech Vecalign: an Embedding-based Method for Aligning Parallel Speech Documents.
A comparison of E2E and Cascading S2ST systems on the CVSS-C Spanish to English dataset (CommonVoice 4.0)
Speech to Speech Translation Python
MedSpeak - bidirectional Yoruba-English speech-to-speech translation for medical consultations. Fine-tuned Whisper ASR + MarianMT NMT + VITS TTS, Flask API, React frontend.
Hệ thống tự động tải video đa nền tảng, trích xuất phụ đề (Faster-Whisper), dịch thuật AI (Gemini/OpenAI), lồng tiếng Việt (CapCut/vieNeuTTS), xóa sub gốc (PaddleOCR) và render video qua Telegram Bot.
Supplementary data, Parsifal export, and references for the systematic literature review on end-to-end Audio-Visual Speech-to-Speech Translation (AV-S2ST) models.
End-to-End Speech-to-Speech Translation System for Balti (بلتی) using Fine-Tuned Whisper (17.4% WER), NLLB, and Kokoro TTS.
Speech-to-speech translation with Meta's SeamlessM4T v2 — CLI, Vue 3 web dashboard and a browser extension that dubs videos as they play. Expressive voice preservation, GPU acceleration, SQLite job queue.
To associate your repository with the speech-to-speech-translation topic, visit your repo's landing page and select "manage topics."