A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
-
Updated
Aug 24, 2026 - Python
A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
说话人分割仓库-聚类分割-谱聚类 || a ready-to-use repo for Speaker Diariazation with Spectral Clustering
Fork of NVIDIA-NeMo/Speech. A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
Explore Arabic speech AI end-to-end: datasets, models, benchmarks, and production tools for STT, TTS, and dialect-focused tasks.
View GitHub-flavored Markdown files with syntax highlighting, diagrams, and math rendering directly in your browser.
A simple and private transcription tool able to segment speakers and convert audio to text.
Build low-latency voice agents using a modular pipeline with an OpenAI-compatible WebSocket API.
Add a description, image, and links to the speaker-diariazation topic page so that developers can more easily learn about it.
To associate your repository with the speaker-diariazation topic, visit your repo's landing page and select "manage topics."