Skip to content
@QwenAudio

QwenAudio

Open-source speech and audio language models from the QwenAudio Team

Popular repositories Loading

  1. CosyVoice CosyVoice Public

    Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.

    Python 22.5k 2.6k

  2. SenseVoice SenseVoice Public

    Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.

    C 9k 799

  3. Fun-ASR Fun-ASR Public

    Open-source LLM-based ASR model family for Chinese, dialect, accent, and multilingual speech, with FunASR, vLLM, streaming, and llama.cpp runtimes.

    C 1.4k 141

  4. FunMusic FunMusic Public

    A fundamental toolkit designed for music, song, and audio generation

    Python 1.4k 138

  5. ThinkSound ThinkSound Public

    [NeurIPS 2025] PyTorch implementation of [ThinkSound], a unified framework for generating audio from any modality, guided by Chain-of-Thought (CoT) reasoning.

    Python 1.4k 81

  6. Fun-Audio-Chat Fun-Audio-Chat Public

    Fun-Audio-Chat is a Large Audio Language Model built for natural, low-latency voice interactions.

    Python 985 105

Repositories

Showing 10 of 17 repositories
  • qwen-audio-agent Public

    A realtime voice runtime that keeps Agents talking, working, and present. Real-time Voice Runtime for AI Agents

    QwenAudio/qwen-audio-agent's past year of commit activity
    JavaScript 174 Apache-2.0 11 0 9 Updated Jul 28, 2026
  • SenseVoice Public

    Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.

    QwenAudio/SenseVoice's past year of commit activity
    C 8,954 MIT 799 0 0 Updated Jul 27, 2026
  • QwenAudio/FunAudioLLM.github.io's past year of commit activity
    HTML 59 MIT 11 0 0 Updated Jul 24, 2026
  • Fun-ASR Public

    Open-source LLM-based ASR model family for Chinese, dialect, accent, and multilingual speech, with FunASR, vLLM, streaming, and llama.cpp runtimes.

    QwenAudio/Fun-ASR's past year of commit activity
    C 1,445 Apache-2.0 141 0 0 Updated Jul 24, 2026
  • FunResearch Public

    This repository is maintained by the Speech Team at Alibaba’s Tongyi Lab, serving as an open-source platform for our cutting-edge research in speech, audio, NLP technologies. We believe in accelerating scientific progress through transparent collaboration, and invite the global research community to explore, reproduce, and build upon our work.

    QwenAudio/FunResearch's past year of commit activity
    Python 42 Apache-2.0 6 2 0 Updated Jul 23, 2026
  • llama-index-readers-funasr Public

    FunASR (SenseVoice/Paraformer/Fun-ASR-Nano) audio reader for LlamaIndex

    QwenAudio/llama-index-readers-funasr's past year of commit activity
    Python 2 MIT 0 0 0 Updated Jun 17, 2026
  • langchain-funasr Public

    FunASR (SenseVoice/Paraformer/Fun-ASR-Nano) speech-to-text integration for LangChain

    QwenAudio/langchain-funasr's past year of commit activity
    Python 0 0 0 0 Updated Jun 17, 2026
  • funasr-haystack Public archive

    FunASR (SenseVoice/Paraformer) speech-to-text integration for Haystack

    QwenAudio/funasr-haystack's past year of commit activity
    2 0 0 0 Updated Jun 17, 2026
  • CosyVoice Public

    Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.

    QwenAudio/CosyVoice's past year of commit activity
    Python 22,479 Apache-2.0 2,593 716 30 Updated May 26, 2026
  • ThinkSound Public

    [NeurIPS 2025] PyTorch implementation of [ThinkSound], a unified framework for generating audio from any modality, guided by Chain-of-Thought (CoT) reasoning.

    QwenAudio/ThinkSound's past year of commit activity
    Python 1,372 81 39 3 Updated Apr 3, 2026

Top languages

Loading…

Most used topics

Loading…