The open-source AI voice studio. Clone, dictate, create.
-
Updated
Aug 9, 2026 - TypeScript
The open-source AI voice studio. Clone, dictate, create.
VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.
Production ready toolkit to run AI locally
Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.
Run Claude Code 100% on-device with local AI on Apple Silicon. MLX-native Anthropic-API server. 6 fighters incl. Muse-Glimmer 30B (now multimodal — reads images, abliterated), Gemma 4 31B, Qwen 3.5 122B (65 tok/s), DeepSeek V4 Flash (1M ctx). Private, offline, airgap-ready. Built for NDA / legal / healthcare workflows.
Voice AI SDK is a reusable Android library that gives any app a full voice-driven AI conversation pipeline in minutes. Voice Assistant + Android Voide AI + SDK + MVVM + Kotlin
The Self-Coding System for Your App — Alan AI SDK for Web
A realtime voice runtime that keeps Agents talking, working, and present. Real-time Voice Runtime for AI Agents
The Self-Coding System for Your App — Alan AI SDK for iOS
The Self-Coding System for Your App — Alan AI SDK for Flutter
The Self-Coding System for Your App — Alan AI SDK for Ionic
Formerly Better Chatbot. Navigator is an open-source AI workspace for agents, MCP and workflow automation.
Open-source voice-AI SDK. The Vapi/Retell alternative for builders who want to own the stack. Give your AI agent a phone number in 4 lines — Python and TypeScript, MIT licensed, Twilio, Telnyx, and Plivo.
Optimized Whisper models for streaming and on-device use
Real-time voice assistant — WebRTC streaming, faster-whisper ASR, local LLM, Vui Nano (300M) TTS. OpenAI Realtime API compatible. Voice cloning, barge-in, ~9× realtime on a 4090. Apache 2.0.
Rapida is an open-source, end-to-end voice AI orchestration platform for building real-time conversational voice agents with audio streaming, STT, TTS, VAD, multi-channel integration, agent state management, and observability.
The retrieval layer for production AI systems. Lightning-fast (<10ms) search without vector databases. Built for browser, edge, on-device, and cloud.
TypeScript AI agent framework: cognitive memory, runtime tool forging, multi-agent orchestration, 11 LLM providers.
Add a description, image, and links to the voice-ai topic page so that developers can more easily learn about it.
To associate your repository with the voice-ai topic, visit your repo's landing page and select "manage topics."