这是基于官方 IndexTTS 2.5 制作的 Windows Electron 桌面整合版源码仓库,当前桌面版本为 0.26.2。完整功能、开发和打包说明见 Desktop README。
0.26.2 对齐官方 PR #802:英文、西班牙语等非 CJK 长文本会把过大的单段预算限制到约 86 Token,降低末句读错或编造内容的风险; 0.26.1 新增通用常驻本地 API:兼容 OpenAI
/v1/audio/speech,同时提供 IndexTTS 高级参数与异步任务接口; 桌面界面和 API 共用一份模型、音色库、输出历史和 GPU 推理锁,便携包另带前台常驻的启动/状态/停止 CMD。 同时修复多段长文本自动关闭流式试听后,最终音频已经生成却出现Error 181的问题。 桌面更新拆分为签名应用层与独立运行库层,支持续传、校验、健康检查和自动回滚。 模型、更新和可选组件仍只在用户明确操作后下载或安装。
- 官方开源项目:index-tts/index-tts(感谢官方团队开源)
- T8star-Aix ComfyUI 节点:T8mars/comfyui-indextts25-t8
- B站:T8star-Aix
- YouTube:T8star-Aix
- ComfyUI 整合包:夸克网盘
- IndexTTS 2.5 模型网盘:夸克网盘
- IndexTTS 2.5 Hugging Face 模型:t8star/IndexTTS-2.5-Comfy
仓库不包含约 10 GB 的模型权重、Python 虚拟环境、Electron 构建产物、用户数据或生成音频。
运行便携版时,可在启动器中选择完整的 IndexTTS 2.5 模型目录,也可点击
“Hugging Face 自动下载/修复完整模型”;选择父目录后会创建 IndexTTS-2.5 子目录。源码开发环境与打包步骤见上方
Desktop README。本整合版并非 IndexTTS 官方发行版,模型与基础推理代码的权利和许可归原项目所有。
- “语音生成”和“多角色 / 批量台词 / SRT”页使用固定在窗口底部的生成/停止操作栏,滚动到任何位置都可立即操作;非常用说明、质检和高级参数默认折叠。
- 多角色固定操作栏会持续显示当前行、已完成数量和百分比;时间轴选中任意台词后可直接上移/下移,角色、语言、文本和情感一起移动,已有 SRT 时间槽保持原位。
- 时间轴表格上方直接显示逐句情感速写;折叠速查列出八维顺序与含义。当前可编辑时间轴可单独导出为保真的 JSON 或易于表格编辑的 CSV,之后导入并继续生成,不必打包完整音频工程。
- “任务队列”统一保存单句、多角色和 SRT 任务,桌面程序重启后仍可继续;失败或取消的任务可以重新排队。
- “生成历史”以紧凑列表和详情面板展示结果,可搜索、按类型/语言筛选、试听并下载音频、逐句 ZIP、字幕和报告;路径不会暴露在列表中,且只能在配置的输出目录内定位文件。
- 长文本可在生成前查看实际归一化文本、Token 分段、预计时长和风险;英文、西班牙语等非 CJK 文本的大分段值会自动限制到安全上限,已有较小限制不会被重复缩小;预检结果只有点击“应用到正文”才会替换输入。
- 参考音频会自动给出质量摘要;模型支持手动释放、默认空闲 10 分钟释放和下次生成自动重载。追加候选后可在折叠的 A/B 工作区试听、评分和收藏。
- 启动器可分别更改“生成音频保存目录”和“用户数据目录”,设置会持久保存。音色库、预设、任务、ASR 缓存和日志都跟随用户数据目录,不必放在 C 盘。
- 正式 WebUI 顶部可直接打开输出、用户数据和日志目录,也可点击“返回启动配置(停止模型)”。点击正式界面的窗口关闭按钮时,同样会先停止模型、释放显存并返回启动配置;在启动配置页再次关闭才退出程序。
- 可选加速缺失编译工具或运行库时属于正常安全回退,不等于语音生成失败。页面默认显示实际生效模式和中文说明,完整 JSON 只放在“技术诊断 JSON”折叠区供排错复制。
Desktop 0.26.1 可供 SillyTavern、直播工具、剪辑脚本和其他程序直接调用,不依赖专用适配器。启动器可选择“桌面 + API”或“仅 API”;便携包 EXE 同级的 启动API服务.cmd 会使用内置 Python 在前台持续监听,错误不会一闪而过。默认地址 http://127.0.0.1:7861,OpenAI 兼容接口为 POST /v1/audio/speech,完整 Swagger 文档为 /docs。
API Key 自动生成并可在启动器复制;音色只从桌面角色音色库选择,生成结果保存到相同输出目录并进入生成历史。原生高级接口还支持情感向量、种子、采样/扩散、分段和后处理,长文本可用异步任务查询进度和取消。详细启动方式、请求示例、安全与排错说明见 本地 API 使用说明。
如果音色克隆正常,但 1939年 等阿拉伯数字听起来不对,问题通常来自文本归一化,不是参考音频。
整合包已把 Windows 文本前端 wetext>=0.1.7,<0.2 列为直接依赖,并在完整便携包构建时一并安装;
Linux 开发环境使用 WeTextProcessing>=1.2.0,<2。应用启动时会真实执行一次
1939年 → 一九三九年 自检,并在“语音生成 → 高级生成参数 → 文本归一化(数字/日期)”下方显示状态。
- 年份:
1939年通常读作“一九三九年”。 - 数量:
1939个人通常读作“一千九百三十九个人”。 - 关闭“文本归一化”或自检失败时,请直接把数字写成希望听到的中文口语形式。
完整便携包会包含这项 Python 依赖;新版启动器也支持从 GitHub Release 获取独立签名运行库层,
按小于 2 GiB 的分卷续传并在安装前完成分卷、整包和逐文件 SHA-256 校验。旧版启动器若尚不支持运行库层,
会先安装小体积程序层,重启后再检查一次即可安装运行库。桌面源码环境可执行
uv sync --frozen --extra webui 安装锁定依赖。
ComfyUI 节点的独立安装、状态检查和命令见
节点 README 的数字与年份说明。
- GitHub Release 发布小体积桌面程序层;Python / Torch / CUDA 运行库使用独立
runtime-v*Release,自动拆成每卷小于 2 GiB。 - 启动器只接受本项目 GitHub Release 地址和 Ed25519 签名的更新清单,并依次验证每个分卷、合并归档及解压后的每个文件。
- 安装前会预估下载、解压和回滚备份所需空间;安装后必须通过启动健康检查,否则自动恢复原文件。
- 约 10 GB 的模型始终与程序、运行库分离,继续从 t8star/IndexTTS-2.5-Comfy 按固定版本下载、续传和校验。
维护者构建与发布步骤见 分层更新发布说明。
- 在“语音生成”或“多角色 / 批量台词 / SRT”页点击“加入队列”,会保存当时的全部生成参数和脚本;到“任务队列”页统一顺序执行。
- 队列文件位于用户数据目录的
task_queue.json。意外退出时,运行中的项目会在下次启动恢复为等待状态;不会自动开始推理。 - 将高级参数里的“追加候选数量”设为
1–3,生成后展开“声音 A/B 候选试听、评分与收藏”。收藏文件位于用户数据目录的candidate_favorites/,评分记录位于candidate_reviews.json。 - 以上区域均按需展开;首页只保留音色、文本、语言、时长和生成/停止等常用操作。
.\.venv\Scripts\python.exe -m indextts.cli "欢迎使用 IndexTTS 2.5。" `
--voice .\voice.wav --model-dir .\checkpoints --language ZH `
--precision auto --duration-factor 1.0 --output-path .\gen.wavCLI 支持 ZH / EN / JA / ES / AR、参考情感、八维情感向量、情感文本、正式时长系数、
原生目标时长、完整采样/CFM 参数,以及可选 BigVGAN CUDA、GPT 加速、Torch Compile 和 DeepSpeed。
运行 python -m indextts.cli --help 可查看全部参数。
OpenAI Whisper 固定为 20250625;中英日西使用 base,Arabic 使用实测更准确的 small。
8GB/24GB 脱敏基线保存在节点仓库的 quality_baselines/openai-whisper-mixed-*-gpu.json;
每周任务串行执行两档并生成 CER/WER、RTF、峰值显存趋势图,不让普通 PR/Push 加载大模型。
Torchaudio 2.9+ 使用原生 TorchCodec I/O,并在启动前检查 TorchCodec 与 FFmpeg 共享 DLL;桌面自带的
Torchaudio 2.8 环境继续走兼容后端。
.\.venv\Scripts\python.exe .\desktop\scripts\smoke-multilingual-quality.py `
--model-dir .\checkpoints --voice .\voice.wav `
--asr-backend auto --output-dir .\quality-regression --strict工具固定生成中、英、日、西、阿五组长文本,输出每组 WAV 和 quality-report.json,记录 CER/WER、
分段语速离散度、削波、静音、时长、RTF 与峰值显存。再次运行时可用
--baseline 旧报告.json 检查音质或性能回退;脚本不会自动下载主模型或参考音频,
但启用可选 ASR 时,所选 Whisper 模型可能下载到输出目录的 asr_models 子目录。
以下为官方上游 README:
| Model | Demos | Paper | Modelscope | HuggingFace |
|---|---|---|---|---|
| IndexTTS-2.5 | Demos | Paper | Modelscope | HuggingFace |
| IndexTTS-2 | Demos | Paper | Modelscope | HuggingFace |
| IndexTTS-1.5 | Demos | Paper | Modelscope | HuggingFace |
| IndexTTS | Demos | Paper | Modelscope | HuggingFace |
2026/07/17🔥 We release IndexTTS-2.5- The model now supports Chinese, English, Japanese, Spanish and Arabic, with faster inference speed compared to IndexTTS-2, while maitaining the cross-lingual and timbre-emotion disentanglement capabilities.
- The model improves the controbility of Chinese Pinyin and English CMU phonemes and Japanese Kana.
2025/09/08🔥 We release IndexTTS-2- The first autoregressive TTS model with precise synthesis duration control, supporting both controllable and uncontrollable modes. This functionality is not yet enabled in this release.
- The model achieves highly expressive emotional speech synthesis, with emotion-controllable capabilities enabled through multiple input modalities.
2025/05/14🔥 We release IndexTTS-1.5, significantly improving the model's stability and its performance in the English language.2025/03/25🔥 We release IndexTTS-1.0 with model weights and inference code.2025/02/12🎉 We submitted our paper to arXiv, and released our demos and test sets.
Table 1: Zero-shot TTS evaluation results on CV3-Eval test set (Arabic uses an in-house test set). WER (%) ↓ and Speaker Similarity (SS) ↑ are reported. †Results cited from the original paper.
| Model | Params | test-zh WER (%) ↓ |
test-zh SS (%) ↑ |
test-en WER (%) ↓ |
test-en SS (%) ↑ |
test-es WER (%) ↓ |
test-es SS (%) ↑ |
test-ja WER (%) ↓ |
test-ja SS (%) ↑ |
test-ar WER (%) ↓ |
test-ar SS (%) ↑ |
Average WER (%) ↓ |
Average SS (%) ↑ |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| VoxCPM2 | 2B | 3.88 | 74.99 | 5.13 | 71.57 | 5.49 | 74.67 | 6.69 | 72.90 | 14.94 | 65.99 | 7.22 | 72.02 |
| OmniVoice | 0.8B | 3.41 | 72.99 | 3.62 | 70.13 | 3.52 | 74.14 | 5.38 | 70.49 | 17.88 | 64.22 | 6.76 | 70.39 |
| Moss-TTS 1.5 | 8B | 4.02 | 72.68 | 4.45 | 67.46 | 3.83 | 71.75 | 10.97 | 68.71 | 23.71 | 62.21 | 9.40 | 68.56 |
| CosyVoice3-0.5B | 0.5B | 3.84 | 80.01 | 4.88 | 74.16 | 4.04 | 78.85 | - | 76.36 | - | - | - | - |
| CosyVoice3-1.5B | 1.5B | 3.91† | - | 4.99† | - | 4.47† | - | 7.57† | - | - | - | - | - |
| FireRedTTS-2 | 1.5B | 8.22 | 68.10 | 14.92 | 56.93 | - | - | - | - | - | - | - | - |
| Fish Audio S2 Pro | 4B | 3.62 | 67.79 | 3.83 | 61.66 | 2.93 | 67.44 | 5.15 | 66.15 | 14.15 | 59.43 | 5.94 | 64.49 |
| Qwen3-TTS | 1.7B | 3.27 | 73.02 | 5.06 | 67.17 | 2.87 | 73.17 | 5.89 | 70.18 | - | - | - | - |
| IndexTTS2.5 | 0.8B | 4.36 | 77.10 | 5.12 | 68.06 | 3.75 | 76.39 | 5.66 | 74.62 | 14.88 | 69.74 | 6.75 | 73.18 |
| IndexTTS2.5-RL | 0.8B | 3.93 | 77.92 | 3.89 | 67.79 | 3.33 | 76.68 | 5.30 | 75.41 | 13.58 | 70.36 | 6.00 | 73.63 |
Table 2: Cross-lingual TTS evaluation on CV3-Eval test set (Chinese prompt → target language, Arabic uses an in-house test set). WER (%) ↓ and Speaker Similarity (SS) ↑ are reported.
| Model | Params | zh→en WER (%) ↓ |
zh→en SS (%) ↑ |
zh→es WER (%) ↓ |
zh→es SS (%) ↑ |
zh→ja WER (%) ↓ |
zh→ja SS (%) ↑ |
zh→ar WER (%) ↓ |
zh→ar SS (%) ↑ |
Average WER (%) ↓ |
Average SS (%) ↑ |
|---|---|---|---|---|---|---|---|---|---|---|---|
| VoxCPM2 | 2B | 4.48 | 64.25 | 16.38 | 64.89 | 11.84 | 71.54 | 11.09 | 67.62 | 10.95 | 67.08 |
| OmniVoice | 0.8B | 3.74 | 64.91 | 5.84 | 62.08 | 9.09 | 69.06 | 19.80 | 65.27 | 9.62 | 65.33 |
| Moss-TTS 1.5 | 8B | 6.13 | 59.23 | 4.32 | 56.63 | 11.52 | 65.54 | 17.03 | 62.93 | 9.75 | 61.08 |
| CosyVoice3-0.5B | 0.5B | 3.23 | 62.79 | 4.58 | 64.04 | - | - | - | - | - | - |
| CosyVoice3-1.5B | 1.5B | 4.32 | - | - | - | 13.70 | - | - | - | - | - |
| FireRedTTS-2 | 1.5B | 9.34 | 53.19 | 12.25 | 58.31 | 19.05 | 64.12 | - | - | - | - |
| Fish Audio S2 Pro | 4B | 4.14 | 55.89 | 4.46 | 55.57 | 10.48 | 61.74 | 14.49 | 59.80 | 8.39 | 58.25 |
| Qwen3-TTS | 1.7B | 5.74 | 63.04 | 5.15 | 68.02 | 36.09 | 65.71 | - | - | - | - |
| IndexTTS2.5 | 0.8B | 3.62 | 63.83 | 5.17 | 65.48 | 6.57 | 74.16 | 9.51 | 71.02 | 6.22 | 68.62 |
| IndexTTS2.5-RL | 0.8B | 3.55 | 67.47 | 4.86 | 64.47 | 6.38 | 75.82 | 9.89 | 73.05 | 6.17 | 70.20 |
QQ Group:663272642(No.4) 1013410623(No.5)
Discord:https://discord.gg/uT32E7KDmy
Email:indexspeech@bilibili.com
You are welcome to join our community! 🌏
欢迎大家来交流讨论!
Caution
Thank you for your support of the bilibili indextts project! Please note that the only official channel maintained by the core team is: https://github.com/index-tts/index-tts. Any other websites or services are not official, and we cannot guarantee their security, accuracy, or timeliness. For the latest updates, please always refer to this official repository.
Tips: Please contact the authors for more detailed information. For commercial usage and cooperation, please contact indexspeech@bilibili.com.
The Git-LFS plugin must also be enabled on your current user account:
git lfs install- Download this repository:
git clone https://github.com/index-tts/index-tts.git && cd index-tts
git lfs pull # download large repository files- Install the uv package manager. It is required for a reliable, modern installation environment.
Tip
Quick & Easy Installation Method:
There are many convenient ways to install the uv command on your computer.
Please check the link above to see all options. Alternatively, if you want
a very quick and easy method, you can install it as follows:
pip install -U uvWarning
We only support the uv installation method. Other tools, such as conda
or pip, don't provide any guarantees that they will install the correct
dependency versions. You will almost certainly have random bugs, error messages,
missing GPU acceleration, and various other problems if you don't use uv.
Please do not report any issues if you use non-standard installations, since
almost all such issues are invalid.
Furthermore, uv is up to 115x faster
than pip, which is another great reason to embrace the new industry-standard
for Python project management.
- Install required dependencies:
We use uv to manage the project's dependency environment. The following command
will automatically create a .venv project-directory and then installs the correct
versions of Python and all required dependencies:
uv sync --all-extrasIf the download is slow, please try a local mirror, for example any of these local mirrors in China (choose one mirror from the list below):
uv sync --all-extras --default-index "https://mirrors.aliyun.com/pypi/simple"
uv sync --all-extras --default-index "https://mirrors.tuna.tsinghua.edu.cn/pypi/web/simple"Tip
Available Extra Features:
--all-extras: Automatically adds every extra feature listed below. You can remove this flag if you want to customize your installation choices.--extra webui: Adds WebUI support (recommended).--extra deepspeed: Adds DeepSpeed support (may speed up inference on some systems).
Important
Important (Windows): The DeepSpeed library may be difficult to install for
some Windows users. You can skip it by removing the --all-extras flag. If you
want any of the other extra features above, you can manually add their specific
feature flags instead.
Important (Linux/Windows): If you see an error about CUDA during the installation, please ensure that you have installed NVIDIA's CUDA Toolkit version 12.8 (or newer) on your system.
- Download the required models via uv tool:
Download via huggingface-cli:
uv tool install "huggingface-hub[cli,hf_xet]"
hf download IndexTeam/IndexTTS-2 --local-dir=checkpointsOr download via modelscope:
uv tool install "modelscope"
modelscope download --model IndexTeam/IndexTTS-2 --local_dir checkpointsImportant
If the commands above aren't available, please carefully read the uv tool
output. It will tell you how to add the tools to your system's path.
Note
In addition to the above models, some small models will also be automatically downloaded when the project is run for the first time. If your network environment has slow access to HuggingFace, it is recommended to execute the following command before running the code:
export HF_ENDPOINT="https://hf-mirror.com"If you need to diagnose your environment to see which GPUs are detected, you can use our included utility to check your system:
uv run tools/gpu_check.py# IndexTTS2 (default)
uv run webui.py
# IndexTTS2.5
uv run webui.py --version 2.5 --model_dir ./checkpointsOpen your browser and visit http://127.0.0.1:7860 to see the demo.
You can also adjust the settings to enable features such as BF16(IndexTTS 2.5)/FP16(IndexTTS 2) inference (lower VRAM usage), DeepSpeed acceleration, compiled CUDA kernels for speed, etc. All available options can be seen via the following command:
uv run webui.py -hHave fun!
Important
It can be very helpful to use FP16/BF16 (half-precision) inference. It is faster and uses less VRAM, with a very small quality loss.
DeepSpeed may also speed up inference on some systems, but it could also make it slower. The performance impact is highly dependent on your specific hardware, drivers and operating system. Please try with and without it, to discover what works best on your personal system.
Lastly, be aware that all uv commands will automatically activate the correct
per-project virtual environments. Do not manually activate any environments
before running uv commands, since that could lead to dependency conflicts!
To run scripts, you must use the uv run <file.py> command to ensure that
the code runs inside your current "uv" environment. It may sometimes also be
necessary to add the current directory to your PYTHONPATH, to help it find
the IndexTTS modules.
Example of running a script via uv:
# IndexTTS2
PYTHONPATH="$PYTHONPATH:." uv run indextts/infer_v2.py
# IndexTTS2.5
PYTHONPATH="$PYTHONPATH:." uv run indextts/infer_v2_5.py \
--cfg_path checkpoints/config_v2_5.yaml \
--model_dir checkpoints \
--text "Hello world" \
--lang ENHere are several examples of how to use in your own scripts:
- Initialize IndexTTS
# IndexTTS2
from indextts.infer_v2 import IndexTTS2
tts = IndexTTS2(cfg_path="checkpoints/config.yaml", model_dir="checkpoints", use_fp16=False, use_cuda_kernel=False, use_deepspeed=False)
# IndexTTS2.5
from indextts.infer_v2_5 import IndexTTS2
tts = IndexTTS2(cfg_path="checkpoints_25/config_v2_5.yaml", model_dir="checkpoints_25", use_bf16=True)- Synthesize new speech with a single reference audio file (voice cloning):
text = "Translate for me, what is a surprise!"
# IndexTTS2
tts.infer(spk_audio_prompt='examples/voice_01.wav', text=text, output_path="gen.wav", verbose=True)
# IndexTTS2.5 (multilingual, with language selection)
tts.infer(spk_audio_prompt='examples/voice_01.wav', text=text, lang="EN", output_path="gen.wav", verbose=True)- Using a separate, emotional reference audio file to condition the speech synthesis:
text = "酒楼丧尽天良,开始借机竞拍房间,哎,一群蠢货。"
# IndexTTS2
tts.infer(spk_audio_prompt='examples/voice_07.wav', text=text, output_path="gen.wav", emo_audio_prompt="examples/emo_sad.wav", verbose=True)
# IndexTTS2.5
tts.infer(spk_audio_prompt='examples/voice_07.wav', text=text, lang="ZH", output_path="gen.wav", emo_audio_prompt="examples/emo_sad.wav", verbose=True)- When an emotional reference audio file is specified, you can optionally set
the
emo_alphato adjust how much it affects the output. Valid range is0.0 - 1.0, and the default value is1.0(100%):
text = "酒楼丧尽天良,开始借机竞拍房间,哎,一群蠢货。"
# IndexTTS2
tts.infer(spk_audio_prompt='examples/voice_07.wav', text=text, output_path="gen.wav", emo_audio_prompt="examples/emo_sad.wav", emo_alpha=0.9, verbose=True)
# IndexTTS2.5
tts.infer(spk_audio_prompt='examples/voice_07.wav', text=text, output_path="gen.wav", lang="ZH", emo_audio_prompt="examples/emo_sad.wav", emo_alpha=0.9, verbose=True)- It's also possible to omit the emotional reference audio and instead provide
an 8-float list specifying the intensity of each emotion, in the following order:
[happy, angry, sad, afraid, disgusted, melancholic, surprised, calm]. You can additionally use theuse_randomparameter to introduce stochasticity during inference; the default isFalse, and setting it toTrueenables randomness:
Note
Enabling random sampling will reduce the voice cloning fidelity of the speech synthesis.
text = "对不起嘛!我的记性真的不太好,但是和你在一起的事情,我都会努力记住的~"
# IndexTTS2
tts.infer(spk_audio_prompt='examples/09.wav', text=text, output_path="gen.wav", emo_vector=[0, 0, 0.8, 0, 0, 0, 0, 0], use_random=False, verbose=True)
# IndexTTS2.5
tts.infer(spk_audio_prompt='examples/09.wav', text=text, lang="ZH", output_path="gen.wav", emo_vector=[0, 0, 0.8, 0, 0, 0, 0, 0], use_random=False, verbose=True)- Alternatively, you can enable
use_emo_textto guide the emotions based on your providedtextscript. Your text script will then automatically be converted into emotion vectors. It's recommended to useemo_alphaaround 0.6 (or lower) when using the text emotion modes, for more natural sounding speech. You can introduce randomness withuse_random(default:False;Trueenables randomness):
text = "快躲起来!是他要来了!他要来抓我们了!"
# IndexTTS2
tts.infer(spk_audio_prompt='examples/voice_12.wav', text=text, output_path="gen.wav", emo_alpha=0.6, use_emo_text=True, use_random=False, verbose=True)
# IndexTTS2.5
tts.infer(spk_audio_prompt='examples/voice_12.wav', text=text, lang="ZH", utput_path="gen.wav", emo_alpha=0.6, use_emo_text=True, use_random=False, verbose=True)- It's also possible to directly provide a specific text emotion description
via the
emo_textparameter. Your emotion text will then automatically be converted into emotion vectors. This gives you separate control of the text script and the text emotion description:
text = "快躲起来!是他要来了!他要来抓我们了!"
emo_text = "你吓死我了!你是鬼吗?"
# IndexTTS2
tts.infer(spk_audio_prompt='examples/voice_12.wav', text=text, output_path="gen.wav", emo_alpha=0.6, use_emo_text=True, emo_text=emo_text, use_random=False, verbose=True)
# IndexTTS2.5
tts.infer(spk_audio_prompt='examples/voice_12.wav', text=text, lang="ZH", output_path="gen.wav", emo_alpha=0.6, use_emo_text=True, emo_text=emo_text, use_random=False, verbose=True)Tip
IndexTTS2.5 Pinyin/English phonemes/Japan Kana Usage Notes:
IndexTTS2.5 now can support these character replacement, with better instruction-following capability.
For the full list of valid entries, please refer to checkpoints/pinyin.vocab for Pinyin, and 'https://svn.code.sf.net/p/cmusphinx/code/trunk/cmudict/cmudict-0.7b' for CMU dictionary.
Example:
他在银<行|XING2>里<行|HANG2>走了半天,发现这笔业务办不<行|HANG2>。
He had a <minute|M IH1 . N AH0 T> to examine the <minute|M AY0 . N UW1 T> details of the contract.
彼は料理が<上手|じょうず>だが、囲碁では<上手|うわて>に負けた。
IndexTTS2 Pinyin Usage Notes:
IndexTTS2 still supports mixed modeling of Chinese characters and Pinyin.
When you need precise pronunciation control, please provide text with specific Pinyin annotations to activate the Pinyin control feature.
Note that Pinyin control does not work for every possible consonant–vowel combination; only valid Chinese Pinyin cases are supported.
For the full list of valid entries, please refer to checkpoints/pinyin.vocab.
Example:
之前你做DE5很好,所以这一次也DEI3做DE2很好才XING2,如果这次目标完成得不错的话,我们就直接打DI1去银行取钱。
IndexTTS1 Usage Notes:
You can also use our previous IndexTTS1 model by importing a different module:
from indextts.infer import IndexTTS
tts = IndexTTS(model_dir="checkpoints",cfg_path="checkpoints/config.yaml")
voice = "examples/voice_07.wav"
text = "大家好,我现在正在bilibili 体验 ai 科技,说实话,来之前我绝对想不到!AI技术已经发展到这样匪夷所思的地步了!比>如说,现在正在说话的其实是B站为我现场复刻的数字分身,简直就是平行宇宙的另一个我了。如果大家也想体验更多深入的AIGC功能,可>以访问 bilibili studio,相信我,你们也会吃惊的。"
tts.infer(voice, text, 'gen.wav')For more detailed information, see README_INDEXTTS_1_5, or visit the IndexTTS1 repository at index-tts:v1.5.0.
🌟 If you find our work helpful, please leave us a star and cite our paper.
IndexTTS2.5:
@misc{li2026indextts25technicalreport,
title={IndexTTS 2.5 Technical Report},
author={Yunpei Li and Xun Zhou and Jinchao Wang and Lu Wang and Yong Wu and Siyi Zhou and Yiquan Zhou and Jingchen Shu},
year={2026},
eprint={2601.03888},
archivePrefix={arXiv},
primaryClass={cs.SD},
url={https://arxiv.org/abs/2601.03888},
}
IndexTTS2:
@article{zhou2025indextts2,
title={IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech},
author={Siyi Zhou, Yiquan Zhou, Yi He, Xun Zhou, Jinchao Wang, Wei Deng, Jingchen Shu},
journal={arXiv preprint arXiv:2506.21619},
year={2025}
}
IndexTTS:
@article{deng2025indextts,
title={IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System},
author={Wei Deng, Siyi Zhou, Jingchen Shu, Jinchao Wang, Lu Wang},
journal={arXiv preprint arXiv:2502.05512},
year={2025},
doi={10.48550/arXiv.2502.05512},
url={https://arxiv.org/abs/2502.05512}
}
