INDEXTTS 2.5 TECHNICAL REPORT
A source-based overview of the official IndexTTS 2.5 technical report, its five demonstrated languages, architecture changes and audio examples.
Version notice: the free interactive demo on this site currently embeds IndexTTS2. This page explains the IndexTTS 2.5 report and does not claim that the embedded demo runs 2.5.
Demonstrated languages
Reported RTF improvement
Semantic codec frame rate
The following points are summarized from the official project page and technical report.
The semantic codec frame rate is reduced from 50 Hz to 25 Hz, shortening sequences and lowering training and inference cost.
The semantic-to-mel backbone is changed to a more efficient Zipformer-based architecture.
Boundary-aware alignment, token-level concatenation and instruction-guided generation support cross-lingual synthesis.
GRPO post-training is used to improve pronunciation accuracy and naturalness.
The official sample page demonstrates zero-shot voice cloning and emotional speech across these five languages.
Keeping the versions separate prevents the report's capabilities from being presented as features already available in the embedded demo.
The homepage embeds the existing IndexTTS2 Hugging Face demo for browser-based testing.
The official IndexTTS 2.5 project page publishes research details and audio comparisons. Integration will wait for a suitable official runnable release.
No. The embedded interactive demo currently runs IndexTTS2. This page is an information page for the official IndexTTS 2.5 technical report.
The official report page demonstrates Chinese, English, Japanese, Spanish and Arabic.
Use the official project page, the arXiv technical report and the official IndexTTS repository linked below.