Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning
A new open-source and multilingual TTS leaderboard has been introduced to address the scalability issues of existing arena-style evaluations. Current leaderboards, such as Artificial Analysis and Voice Arena, struggle to keep up with the rapid pace of TTS releases and underrepresent open-source models due to practical hosting challenges and commercial incentives. Additionally, voter consistency is a limitation in arena-style evaluations. The new leaderboard aims to provide a more scalable and consistent evaluation method, with evaluation scripts to be open-sourced for community feedback.
时间与来源
时间显示为 UTC
显示时区:UTC
本地时区尚不可用,暂时显示 UTC。
发布当时偏移:UTC+02026年9月30日 00:00 UTC
收录当时偏移:UTC+02026年9月30日 15:00 UTC
- 发布
- 2026年9月30日 00:00
- 收录
- 2026年9月30日 15:00
- 来源类型
- 官方发布
- 档位
- 当事方
- 信源状态
- 正常
档位是按信源手工设定的编辑判断,不是逐条打分。
讨论趋势
百分比基于采集到的讨论信号,不代表新增评论数或独立参与人数。曲线仅用于同一话题在不同时段的比较。
TLDR 👉 new TTS leaderboard focused on open-source and multilingual
The pace of open-source text-to-speech (TTS) model releases has been incredible. On the Hugging Face Hub (as of Sep 30, 2026) there are more than 8K TTS models available 🚀
Evaluation, however, hasn't kept pace: it remains fragmented and unstandardized. The gold standard is human preference scores such as MOS or MUSHRA (more on metrics ). To this end, several arena-based leaderboards have established themselves as useful reference points for the community:
- TTS Arena v2
- Artificial Analysis
- Voice Arena
These arenas compare models by presenting users with TTS outputs from two models, and asking them to choose one over the other. After collecting a sufficient number of votes, an Elo score is computed to rank models, typically with the Bradley–Terry model (see Voice Arena methodology ).
While human preference is the ultimate decider, arenas cannot scale to keep up with the pace of TTS releases. This may partly explain why open-source models are underrepresented on arena-style leaderboards: as of Sep 30, 2026, only 16 of the 92 models on Artificial Analysis are open-weights, with a similar skew on Voice Arena . This likely reflects practical factors: adding an API model requires little more than an API key, whereas an open model must be hosted and served by the arena operator, and commercial providers have more reason to seek placement than open-source authors. Another limitation with arena-style evaluation is voter consistency: no arena can ensure that the same voters with the same criteria of “better” can consistently evaluate models over time. Even the preferences of a single person change over time (“A man cannot step into the same river twice” as famously said by Heraclitus).
To this end, we've built the Open TTS Leaderboard , which uses objective metrics to evaluate models on complementary aspects of performance:
- Intelligibility: word/character error rate (WER and CER) between the prompt and the generated audio's transcript, using Qwen3 ASR (top ranking open-source model on the Open ASR Leaderboard ).
- Speed: inverse real-time factor (RTFx) for batched offline inference on an H200 GPU, and time-to-first-audio (TTFA) for quantifying streaming batch size 1 latency on an H200 GPU and CPU.
- Speaker similarity by computing the cosine similarity (SIM) between WavLM speaker embeddings of the generated audio and the reference clip.
By relying on objective metrics evaluating a model drops from a couple weeks (for collecting votes) to a couple hours ⚡
Importantly, the Open TTS Leaderboard does not replace human preference ranking. ASR-based WER provides a proxy for intelligibility, while speaker similarity estimates voice identity preservation. Neither directly measures naturalness, expressiveness, or listener preference. Nevertheless, they can even inform voting-based leaderboards which models to include in their evaluations.
Our intention with this leaderboard is for it to be shaped by the community; we want to hear your feedback so the evaluations stay relevant and insightful. The next few sections give an overview of main features of the Open TTS Leaderboard.
Multilingual + voice cloning evaluation
From the default view of the leaderboard, models are ranked by macro-average WER on the English splits of Seed TTS Eval ( paper ) and CV3 Eval (zero shot) ( paper ).
hexgrad/Kokoro-82M , Supertone/supertonic-3 , and fishaudio/s2-pro lead the pack on English WER when averaged on these two splits, while the Pareto plots visualize which models strike a good balance between WER, batched inference (RTFx), and size.
English performance doesn't necessarily translate to other languages. Multiple languages can be toggled to rank models on multilingual performance. Seed TTS Eval only has audio for English and Chinese, so the other languages are simply the score on CV3 Eval (zero shot). Note that Chinese, Japanese, and Korean are character-based languages and so character error rate (CER) is reported, and the “Average WER” across languages is a macro-average across languages.
k2-fsa/OmniVoice , fishaudio/s2-pro , and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 are strong multilingual models.
By toggling “Voice cloning”, the models that support this functionality (on the selected languages) can be compared.
Moreover, a SIM column for speaker similarity now appears in the table, as well as two more Pareto plots for visualizing the tradeoff between SIM, batched inference, and size.
The average WER of some models, such as bosonai/higgs-tts-3-4b and openbmb/VoxCPM2 , improve under voice cloning, namely when a reference audio is provided.
Compare and vote on TTS outputs