Hugging Face launches Open TTS Leaderboard for open-source speech models

Hugging Face has published the Open TTS Leaderboard, a ranking of open-source and multilingual text-to-speech (TTS) models built on objective metrics rather than human-vote arenas. The post opens with the problem: as of Sep 30, 2026, the Hugging Face Hub hosts more than 8K TTS models, but evaluation is fragmented and unstandardized. Human preference scores such as MOS or MUSHRA are the gold standard, and several arena-style leaderboards have become reference points. Those arenas show users outputs from two models, ask which is better, and compute an Elo score after enough votes, typically with the Bradley-Terry model.
The authors argue that arenas cannot scale to keep up with the pace of TTS releases. They say this may partly explain why open-source models are underrepresented: as of Sep 30, 2026, only 16 of the 92 models on Artificial Analysis are open-weights, with a similar skew on Voice Arena. They suggest this likely reflects practical factors: an API model needs little more than an API key, while an open model must be hosted and served by the arena operator, and commercial providers have more reason to seek placement than open-source authors. A second limitation is voter consistency: no arena can ensure the same voters apply the same criteria of "better" over time, and even one person's preferences shift (the post cites Heraclitus: "A man cannot step into the same river twice").
The leaderboard measures three complementary things. Intelligibility is the word or character error rate (WER and CER) between the prompt and the transcript of the generated audio, transcribed with Qwen3 ASR, described as the top-ranking open-source model on the Open ASR Leaderboard. Speed is inverse real-time factor (RTFx) for batched offline inference on an H200 GPU, plus time-to-first-audio (TTFA) for streaming latency at batch size 1, on an H200 GPU and on CPU. Speaker similarity (SIM) is the cosine similarity between WavLM speaker embeddings of the generated audio and the reference clip. The authors say this cuts the time to evaluate a model from a couple of weeks of vote collection to a couple of hours.
They are explicit that the leaderboard does not replace human preference ranking. WER is a proxy for intelligibility and speaker similarity estimates voice identity preservation; neither directly measures naturalness, expressiveness or listener preference. They add that the metrics can even inform voting-based leaderboards about which models to include.
By default, models are ranked by macro-average WER on the English splits of Seed TTS Eval and CV3 Eval (zero shot). On that ranking, hexgrad/Kokoro-82M, Supertone/supertonic-3 and fishaudio/s2-pro lead on English WER averaged over the two splits. Pareto plots show which models balance WER, batched inference speed (RTFx) and size. Because English performance does not necessarily carry over to other languages, users can toggle languages to rank models on multilingual results. Seed TTS Eval has audio only for English and Chinese, so other languages are scored on CV3 Eval (zero shot) alone. Chinese, Japanese and Korean are character-based, so CER is reported for them, and the "Average WER" across languages is a macro-average across languages. The authors name k2-fsa/OmniVoice, fishaudio/s2-pro and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 as strong multilingual models.
A "Voice cloning" toggle compares models that support cloning in the selected languages. It adds a SIM column and two more Pareto plots showing the tradeoff between SIM, batched inference speed and size. The average WER of some models, such as bosonai/higgs-tts-3-4b and openbmb/VoxCPM2, improves under voice cloning, that is, when a reference audio is provided.
A "Listen" tab lets users compare the generated outputs behind the metrics: pick a language and dataset, choose whether to compare voice cloning, and optionally pick models or listen to a random selection. Users can also vote on outputs. The authors say they may include that community vote data on the leaderboard later, and ask voters to log in with a Hugging Face account to help weed out spam and bots.
A "Streaming" tab ranks models by TTFA, the wait between prompting a model and getting playable audio, which matters for voice agents and other interactive apps. For streaming models it is the time until the first audio chunk arrives; for non-streaming models it is the time until the whole utterance is generated, since playback cannot start earlier. Every model runs one audio at a time (batch size 1) on the same 50 English prompts from CV3-Eval, on the same hardware and in its default voice. The first 3 runs are dropped as warm-up and the median TTFA across the rest is reported. The default view uses an H200 GPU; CPU results exist for a small but growing set of models. The authors single out kyutai/pocket-tts as a great model for streaming on both GPU and CPU.
The team says it wants the leaderboard shaped by the community and asks for feedback on which datasets, models and metrics to add. For now it focuses on open-source models and on multilingual evaluation. The evaluation scripts are to be open-sourced soon, in the manner of the Open ASR Leaderboard repo, so feedback can come through GitHub Issues and pull requests.
Key facts
- Hugging Face's Open TTS Leaderboard ranks open-source and multilingual TTS models on objective metrics: WER/CER via Qwen3 ASR, speed (RTFx and TTFA on an H200 GPU, with some CPU results), and speaker similarity (SIM) from WavLM embeddings.
- The authors say evaluating a model takes a couple of hours instead of the couple of weeks needed to collect votes for an arena.
- As of Sep 30, 2026, the Hub has more than 8K TTS models, while only 16 of the 92 models on Artificial Analysis are open-weights.
- Default ranking is macro-average English WER on Seed TTS Eval and CV3 Eval (zero shot); hexgrad/Kokoro-82M, Supertone/supertonic-3 and fishaudio/s2-pro lead.
- It does not replace human preference ranking: the authors say WER and speaker similarity do not directly measure naturalness, expressiveness or listener preference.
Why it matters
Open TTS models are released faster than anyone can collect enough votes to rank them. The post says more than 8K TTS models sit on the Hugging Face Hub as of Sep 30, 2026, yet only 16 of the 92 models on Artificial Analysis are open-weights, with a similar skew on Voice Arena. The authors argue arenas cannot scale to this pace. A leaderboard that scores models automatically, in hours, gives open models a place in the comparison they otherwise tend to miss, and adds multilingual and voice cloning views that English-only results do not cover.
Who it affects
Developers choosing an open TTS model for an app, especially one that needs a non-English language, voice cloning or low streaming latency for a voice agent. Open-source model authors, whose releases have been neglected by arena-style evaluation according to the post. Operators of voting-based leaderboards, who the authors say could use the metrics to decide which models to include.
How to use it
The default view ranks models by English WER; toggle languages to rank on multilingual results, and toggle "Voice cloning" to compare models that support it in the selected languages, which also shows the SIM column. The Pareto plots show the balance between WER, batched speed and model size. The "Listen" tab lets you hear the generated outputs behind the metrics and vote on them; voting requires logging in with a Hugging Face account. The "Streaming" tab ranks models by time-to-first-audio on an H200 GPU, with CPU results for a small set of models. The evaluation scripts are to be open-sourced soon, and the authors invite feedback on datasets, models and metrics through GitHub Issues and pull requests.
How solid is it
The setup is described in detail: a fixed ASR model (Qwen3 ASR) for transcripts, named datasets (Seed TTS Eval and CV3 Eval), fixed hardware (H200 GPU), and for streaming the same 50 English CV3-Eval prompts at batch size 1, default voice, with the first 3 runs dropped as warm-up and the median reported. This is the publisher's own description of its own leaderboard, and the evaluation scripts are not yet open-sourced, so the results cannot yet be reproduced independently. The post gives no numeric WER, CER, SIM, RTFx or TTFA values for any model, and does not state how many models the leaderboard covers. Its explanation for why open models are scarce on arenas is hedged by the authors themselves ("may partly explain", "likely reflects").
Risks and caveats
The authors say plainly that the leaderboard does not replace human preference ranking. WER and speaker similarity do not directly measure naturalness, expressiveness or listener preference, so a model that tops the table may not be the one listeners like best. The default ranking rests on English WER, which the post itself says does not necessarily translate to other languages. Chinese, Japanese and Korean are scored by CER, and the multilingual average is a macro-average across languages. For non-streaming models TTFA is the time to generate the whole utterance, so it is not directly comparable to a streaming model's first-chunk time. Community votes are not yet part of the ranking; the authors say they may include them later.
“While human preference is the ultimate decider, arenas cannot scale to keep up with the pace of TTS releases.”
— Hugging Face, Open TTS Leaderboard blog post