Hugging Face has launched an Open TTS Leaderboard to compare text-to-speech and voice-cloning models with repeatable automated tests. The project responds to a widening evaluation gap: the Hub lists more than 8,000 speech models, while human-voting arenas can take weeks to add and rank a new entry.
The leaderboard measures intelligibility through word or character error rates after Qwen3 ASR transcribes generated speech. It reports batched throughput and time to first audio on an Nvidia H200, with CPU results for a smaller set of models. Voice-cloning entries also receive a speaker-similarity score based on WavLM embeddings.
Using objective metrics can reduce an evaluation from a couple of weeks to a couple of hours. It also gives open-weight models a more practical path to inclusion because an arena operator does not need to host each one long enough to collect many votes. Users can filter by language, compare cloning results and listen to the underlying clips.
The numbers do not directly measure naturalness, expressiveness or listener preference. Hugging Face therefore presents the tool as a complement to human voting, not a replacement. The team plans to open-source its evaluation scripts and is asking the community to propose additional datasets, models and metrics.