AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get tech for your team delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Hugging Face has introduced the Open TTS Leaderboard to evaluate text-to-speech models using speech accuracy, inference speed and speaker similarity metrics. The project says these measurements can be produced in hours, but stresses that they do not replace human judgments of naturalness or listener preference.

Hugging Face has launched the Open TTS Leaderboard, a new system for comparing text-to-speech models on speech accuracy, inference speed and speaker similarity. The project is aimed at a fast-growing field where, according to Hugging Face, the Hub held more than 8,000 TTS models as of September 30, 2026, while existing comparisons remained fragmented and slow to update.

The leaderboard reports word and character error rates by comparing generated speech transcripts with their prompts, using Qwen3 automatic speech recognition. It also measures offline generation speed with inverse real-time factor on an H200 GPU, and streaming responsiveness through time-to-first-audio on both an H200 GPU and a CPU. For voice cloning, it adds a speaker-similarity score based on WavLM embeddings of generated speech and reference audio.

Its default rankings use macro-average word error rate across English splits from Seed TTS Eval and CV3 Eval. Users can select other languages and turn on a voice-cloning view. Hugging Face names k2-fsa/OmniVoice, fishaudio/s2-pro and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 as strong multilingual models; for English error rates, it lists hexgrad/Kokoro-82M, Supertone/supertonic-3 and fishaudio/s2-pro among the leaders. These are leaderboard results, not a general verdict on which systems sound best.

A separate “Listen” tab lets users compare generated audio and submit preferences. Hugging Face says an evaluation can take a couple of hours with its objective metrics, compared with weeks for arena voting. Community votes may be included later, and users are asked to sign in with a Hugging Face account to limit spam and bot submissions.

At a glance
announcementWhen: Announced in material dated September 3…
The developmentHugging Face has launched an open leaderboard that evaluates text-to-speech models across multiple languages using objective performance metrics and offers audio comparisons for community feedback.

A Faster View of Open TTS

The leaderboard addresses a coverage problem in model comparisons. Human-vote arenas can reveal which outputs people prefer, but collecting enough votes takes time, and operating an arena requires hosting the models it tests. Hugging Face says open-weight models are underrepresented on some existing rankings: as of September 30, it counted 16 open-weight models among 92 on Artificial Analysis, with a similar skew on Voice Arena. It attributes the imbalance in part to the effort required to host open models and the stronger incentive commercial providers may have to seek placement.

Faster, repeatable measurements could help researchers and developers identify models worth testing, including across languages and latency constraints. The streaming tab may be useful to teams building voice agents, where the delay before playback starts affects responsiveness. The metrics also expose tradeoffs: a model that performs well on transcription accuracy may be larger or slower, while speaker similarity and cloning performance add another dimension to comparisons.

The scores cannot settle every practical choice. Error rates estimate intelligibility through an ASR system, and speaker similarity estimates voice identity preservation; neither directly measures naturalness, expressiveness or human preference. The leaderboard is most informative as a set of complementary measurements alongside listening tests.

Amazon

multilingual text-to-speech device

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How TTS Rankings Have Worked

Text-to-speech models turn written prompts into spoken audio, and their rapid release has made consistent comparison difficult. Hugging Face points to TTS Arena v2, Artificial Analysis and Voice Arena as existing arena-style reference points. These systems present two outputs for listeners to compare, collect votes and use an Elo score, often calculated with a Bradley–Terry model, to rank models.

Such rankings capture subjective preferences directly, but they depend on enough votes and consistent evaluation criteria. Hugging Face argues that preferences can vary between voters and over time. Its new leaderboard takes a different route for much of the evaluation: standardized datasets and automated measurements are intended to make model comparisons quicker, while its listening feature leaves room for users to judge the audio themselves.

“The Open TTS Leaderboard does not replace human preference ranking.”

— Hugging Face, describing the leaderboard’s role

Amazon

voice cloning microphone

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limits of the Current Scores

The supplied announcement does not give the full model list, score tables, evaluation sample sizes or uncertainty ranges for individual results. It also does not establish how well the automatic metrics align with listener judgments across languages, accents or speaking styles. Chinese, Japanese and Korean use character error rate in the leaderboard; for languages beyond English and Chinese, the announcement says Seed TTS Eval has no audio and scores come from CV3 Eval alone.

Hugging Face says community votes may be incorporated as more feedback arrives, but does not specify a threshold or schedule. It is also unclear from the announcement how often rankings will be refreshed or how changes to models and evaluation data will be handled. The published comparisons should be read with those limits in mind.

Amazon

high accuracy speech recognition software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

More Listening and Community Votes

Users can explore model outputs in the “Listen” tab, select a language and dataset, choose whether to compare voice cloning, and vote on audio. Hugging Face asks participants to log in so it can reduce spam and bot activity. The project says it may use accumulated votes on the leaderboard, though it has not announced when or how that would happen.

Future usefulness will depend on the breadth and upkeep of the evaluations: which models are added, how often results are rerun, and whether community listening data can complement the automated scores. For now, the leaderboard provides a new set of comparisons and a route for users to assess the actual audio themselves.

Amazon

real-time speech synthesis hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Where I land

I see the Open TTS Leaderboard as a useful way to make model evaluation faster and more comparable, especially for open models and multilingual use. Its clearest contribution is not a definitive ranking of voice quality, but a way to surface accuracy, speed and cloning tradeoffs that developers can investigate further.

The strongest counterargument is that automated scores can reward systems that perform well on narrow benchmarks without sounding natural or expressive to people. ASR error rates and speaker embeddings are proxies, and the source material does not provide evidence that they reliably predict listener preference across languages. I would give the rankings more weight if Hugging Face published repeatable evaluation details and showed how the scores relate to blinded human judgments across languages and tasks. Until then, I would use them to shortlist models and make the final choice by listening.

Key Questions

What does the Open TTS Leaderboard measure?

It reports word or character error rate as a proxy for intelligibility, generation speed, time-to-first-audio for streaming responsiveness, and speaker similarity for voice cloning.

Does it identify the best-sounding TTS model?

No single metric in the announcement measures overall listener preference. The leaderboard provides automated comparisons, and its “Listen” tab lets people compare audio and submit votes.

Which models does Hugging Face highlight?

For English error rates, it lists hexgrad/Kokoro-82M, Supertone/supertonic-3 and fishaudio/s2-pro among the leaders. For multilingual performance, it names k2-fsa/OmniVoice, fishaudio/s2-pro and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 as strong models.

How are Chinese, Japanese and Korean evaluated?

The leaderboard uses character error rate for Chinese, Japanese and Korean. The announcement says Seed TTS Eval provides audio only for English and Chinese, so the other language scores come from CV3 Eval.

Can users influence the rankings?

Users can compare outputs and submit preferences through the “Listen” tab. Hugging Face says it may include the votes in the leaderboard as more are collected, but has not published a timeline or method for doing so.

Source: Hugging Face

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Grok Bot For Engineering – xAI

xAI introduces a Grok bot aimed at engineering tasks; details are limited. Here is what is confirmed and what remains unclear.

New X47.c Windows Botnet Weaponizes xAI Grok, AI API Draining – SecurityWeek

A SecurityWeek headline reports that the x47.c Windows botnet uses xAI’s Grok API. The article body and supporting details were unavailable.

Grok 4.6 In GitHub Copilot – X.ai

GitHub has added xAI’s Grok 4.6 to Copilot for paid plans, with a gradual rollout, usage-based billing and administrator controls.

“No Reason Why Everyone Should Have An Identical Claude Experience”: Anthropic’s Mods Let You Change Claude Code’s Look And Behavior – The New Stack

Anthropic’s Claude Code mods let users alter the tool’s appearance and behavior. Details on availability, scope and safeguards remain limited.