Article
How does Hume Octave compare to other leading TTS models like Elevenlabs?

Our state-of-the-art speech-language model Octave is the first LLM explicitly designed for text-to-speech (TTS). Unlike traditional TTS systems, Octave understands the meaning and context of the content it generates, leading to more natural, fluid, and contextually appropriate speech.
But how does Octave stack up against other leading TTS models available today?
To answer this question, we conducted a study with 1,200 participants, comparing Octave against other prominent models using the same voices featured in Huggingface's popular TTS Arena. Each participant created their own unique prompts, ensuring diversity and realism reflective of actual user interactions.
Participants then listened to pairs of speech samples generated randomly by two competing models, rating them head-to-head in a blind evaluation. Every participant rated at least five pairs of speech samples.
In head-to-head matchups, Hume won 68% of the time. Elevenlabs came in second with a 60.9% win rate, and Papla in third with 54.9% win rate.

These results indicate that Octave produces higher-quality, more natural-sounding speech, even when handling short, isolated pieces of user-generated text. However, this study didn't even take advantage of Octave's greatest differentiator -- maintaining coherent and engaging speech across longer texts or narratives. For long-form content, such as audiobooks or interactive voice experiences with complex character dialogues and distinct emotional expressions, we anticipate Octave’s contextual intelligence to demonstrate an even greater advantage.
Hume remains committed to building models that can better anticipate users’ needs by deeply understanding and adapting to human expression. Explore Octave yourself at the TTS playground, and keep an eye out for future updates.
Keep reading

Behind Hume’s Expression Measurement Models for Face and Voice
Editor's note: This post covers the science, development and evaluation of Hume’s state of the art Expression API, that enables developers to enrich their applications or research pipelines with real-time measurement of voice and facial expressions. Facial expressions and vocal tone are central to how we communicate. They convey emotions such as amusement, interest, frustration, and surprise, often without naming those feelings in words. Understanding these signals is part of how we connect with one another and respond to what others express.
Oct 6, 2026

Evaluating Google’s multi-speaker TTS: A case study in why private evaluations matter
Text-to-speech systems were originally built to read text aloud in a single voice. As voice AI expands into audiobooks, game dialogue, and advertising, these systems are being asked to do more: generate conversations between multiple speakers. That takes more than generating two distinct voices and stitching their lines together. The speakers need to sound like they are responding to one another, with natural timing, changes in tone, and smooth handoffs. Each voice must remain distinct and consistent while contributing to a believable conversation.
Sharath Rao, Kimberly Lo, Alice Baird / Sep 29, 2026
