Research

The science of emotion

Explore our publications, models, and datasets pushing the boundaries of empathic AI.

#1

in naturalness and expressivity

600+tags

of emotions and voice characteristics detected

250ms

speech LLM latency

Our Research

What the benchmarks reveal about voice AI quality

Every finding comes from the same human-grounded methodology we use for customers, the dimensions that determine whether voice AI works in the real world.

Naturalness

Naturalness is the dimension most voice AI gets wrong, and the one users feel immediately

Our human studies show consistent, measurable differences in how natural voice AI models sound. These gaps predict whether users keep engaging or disengage within the first few turns.

  • Authentic speech rhythms and pauses signal presence, not processing
  • Natural intonation varies with content, flat delivery breaks trust
  • Breathing and cadence distinguish voice AI that sounds alive from voice AI that sounds read
Voice naturalness, RW-Voice-EQ Bench
0.01.02.03.04.05.0cascade-2cascade-2gemini-3.1-flas…gemini-3.1-flash-livegemini-2.5-flas…gemini-2.5-flash-native-latestcascade-1cascade-1gpt-realtime-mi…gpt-realtime-minigpt-realtime-2gpt-realtime-2
Emotional Calibration

Most voice AI can detect emotion. Fewer respond to it appropriately

Our research measures the gap between a model's emotional perception and its behavioral response, a gap that's consequential in production and almost universally underestimated.

  • Does the model recognize frustration and respond with patience, not scripts?
  • Does it match energy when a caller is engaged or excited?
  • Does it offer reassurance when uncertainty is detected in the voice?
Emotion alignment, RW-Voice-EQ Bench
0.01.02.03.04.05.0gemini-3.1-flas…gemini-3.1-flash-livegpt-realtime-2gpt-realtime-2gemini-2.5-flas…gemini-2.5-flash-native-latestagentagentqwen3-omni-30b-…qwen3-omni-30b-a3bcascade-1cascade-1
Expressiveness

Expressiveness is what separates voice AI that connects from voice AI that performs

Our leaderboard measures the range and nuance of emotional expression across models, not whether the voice sounds good, but whether it conveys the right feeling at the right moment, with the right intensity.

  • Appropriate warmth for good news, not generic positivity
  • Genuine concern when discussing problems, not performed sympathy
  • Tonal range across humor, empathy, and authority within a single conversation
Expressiveness, RW-Voice-EQ Bench
0.01.02.03.04.05.0gemini-3.1-flashgemini-3.1-flashgemini-2.5-flashgemini-2.5-flashgemini-2.5-progemini-2.5-prolightning_v3.1lightning_v3.1gpt-4o-mini-ttsgpt-4o-mini-ttsaura-2aura-2
Precision Reliability

The most consequential failures happen on the inputs that matter most in production

Our precision benchmarks test pronunciation accuracy on the content types that actually appear in real deployments: financial figures, dates, measurements, medical terminology. This is where the gap between models is most costly.

  • The local mycologist explained that consuming just one fourth plus one fourth equals one half ounce of the misidentified death caps could prove fatal within forty eight hours.
  • Most businesses close in the late afternoon from between two until four thirty or five o'clock when it can get hot.
  • On december fifteenth two thousand seven, Dennis Kucinich raised one hundred thirty one thousand four hundred dollars from approximately one thousand six hundred donors.
Reliability, RW-Voice-EQ Bench
0.01.02.03.04.05.0gemini-2.5-progemini-2.5-progemini-3.1-flashgemini-3.1-flashgemini-2.5-flashgemini-2.5-flashaura-2aura-2eleven_v3eleven_v3sonic-3.5sonic-3.5
Expression Measurement

How accurately can a model identify the emotion being expressed, not just the words?

Our expression measurement research measures whether models can name the feeling behind a voice across 48+ emotion dimensions. This is the scientific foundation behind our Expression Measurement API and the basis for grounding all our evaluations in emotional expression.

  • Emotion categories: joy, sadness, anger, fear, surprise
  • Subtle states: hesitation, relief, mild amusement, tense calm
  • Complex emotions: bittersweet nostalgia, guarded optimism, conflicted urgency
Speech Emotion Recognition, RW-Voice-EQ Bench
0.00.20.40.60.81.0gemini-3.1-pro-…gemini-3.1-pro-previewgemini-3.5-flashgemini-3.5-flashgpt-audio-minigpt-audio-minigemini-2.5-progemini-2.5-progpt-audiogpt-audiogpt-audio-1.5gpt-audio-1.5
Instruction Following

Instruction following is the test of whether a model understands voice, not just text

Asking for “a whisper, like sharing a secret” requires understanding what that means acoustically, not just semantically. Our benchmark tests vocal instructions across emotion, style, and character, and measures how faithfully each model delivers.

Acting / Role-fit, RW-Voice-EQ Bench
0.01.02.03.04.05.0gemini-3.1-flashgemini-3.1-flashgemini-2.5-flashgemini-2.5-flashgemini-2.5-progemini-2.5-protts-1tts-1Qwen3-TTS-12Hz-…Qwen3-TTS-12Hz-1.7BVibeVoice-1.5BVibeVoice-1.5B

Recent Publications

Peer-reviewed insights

View all
Jul 2026

RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems

DA
Alice Baird
Jeff
+11
David Ayllon, Alice Baird, Jeffrey Brooks and 11 more

Current voice AI benchmarks typically evaluate isolated capabilities such as speech intelligibility, word error rate, or text-based dialogue quality, but they rarely test whether systems harness the acoustic information that distinguishes spoken language from its textual representation. To this end, we introduce the Real World Voice EQ Bench, a multidimensional benchmark for evaluating voice AI across text-to-speech (TTS), speech-to-speech (STS), speech understanding (SU), and automatic speech recognition (ASR). Our evaluations indicate that performance is highly dimension-specific. For TTS, naturalness, expressiveness, identity stability, and reliability are largely independent evaluation dimensions. For STS, access to audio does not guarantee use of vocal affect, and some agents remain largely transcript-driven. For SU, models perform unevenly across paralinguistic tasks. For ASR, real world accent, emotion, noise, and conversational conditions expose failures that are not captured by established clean-speech benchmarks. Together, these results show that voice AI should be evaluated as a profile of acoustic, expressive, interactional, and robustness capabilities rather than by a single aggregate score.

arXiv·May 2026

The 2026 ACII Dyadic Conversations (DaiKon) Workshop & Challenge

PT
Alice Baird
Jeff
+6
Panagiotis Tzirakis, Alice Baird, Jeffrey Brooks and 6 more

The 2026 ACII Dyadic Conversations (ACII-DaiKon) Workshop & Challenge introduces a benchmark for modeling interpersonal affect and social dynamics in dyadic conversations. Although conversational affect modeling has advanced rapidly, most benchmarks remain speaker-centric and underrepresent coupled, time-evolving processes between partners, including directional influence, conversational timing coordination, and rapport development. To address this gap, ACII-DaiKon presents three coordinated sub-challenges built on a shared dataset: (1) directional interpersonal influence prediction, (2) turn-taking prediction (next-speaker and time-to-next-speech), and (3) rapport trajectory prediction across full interactions. The challenge is built on the Hume-DaiKon dataset, comprising 945 dyadic conversations (743.4 hours of audiovisual data) collected under naturalistic conditions across five languages. The benchmark supports multimodal modeling, temporal reasoning, and cross-context generalization through fixed train/validation/test splits, standardized metrics, and released baseline systems. Evaluation uses Concordance Correlation Coefficient (CCC), Pearson correlation, Macro-F1, and Mean Absolute Error (MAE) depending on the sub-challenge. Baseline experiments establish initial reference performance, with best test results of 0.40 CCC and 0.50 Pearson for influence prediction, 0.66 Macro-F1 and 1.50~s MAE for turn-taking, and 0.68 CCC and 0.70 Pearson for rapport trajectory modeling. These results indicate that while current methods capture coarse dyadic patterns, robust modeling of directional dependence and long-horizon interpersonal dynamics remains challenging. The workshop provides a shared platform for rigorous comparison and cross-disciplinary discussion on data validity, evaluation protocols, and culturally aware modeling for dyadic interaction.

arXiv·Feb 2026

TADA: A Generative Framework for Speech Modeling via Text-Acoustic Dual Alignment (Under Review)

TD
SR
AG
+6
Trung Dang, Sharath Rao, Ananya Gupta and 6 more

Modern Text-to-Speech (TTS) systems increasingly leverage Large Language Model (LLM) architectures to achieve scalable, high-fidelity, zero-shot generation. However, these systems typically rely on fixed-frame-rate acoustic tokenization, resulting in speech sequences that are significantly longer than, and asynchronous with their corresponding text. Beyond computational inefficiency, this sequence length disparity often triggers hallucinations in TTS and amplifies the modality gap in spoken language modeling (SLM). In this paper, we propose a novel tokenization scheme that establishes one-to-one synchronization between continuous acoustic features and text tokens, enabling unified, single-stream modeling within an LLM. We demonstrate that these synchronous tokens maintain high-fidelity audio reconstruction and can be effectively modeled in a latent space by a large language model with a flow matching head. Moreover, the ability to seamlessly toggle speech modality within the context enables text-only guidance--a technique that blends logits from text-only and text-speech modes to flexibly bridge the gap toward text-only LLM intelligence. Experimental results indicate that our approach achieves performance competitive with state-of-the-art TTS and SLM systems while virtually eliminating content hallucinations and preserving linguistic integrity, all at a significantly reduced inference cost.

Everything your model needs

Why Our Datasets

World-class data for pre-training and fine-tuning your emotion AI models, backed by years of scientific research.

Contact Research

Ethically Sourced

All data collected with informed consent and rigorous privacy protections.

Globally Diverse

Representative samples across cultures, ages, genders, and demographics.

Expert Annotated

Labeled by trained researchers using validated scientific frameworks.

Research Ready

Clean, structured formats optimized for modern ML pipelines.

Explore Our Data

See the structure behind the emotion science

From voice AI training data to multimodal expression datasets, explore our full data catalog.

Browse all datasets

Scientific Foundation

Built on decades of research

Our datasets are grounded in peer-reviewed emotion science, developed in collaboration with leading researchers in psychology, affective computing, and machine learning.

View publications

Peer-reviewed methods

Built on 53+ publications in affective science and validated by independent researchers.

Neuroscience-informed

Datasets designed around how the brain actually processes and expresses emotion.

Validated accuracy

Benchmarked against gold-standard datasets with documented performance metrics.

Continuous updates

Regularly refined with new research findings and expanded training data.

Dataset Validation

Proven in production

View all case studies
Screenshot 2025 04 07 at 4.25.08 Pm 3

Niantic Spatial × Hume AI: Creating Interactive & Spatially Aware AI Companions

In partnership with Snap Inc. (hardware) and Hume AI (voice), Niantic Spatial has developed location-aware companions for Spectacles, blending Snap Inc.’s AR glasses, Niantic Spatial’s Large Geospatial Model, and Hume’s Empathic Voice Interface (EVI) for natural, emotionally intelligent conversation. Niantic Spatial, the team pioneering AI that understands the physical world, is showcasing a compelling glimpse of what can happen when spatial intelligence and augmented reality meet.

Read case study
Gaf Logo

GAF Powers Professional Training with Hume’s Text-to-Speech

To support their extensive training programs and marketing initiatives, GAF leverages Hume's text-to-speech technology to make internal training videos and marketing voiceovers. Our partnership addresses several key needs: Professional training content: Delivering consistent, high-quality audio for thousands of contractors and employees. Marketing collateral: Producing engaging voiceovers for promotional content and product demonstrations. Scalable production: Generating content without the logistics and cost of traditional voice recording. Hume's voice design also proved ideal for GAF. The platform's natural, expressive voices maintain the authoritative yet approachable tone that GAF needs to communicate with contractors, retailers, and customers. Unlike synthetic voices that can sound robotic or overly casual, Hume's TTS technology delivers the polished, trustworthy quality expected from an industry leader.

Read case study
Coconot Logo 3.0

Hume AI powers conversational learning with Coconote

While traditional note-taking apps require students to manually scroll and search through content, Coconote is creating interactive study experiences through conversational AI. Coconote’s voice chat feature, powered by Hume's EVI, helps users transform static notes into dynamic conversations. Students can: Ask natural questions about their lecture content Receive contextual explanations referencing specific notes, and Engage in quiz-style conversations for active learning—all through natural voice interaction.

Read case study

Research Areas

Where Hume enables research

From fundamental affective computing to applied behavioral research, our datasets power studies across the full spectrum of emotion science.

Affective Computing

Study how AI systems can recognize, interpret, and respond to human emotions across modalities.

Human-AI Interaction

Research the dynamics of emotional exchange between humans and AI systems.

Psychology & Behavior

Use expression analysis to study human behavior, mental health, and psychological phenomena.

Speech & Language

Analyze prosodic features, sentiment, and emotional expression in human communication.

Multimodal Learning

Explore how emotion manifests simultaneously across face, voice, and language.

Ethics & AI Safety

Study the ethical implications of emotionally-aware AI systems and develop guidelines.

From the Blog

Latest research updates

View all

License our datasets

Access world-class expression datasets and collaborate with our team on advancing emotion AI.

Stay in the loop

Get the latest on empathic AI research, product updates, and company news.

Join the community

Connect with other developers, share projects, and get help from the team.

Join our Discord