The science of emotion
Explore our publications, models, and datasets pushing the boundaries of empathic AI.
in naturalness and expressivity
of emotions and voice characteristics detected
speech LLM latency
Our Research
What the benchmarks reveal about voice AI quality
Every finding comes from the same human-grounded methodology we use for customers, the dimensions that determine whether voice AI works in the real world.
Naturalness is the dimension most voice AI gets wrong, and the one users feel immediately
Our human studies show consistent, measurable differences in how natural voice AI models sound. These gaps predict whether users keep engaging or disengage within the first few turns.
- •Authentic speech rhythms and pauses signal presence, not processing
- •Natural intonation varies with content, flat delivery breaks trust
- •Breathing and cadence distinguish voice AI that sounds alive from voice AI that sounds read
Most voice AI can detect emotion. Fewer respond to it appropriately
Our research measures the gap between a model's emotional perception and its behavioral response, a gap that's consequential in production and almost universally underestimated.
- •Does the model recognize frustration and respond with patience, not scripts?
- •Does it match energy when a caller is engaged or excited?
- •Does it offer reassurance when uncertainty is detected in the voice?
Expressiveness is what separates voice AI that connects from voice AI that performs
Our leaderboard measures the range and nuance of emotional expression across models, not whether the voice sounds good, but whether it conveys the right feeling at the right moment, with the right intensity.
- •Appropriate warmth for good news, not generic positivity
- •Genuine concern when discussing problems, not performed sympathy
- •Tonal range across humor, empathy, and authority within a single conversation
The most consequential failures happen on the inputs that matter most in production
Our precision benchmarks test pronunciation accuracy on the content types that actually appear in real deployments: financial figures, dates, measurements, medical terminology. This is where the gap between models is most costly.
- The local mycologist explained that consuming just one fourth plus one fourth equals one half ounce of the misidentified death caps could prove fatal within forty eight hours.
- Most businesses close in the late afternoon from between two until four thirty or five o'clock when it can get hot.
- On december fifteenth two thousand seven, Dennis Kucinich raised one hundred thirty one thousand four hundred dollars from approximately one thousand six hundred donors.
How accurately can a model identify the emotion being expressed, not just the words?
Our expression measurement research measures whether models can name the feeling behind a voice across 48+ emotion dimensions. This is the scientific foundation behind our Expression Measurement API and the basis for grounding all our evaluations in emotional expression.
- •Emotion categories: joy, sadness, anger, fear, surprise
- •Subtle states: hesitation, relief, mild amusement, tense calm
- •Complex emotions: bittersweet nostalgia, guarded optimism, conflicted urgency
Instruction following is the test of whether a model understands voice, not just text
Asking for “a whisper, like sharing a secret” requires understanding what that means acoustically, not just semantically. Our benchmark tests vocal instructions across emotion, style, and character, and measures how faithfully each model delivers.
Our Research
What the benchmarks reveal about voice AI quality
Every finding on this page comes from the same methodology we use for customers, human-grounded, reproducible, and scored against real emotional experience. These are the dimensions that actually determine whether voice AI works in the real world.
Naturalness is the dimension most voice AI gets wrong, and the one users feel immediately
Our human studies show consistent, measurable differences in how natural voice AI models sound. These gaps predict whether users keep engaging or disengage within the first few turns.
- •Authentic speech rhythms and pauses signal presence, not processing
- •Natural intonation varies with content, flat delivery breaks trust
- •Breathing and cadence distinguish voice AI that sounds alive from voice AI that sounds read
Most voice AI can detect emotion. Fewer respond to it appropriately
Our research measures the gap between a model's emotional perception and its behavioral response, a gap that's consequential in production and almost universally underestimated.
- •Does the model recognize frustration and respond with patience, not scripts?
- •Does it match energy when a caller is engaged or excited?
- •Does it offer reassurance when uncertainty is detected in the voice?
Expressiveness is what separates voice AI that connects from voice AI that performs
Our leaderboard measures the range and nuance of emotional expression across models, not whether the voice sounds good, but whether it conveys the right feeling at the right moment, with the right intensity.
- •Appropriate warmth for good news, not generic positivity
- •Genuine concern when discussing problems, not performed sympathy
- •Tonal range across humor, empathy, and authority within a single conversation
The most consequential failures happen on the inputs that matter most in production
Our precision benchmarks test pronunciation accuracy on the content types that actually appear in real deployments: financial figures, dates, measurements, medical terminology. This is where the gap between models is most costly.
- The local mycologist explained that consuming just one fourth plus one fourth equals one half ounce of the misidentified death caps could prove fatal within forty eight hours.
- Most businesses close in the late afternoon from between two until four thirty or five o'clock when it can get hot.
- On december fifteenth two thousand seven, Dennis Kucinich raised one hundred thirty one thousand four hundred dollars from approximately one thousand six hundred donors.
How accurately can a model identify the emotion being expressed, not just the words?
Our expression measurement research measures whether models can name the feeling behind a voice across 48+ emotion dimensions. This is the scientific foundation behind our Expression Measurement API and the basis for grounding all our evaluations in emotional expression.
- •Emotion categories: joy, sadness, anger, fear, surprise
- •Subtle states: hesitation, relief, mild amusement, tense calm
- •Complex emotions: bittersweet nostalgia, guarded optimism, conflicted urgency
Instruction following is the test of whether a model understands voice, not just text
Asking for “a whisper, like sharing a secret” requires understanding what that means acoustically, not just semantically. Our benchmark tests vocal instructions across emotion, style, and character, and measures how faithfully each model delivers.
Recent Publications
Peer-reviewed insights
RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems


Current voice AI benchmarks typically evaluate isolated capabilities such as speech intelligibility, word error rate, or text-based dialogue quality, but they rarely test whether systems harness the acoustic information that distinguishes spoken language from its textual representation. To this end, we introduce the Real World Voice EQ Bench, a multidimensional benchmark for evaluating voice AI across text-to-speech (TTS), speech-to-speech (STS), speech understanding (SU), and automatic speech recognition (ASR). Our evaluations indicate that performance is highly dimension-specific. For TTS, naturalness, expressiveness, identity stability, and reliability are largely independent evaluation dimensions. For STS, access to audio does not guarantee use of vocal affect, and some agents remain largely transcript-driven. For SU, models perform unevenly across paralinguistic tasks. For ASR, real world accent, emotion, noise, and conversational conditions expose failures that are not captured by established clean-speech benchmarks. Together, these results show that voice AI should be evaluated as a profile of acoustic, expressive, interactional, and robustness capabilities rather than by a single aggregate score.
The 2026 ACII Dyadic Conversations (DaiKon) Workshop & Challenge


The 2026 ACII Dyadic Conversations (ACII-DaiKon) Workshop & Challenge introduces a benchmark for modeling interpersonal affect and social dynamics in dyadic conversations. Although conversational affect modeling has advanced rapidly, most benchmarks remain speaker-centric and underrepresent coupled, time-evolving processes between partners, including directional influence, conversational timing coordination, and rapport development. To address this gap, ACII-DaiKon presents three coordinated sub-challenges built on a shared dataset: (1) directional interpersonal influence prediction, (2) turn-taking prediction (next-speaker and time-to-next-speech), and (3) rapport trajectory prediction across full interactions. The challenge is built on the Hume-DaiKon dataset, comprising 945 dyadic conversations (743.4 hours of audiovisual data) collected under naturalistic conditions across five languages. The benchmark supports multimodal modeling, temporal reasoning, and cross-context generalization through fixed train/validation/test splits, standardized metrics, and released baseline systems. Evaluation uses Concordance Correlation Coefficient (CCC), Pearson correlation, Macro-F1, and Mean Absolute Error (MAE) depending on the sub-challenge. Baseline experiments establish initial reference performance, with best test results of 0.40 CCC and 0.50 Pearson for influence prediction, 0.66 Macro-F1 and 1.50~s MAE for turn-taking, and 0.68 CCC and 0.70 Pearson for rapport trajectory modeling. These results indicate that while current methods capture coarse dyadic patterns, robust modeling of directional dependence and long-horizon interpersonal dynamics remains challenging. The workshop provides a shared platform for rigorous comparison and cross-disciplinary discussion on data validity, evaluation protocols, and culturally aware modeling for dyadic interaction.
TADA: A Generative Framework for Speech Modeling via Text-Acoustic Dual Alignment (Under Review)
Modern Text-to-Speech (TTS) systems increasingly leverage Large Language Model (LLM) architectures to achieve scalable, high-fidelity, zero-shot generation. However, these systems typically rely on fixed-frame-rate acoustic tokenization, resulting in speech sequences that are significantly longer than, and asynchronous with their corresponding text. Beyond computational inefficiency, this sequence length disparity often triggers hallucinations in TTS and amplifies the modality gap in spoken language modeling (SLM). In this paper, we propose a novel tokenization scheme that establishes one-to-one synchronization between continuous acoustic features and text tokens, enabling unified, single-stream modeling within an LLM. We demonstrate that these synchronous tokens maintain high-fidelity audio reconstruction and can be effectively modeled in a latent space by a large language model with a flow matching head. Moreover, the ability to seamlessly toggle speech modality within the context enables text-only guidance--a technique that blends logits from text-only and text-speech modes to flexibly bridge the gap toward text-only LLM intelligence. Experimental results indicate that our approach achieves performance competitive with state-of-the-art TTS and SLM systems while virtually eliminating content hallucinations and preserving linguistic integrity, all at a significantly reduced inference cost.
Everything your model needs
Why Our Datasets
World-class data for pre-training and fine-tuning your emotion AI models, backed by years of scientific research.
Ethically Sourced
All data collected with informed consent and rigorous privacy protections.
Globally Diverse
Representative samples across cultures, ages, genders, and demographics.
Expert Annotated
Labeled by trained researchers using validated scientific frameworks.
Research Ready
Clean, structured formats optimized for modern ML pipelines.
Explore Our Data
See the structure behind the emotion science
From voice AI training data to multimodal expression datasets, explore our full data catalog.
Browse all datasetsScientific Foundation
Built on decades of research
Our datasets are grounded in peer-reviewed emotion science, developed in collaboration with leading researchers in psychology, affective computing, and machine learning.
Peer-reviewed methods
Built on 53+ publications in affective science and validated by independent researchers.
Neuroscience-informed
Datasets designed around how the brain actually processes and expresses emotion.
Validated accuracy
Benchmarked against gold-standard datasets with documented performance metrics.
Continuous updates
Regularly refined with new research findings and expanded training data.
Dataset Validation
Proven in production
Niantic Spatial × Hume AI: Creating Interactive & Spatially Aware AI Companions
In partnership with Snap Inc. (hardware) and Hume AI (voice), Niantic Spatial has developed location-aware companions for Spectacles, blending Snap Inc.’s AR glasses, Niantic Spatial’s Large Geospatial Model, and Hume’s Empathic Voice Interface (EVI) for natural, emotionally intelligent conversation. Niantic Spatial, the team pioneering AI that understands the physical world, is showcasing a compelling glimpse of what can happen when spatial intelligence and augmented reality meet.
GAF Powers Professional Training with Hume’s Text-to-Speech
To support their extensive training programs and marketing initiatives, GAF leverages Hume's text-to-speech technology to make internal training videos and marketing voiceovers. Our partnership addresses several key needs: Professional training content: Delivering consistent, high-quality audio for thousands of contractors and employees. Marketing collateral: Producing engaging voiceovers for promotional content and product demonstrations. Scalable production: Generating content without the logistics and cost of traditional voice recording. Hume's voice design also proved ideal for GAF. The platform's natural, expressive voices maintain the authoritative yet approachable tone that GAF needs to communicate with contractors, retailers, and customers. Unlike synthetic voices that can sound robotic or overly casual, Hume's TTS technology delivers the polished, trustworthy quality expected from an industry leader.
Hume AI powers conversational learning with Coconote
While traditional note-taking apps require students to manually scroll and search through content, Coconote is creating interactive study experiences through conversational AI. Coconote’s voice chat feature, powered by Hume's EVI, helps users transform static notes into dynamic conversations. Students can: Ask natural questions about their lecture content Receive contextual explanations referencing specific notes, and Engage in quiz-style conversations for active learning—all through natural voice interaction.
Research Areas
Where Hume enables research
From fundamental affective computing to applied behavioral research, our datasets power studies across the full spectrum of emotion science.
Affective Computing
Study how AI systems can recognize, interpret, and respond to human emotions across modalities.
Human-AI Interaction
Research the dynamics of emotional exchange between humans and AI systems.
Psychology & Behavior
Use expression analysis to study human behavior, mental health, and psychological phenomena.
Speech & Language
Analyze prosodic features, sentiment, and emotional expression in human communication.
Multimodal Learning
Explore how emotion manifests simultaneously across face, voice, and language.
Ethics & AI Safety
Study the ethical implications of emotionally-aware AI systems and develop guidelines.
From the Blog
Latest research updates

Introducing Real World VoiceEQ: Measuring the Human Quality of Voice AI
Jul 14, 2026

Disentangling Emotion from Voice: A Cross-Product Sampling Approach for Expressive Voice Data
Humans can separate how they feel from how they speak. Voice models still struggle to do the same.
May 27, 2026

Introducing the ACII 2026 Dyadic Contest (DaiKon) Workshop & Challenge
Apr 9, 2026
License our datasets
Access world-class expression datasets and collaborate with our team on advancing emotion AI.