Now liveExplore Expression APIs
hume.ai logo

Infrastructure to improve voice AI.

The bar is human. Hume finds the conversations that break your system, diagnoses why, and feeds each failure into a continuous improvement loop.

ToneWarmClippedMeasuredUrgentFlatGentleBrightFirmTiming18% overlap240 ms gap6% overlap910 ms gap2% overlap1.4 s gap24% overlap620 ms gapEmotionAmusement 0.62Annoyance 0.84Anxiety 0.79Calm 0.71Relief 0.46Frustration 0.92Joy 0.55Concern 0.70Interruption7 barge-ins0 barge-ins11 barge-ins1 barge-in5 barge-ins2 barge-ins8 barge-ins1 barge-inBackground noiseSNR 6 dBSNR 31 dBSNR 4 dBSNR 14 dBSNR 9 dBSNR 11 dBSNR 12 dBSNR 28 dBAccentMidwest USGlaswegianLagosChileanKansaiQuébécoisPunjabiEstuaryLanguageEnglishSpanishMandarinHindiArabicPortugueseFrenchYorubaContextFollow-upFirst callEscalationHandoverRenewalComplaintCheck-inRepeatAdaptation2 turnsNo shift4 turns1 turn3 turnsNo shift2 turns1 turnIntentDisputeConfirmCancelAskEscalateRescheduleVerifyComplain

The measurement gap

There is more to voice than words.

A task can succeed while the conversation fails. Humans judge voice AI on timing, tone, emotional expression, interruption, accent, context, and adaptation. Hume measures all of it.

  • Does the system understand the user?
  • Does the user feel understood?

Trusted by

Leading AI labs
5 of 7
Languages supported
50+
Speech samples
100M+
Emotion science research
10+ years

From defining good to knowing what to improve

Generate targeted tests, stress-test your voice AI, diagnose failures, and validate improvements, all in one evaluation platform.

01 / Generate

Automatically turn real-world use cases into targeted evaluation suites.

VoiceEQ analyzes your conversation data to identify use cases, generate test cases, and define the scenarios and success criteria that matter.

Your evaluation suite

Criteria list

Metric set

Experiment plan

Explore the VoiceEQ evaluation platform

See how leading voice models perform in the real world.

Real-World VoiceEQ evaluates voice systems under realistic conditions across the dimensions people experience directly, combining objective measures with human judgment to reveal model strengths and performance tradeoffs.

Select a model

Illustrative, not real data

Dimension detail · GPT-4o Realtime

Expression
61
Naturalness
77
Timing & turn-taking
71
Contextual appropriateness
54
Transcript accuracy
84

One evaluation platform for models and agents.

Develop better voice models

Create and enrich data, compare checkpoints, benchmark alternatives, and diagnose weaknesses across TTS, speech-to-speech, ASR, and cascaded systems.

  • TTS

    Baseline
    0.62
    Candidate
    0.81
  • Speech-to-speech

    Baseline
    0.49
    Candidate
    0.64
  • ASR

    Baseline
    0.80
    Candidate
    0.88
  • Cascaded systems

    Baseline
    0.61
    Candidate
    0.57
Explore model development evaluation

Deploy better voice agents

Select the right model, define use-case-specific quality, validate realistic interactions with automated and human evaluation, and turn findings into focused improvements.

  • Healthcare

    Auto
    0.84
    Human
    0.76
  • Contact Center

    Auto
    0.88
    Human
    0.91
  • Financial Services

    Auto
    0.71
    Human
    0.68
  • Retail

    Auto
    0.66
    Human
    0.59
Explore voice-agent evaluation

Voice-native infrastructure for better evaluation and training

Access the research systems Hume built to develop its own voice models. Use them through APIs or in your environment to enrich audio, measure expression, and incorporate human judgment.

Prism Pipeline

Prepare audio for evaluation and training

Transform unstructured audio into searchable, expression-rich inputs.

Expression APIs

Measure what transcripts miss

Measure vocal expression, timing, and conversational dynamics.

Human Feedback API

Ground evaluations in human judgment

Run structured evaluations with trained, vetted raters to assess the qualities people experience directly.

Explore the infrastructure

Learn how real-world voice performance is measured.

Introducing Real World VoiceEQ

Read the benchmark overview

Real World VoiceEQ Methodology

Read the technical report(opens in a new tab)

Research & Insights

Explore the work behind Hume

Make emotionally intelligent voice AI measurable.

Define what good means for your use case, evaluate the moments that matter, and turn evidence into focused improvement.