Article
EVI 2 + Claude Computer Use
You can now control a computer with just your voice. In just a few hours, we combined the EVI 2 API with Claude's new computer use functionality. Here’s how it works.

You can now control a computer with just your voice. In just a few hours, we combined the EVI 2 API with Claude's new computer use functionality.
Here’s how we did it.
1. We start with Replit's Anthropic Computer Use template
Replit provided an excellent demo of how to use Anthropic's experimental computer use capabilities to control Firefox. We only had to make a few modifications.
2. We have EVI process speech in real-time
We replace text input with the EVI API, capture the user’s voice.
3. We send instructions to the agentic computer control loop
When we get transcriptions to EVI, we send them to Claude.
4. We have EVI explain its actions with voice
When we get back responses from Claude, EVI uses it to voice Claude’s actions as they’re being carried out.
5. We allow EVI to interrupt Claude to change course
We take advantage of EVI’s native interruptibility, allowing it to update Claude’s instructions in real time.
You can find all of our code here.
With new tools that allow LLMs to control devices, we're seeing a glimpse of the future of AI interfaces and agents.
Keep reading

Behind Hume’s Expression Measurement Models for Face and Voice
Editor's note: This post covers the science, development and evaluation of Hume’s state of the art Expression API, that enables developers to enrich their applications or research pipelines with real-time measurement of voice and facial expressions. Facial expressions and vocal tone are central to how we communicate. They convey emotions such as amusement, interest, frustration, and surprise, often without naming those feelings in words. Understanding these signals is part of how we connect with one another and respond to what others express.
Oct 6, 2026

Evaluating Google’s multi-speaker TTS: A case study in why private evaluations matter
Text-to-speech systems were originally built to read text aloud in a single voice. As voice AI expands into audiobooks, game dialogue, and advertising, these systems are being asked to do more: generate conversations between multiple speakers. That takes more than generating two distinct voices and stitching their lines together. The speakers need to sound like they are responding to one another, with natural timing, changes in tone, and smooth handoffs. Each voice must remain distinct and consistent while contributing to a believable conversation.
Sharath Rao, Kimberly Lo, Alice Baird / Sep 29, 2026