
Gemini 3.8 Live: Build a Multilingual Voice Agent
Summary
Build a real-time voice agent with Gemini 3.8 Live Extended Thinking.
Google shipped two new voice models this week, and one of them just took the #1 spot on Artificial Analysis' Speech to Speech Quality Index with a score of 82.6. Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are built for near real-time voice conversations that can reason mid-sentence, switch between 97 languages without restarting the call, and call your backend functions in the background while still talking to the user.
That last part is the interesting bit for builders. Most "voice AI" demos you've seen either freeze while the model thinks, or skip thinking entirely and hallucinate an answer. The Extended Thinking variant does neither: it says something like "let me check that" out loud, keeps reasoning in the background, and only interrupts itself when the answer is ready. That's a genuinely new capability, not a repackaged one, and it changes how you'd architect a voice agent.
Keep reading — it's free
Enter your email to keep reading — plus the best of AI & tech, daily. Free, forever.
Already a member? Sign in
Comments
Be the first to comment