Developers can create real-time voice applications with Gemini 3.8 Live and 3.5 Transcribe
New Gemini Audio models have been introduced for developers aiming to enhance conversational experiences utilizing the Gemini API and Google AI Studio.
Today, Google DeepMind has launched the Gemini Live models within the Gemini API and Google AI Studio, broadening their developer tools for crafting real-time, voice-centric product experiences:
Gemini 3.8 Live and 3.8 Live Extended Thinking
Gemini 3.8 Live marks a significant advancement in speech-to-speech capabilities, efficiently managing tasks while maintaining a dialogue. For more complex inquiries, the 3.8 Live Extended Thinking model offers in-depth reasoning, leading the Artificial Analysis’ Speech-to-Speech leaderboard.
Gemini 3.5 Transcribe
This dedicated speech-to-text model delivers precise transcription in over 85 languages, achieving a Word Error Rate (WER) of 4.0% for streaming and 2.6% for non-streaming since its release last month.
The newly launched models, Gemini 3.8 Live and 3.8 Live Extended Thinking, enable developers to create voice agents capable of executing tasks while maintaining conversational continuity. Key features include:
- 3.8 Live Extended Thinking supports configurable thinking for managing complex, multi-step reasoning in the background, narrating progress within the main dialogue.
These models signify a substantial enhancement from prior live models, offering a more efficient alternative to layered architectures. Available through the Live API, they are cost-effectively priced at $0.005 per minute for audio input and $0.018 per minute for audio output, facilitating scalable voice application development with exceptional performance.
Developers can also integrate these models via platforms such as Agora, Fishjam, LiveKit, LangChain, Pipecat, Vercel, and Vision Agents, which handle media streaming infrastructure for practical deployment:
Real-time speech comprehension is vital for voice-first interfaces. Last month, Gemini 3.5 Transcribe was launched to provide low-latency transcription with high accuracy, achieving a 4.0% WER alongside notable features:
- Supporting over 85 languages, it offers a robust listening engine for various voice applications and tasks like rapid captioning, call center operations, and real-time audio analysis. The Interactions API allows users to transcribe audio files up to one hour with structured timestamps and speaker identification. Our developer guide offers further insight.
To begin, explore the models in ai.studio/live, clone sample applications from GitHub, or enhance your agent with our live API skills.
Additionally, you can create audio experiences using our speech and music generation models, all accessible via the Gemini API:
We look forward to seeing the innovative applications you will build!
Sign up for our newsletters for updates on products, events, special offers, and more.
Check your inbox to confirm your subscription. You can choose to subscribe with a different email address at any time.
Your information will be managed in compliance with Google’s privacy policy, and you may opt-out whenever you wish.