Google DeepMind’s Gemini 3.8 Live Models: What Visual Creators Should Watch
Google DeepMind introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two near-real-time voice conversation models aimed at voice agents, visual grounding, and complex multi-step tasks. For visual creators, the notable angle is how voice, vision context, and background tool use may shape faster creative workflows.
Google DeepMind has introduced two near-real-time voice conversation models: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. The announcement, published on the Google DeepMind Blog, positions the models as building blocks for more capable voice agents and more natural AI conversations across developer, enterprise, and consumer surfaces.
Source Facts
According to Google DeepMind, Gemini 3.8 Live is designed for scale and cost efficiency, combining conversational intelligence with fluid dialogue and visual grounding. Gemini 3.8 Live Extended Thinking is positioned for higher-complexity tasks, with stronger intelligence and multi-step reasoning.
The company says Gemini 3.8 Live Extended Thinking scored 82.6 on Artificial Analysis’s Speech to Speech Quality Index, ranking first overall in that index. Google DeepMind also reports 68.6% on τ-Voice, 35.1% on Sierra’s τ-Voice-banking benchmark, and 97.7% on Big Bench Audio. For Gemini 3.8 Live, the source states that it ranked second in Speech Agent Arena and remains cost efficient for developers and enterprises.
Google DeepMind also says Gemini 3.8 Live can process visual input at near-real-time speed, use that context in conversation, automatically detect and switch among 97 supported languages, and continue talking while tool calls and API calls run in the background. The source describes examples such as employee onboarding with visual context, chess guidance using visual context and reasoning, and business-plan creation through natural speech.
For Extended Thinking, Google DeepMind describes a model that can speak while reasoning, use early verbal signals such as “let me check…,” and narrate progress during multi-step background tasks. The source gives examples including turning raw sketches and near-real-time spoken feedback into usable React components, and coordinating multi-step bookings with asynchronous function calls.
Availability is described as a gradual rollout from today. Gemini 3.8 Live is listed for developers in Gemini API and Google AI Studio, enterprises through private preview in Gemini Enterprise and later Gemini Enterprise for Customer Experience, and everyone in Search Live. Gemini 3.8 Live Extended Thinking is listed for developers in Gemini API and Google AI Studio, enterprises through private preview in Gemini Enterprise and later Gemini Enterprise for Customer Experience and Google Workspace enterprise customers, and everyone in Gemini Live, plus Google AI Pro and Ultra subscribers for Docs in Workspace and all Google AI subscribers for Gmail and Keep.
Google DeepMind says all audio generated by its AI products is watermarked with SynthID.
AiPix Analysis
For visual creators, the most relevant shift is not just voice chat. It is the combination of spoken iteration, visual context, and background execution. If a model can look at a sketch, discuss intent in real time, and trigger tools without stopping the conversation, creative workflows may become less menu-driven and more collaborative.
That matters for image and video creation because many creator tasks are iterative: refining a portrait, adjusting an ID photo, removing a background, improving image quality, or exploring a wedding concept. AiPix’s own product direction centers on AI Canvas and one-click creation paths, so the broader industry move toward voice-guided, visually grounded agents reinforces a clear pattern: creators increasingly expect to direct work naturally while the system handles technical steps in the background.
The source does not show AiPix testing Gemini 3.8 Live or Gemini 3.8 Live Extended Thinking, so no performance comparison should be inferred. The practical takeaway is strategic: voice-first creative agents are becoming more credible when they combine visual awareness, tool orchestration, and transparent progress feedback.
Sources
- Google DeepMind 发布 Gemini 3.8 Live 和 3.8 Live Extended ThinkingGoogle DeepMind:Blog(RSS)
Continue in AiPix
Turn this update into a practical image or canvas workflow.
Open AI Canvas