Google moved two families of models out of preview in the same week, both aimed at media rather than text.
Gemini Omni Flash
gemini-omni-1.1-flash reached general availability on August 27 with three new capabilities: video extension, interpolation between images, and resolution control up to 4K. The preview endpoint it replaces is scheduled to shut down on September 30, 2026, so anyone still calling it has a hard date.
Resolution control up to 4K is the practical headline. Generated video that tops out below broadcast resolution is a demo; video that reaches 4K is something a production team can actually use without an upscaling step.
Speech to text
A day earlier, on August 26, gemini-3.5-transcribe and gemini-3.5-transcribe-live both reached GA. They support 85+ languages with speaker diarization and word-level timestamps.
Those last two features are what separate a transcription toy from a transcription product. Diarization tells you who spoke; word-level timestamps let you build search, captions and editing on top. The live variant covers real-time use, where the constraint is latency rather than accuracy.
Why it matters
Speech-to-text with diarization across 85+ languages, available as a general API, puts direct pressure on the standalone transcription market. Tools in that category have historically competed on exactly those two features.
The same applies to video generation. Every capability that moves from a specialised vendor into a general-purpose API from Google, OpenAI or Anthropic narrows what the specialist has left to sell — usually workflow, integrations and support rather than the underlying model.
What to check
If you use the Omni Flash preview endpoint, migrate before September 30. And if you pay for a transcription service, it is worth pricing the same job through the API — the answer may not favour switching, but the gap is smaller than it was.