What Is Gemini Omni?
Gemini Omni is Google’s new family of models that fuse the company’s Gemini reasoning capabilities with creative generation. The first model, Omni Flash, focuses on video — accepting mixed inputs (images, audio, existing video clips, text) and producing coherent video output. Unlike Veo, which is primarily text-to-video generation, Omni is designed as a creation-plus-editing system with persistent conversational context.
How Does Omni Video Editing Work?
Omni lets users edit videos conversationally across multiple turns. Each instruction builds on the last — a user can start by changing the environment, then shift the camera angle, then add objects, all within the same chat session without regenerating from scratch. This “vibe coding for video” approach is Omni’s key differentiator from tools like Runway or Kling that require dedicated editing interfaces.
What Can Omni Actually Generate?
In demos, Omni showed the ability to transform a simple video of someone drawing a circle into complex animations, change a violinist’s environment while keeping the musician intact, and generate specific objects for each letter of the alphabet using Gemini’s world knowledge. Clips are capped at 10 seconds in the initial rollout, with longer durations planned.
Who Gets Access?
Omni Flash is rolling out to Google AI Plus, Pro, and Ultra subscribers globally through the Gemini app and Google Flow. It’s also available at no cost on YouTube Shorts and the YouTube Create App. Developer API access is expected in the coming weeks. All Omni-generated videos include SynthID digital watermarks.