Gemini Omni: from "text-to-video" to "anything-to-video"
Google already had Veo for video generation. But at I/O 2026, DeepMind CEO Demis Hassabis introduced something fundamentally different: Gemini Omni โ a model that combines Gemini's reasoning intelligence with the rendering capabilities of Google's best media models.
The difference from Veo is not just branding. Veo accepted text and images. Omni accepts text, images, audio, and video together in a single prompt. More importantly, it reasons across all of them โ understanding physics, gravity, kinetic energy, cultural context, and real-world knowledge to produce video that behaves like the real world.
Google DeepMind director Nicole Brichtova put it clearly: "It's the next step towards combining the intelligence of Gemini with the rendering capabilities of our media models."
Omni is an umbrella that brings together the capabilities of Veo, Nano Banana, and the core Gemini reasoning models under one family. The first release, Gemini Omni Flash, is available today in the Gemini app, Google Flow, and YouTube Shorts.
Conversational video editing โ describe it, and the model applies it
The most striking feature of Omni is not just generation โ it is editing through conversation. Upload a video, then describe the changes you want. Change the point of view. Adjust the lighting. Add a new character. Transform the environment into something entirely different.
Every instruction builds on the last. Characters stay consistent, physics holds up, and the scene remembers what came before. Google demoed this by starting with raw footage of a person walking down a hallway, then moving them through different scenes โ all without changing the original movement or pacing.
In the keynote demo, a prompt for "a claymation explainer of protein folding" produced a stop-motion-style educational video with an AI-generated voice-over that accurately described amino acids and protein structure. It combined scientific knowledge with creative rendering โ something previous video models could not do.
Where does this leave Veo?
Our earlier analysis of the Omni leak predicted that Veo would be the engine and Omni the experience. That is essentially what happened. The Verge reported it directly: "Unlike Google's Veo model, which is only text to video," Omni accepts multimodal inputs and reasons across them.
Omni is not a Veo rebrand โ it is a step above. Veo was a video generation model. Omni is a world model that can generate video as one of its outputs. The distinction matters because Omni's long-term vision extends beyond video: eventually, it will be able to generate images from audio, audio from video, and more.
For creators, the practical split looks like this:
- Gemini Omni โ consumer-facing, chat-native, available in the Gemini app. Best for quick video creation and conversational editing.
- Google Flow โ the professional creative studio. More granular controls, batch processing, multi-agent workflows.
- YouTube Shorts + YouTube Create โ Omni Flash available directly in YouTube for creators making short-form content.
Flow gets agents: one image, sixteen videos, batch editing
Google Flow also got significant upgrades at I/O 2026. The biggest: agentic capabilities. Flow's agent can now run multiple requests simultaneously โ take a single image and transform it into 16 different video clips. Then batch-edit all 16 at once, turning them all into nighttime shots or applying a unified style change.
With Omni powering the generation, Flow can now completely change the environment, add visual effects, and insert new characters โ all while preserving the original performance. Google demonstrated this by transforming a live-action video of a guitarist into different artistic styles while the music and performance stayed intact.
Flow is also now available as a mobile app (Android in beta, iOS coming soon), and Google Flow Music โ the audio companion โ can transform a simple piano recording into a professionally mixed, complete musical track.
One input image โ 16 unique video clips. Flow's agent runs them simultaneously instead of one at a time.
Apply a style or environment change across all 16 clips at once. Turn everything to nighttime, change the season, or shift the palette.
Change the world around the performer without altering their movement. Swap the scene, keep the action.
Record a simple riff or melody, and Flow Music transforms it into a complete, professionally mixed track.
AI watermarking goes industry-wide
With AI video becoming this realistic, the verification problem becomes urgent. Google addressed this head-on at I/O: SynthID, its AI watermarking technology, is expanding beyond Google's own products.
OpenAI, Kakao, and ElevenLabs are now adopting SynthID. That is significant โ Google's biggest AI rival is voluntarily using Google's watermarking standard. All videos created with Omni are automatically watermarked with SynthID.
On the consumer side, SynthID verification is coming to Google Search and Chrome. You will be able to use Circle to Search or right-click on an image or video to check whether it was generated by AI. Google is also supporting C2PA Content Credentials, which let you verify whether content is an unaltered original from a camera or has been modified by AI tools.
YouTube search gets an AI overhaul
A smaller but notable announcement for video creators: Ask YouTube. Instead of keyword-based search, you can now ask YouTube complex, natural-language questions โ and it returns relevant video clips, not just video links.
For example, "how do I clean the sensor on a Canon R5?" will surface the exact moment in a video where that specific question is answered, and start playback there. It handles follow-up questions too. Currently rolling out to YouTube Premium subscribers in the US.
What is available right now
The constraint creators should know about
Omni Flash currently generates up to 10 seconds of video per clip. Google says this is not a model limitation but a deliberate choice for the initial rollout โ longer durations are coming. But for now, if you need 30- or 60-second clips, you will need to stitch multiple generations together.
The usage limit question from the leak is also confirmed. Google is moving from daily prompt limits to a "compute-used" model that factors the complexity of your prompt, the features you use, and the length of your chat. Limits refresh every five hours until you reach your weekly cap. Video generation will eat through those limits fast.
The new AI Ultra plan at $100/month gives 5x higher limits than AI Pro. The previous $250 plan is now $200 with the same capabilities. If you are doing serious video work with Omni, the $100 tier is likely the minimum you will need.
What this means for video creators
- Omni replaces Veo as the primary video model. It is multimodal, conversational, and understands physics. If you were waiting for Google's answer to Seedance 2, Kling, and Sora โ this is it.
- Conversational editing is the real feature. Generating a video is one step. Being able to say "now change the lighting to golden hour" and having the model apply it while keeping everything else intact โ that is the workflow shift.
- Flow is now the pro layer. Batch generation (1 image โ 16 videos) and batch editing make Flow genuinely useful for content at scale. The mobile app lowers the barrier further.
- Budget for usage limits. The compute-used model means video generation costs more than text or image generation. Plan your iterations.
- SynthID matters. Every Omni-generated video is watermarked. As verification tools roll out to Search and Chrome, AI-generated content will become increasingly identifiable. Build that into your content strategy.
The spaghetti test is fully passed. The question is no longer quality โ it is cost, speed, and whether chat-native editing can replace a real editing timeline. For short-form creators, it probably can. For long-form, Omni Flash is the starting gun. The race is on.
One gap Omni does not close: format. You still generate in one aspect ratio, but need to publish across YouTube (16:9), Shorts and TikTok (9:16), and Instagram (4:5, 1:1). That last mile โ resizing, reframing, and padding without losing the subject โ is exactly what OmniFit handles, and it pairs naturally with anything coming out of Omni or Flow.