OmniFit logo OmniFit Blog home
Google I/O 2026 ยท News + analysis

Gemini Omni is here. Video editing just moved into chat.

Google I/O 2026 dropped the biggest AI video announcement of the year: Gemini Omni, a world model that generates and edits video from any input โ€” text, image, audio, or existing footage โ€” all through conversation. Google Flow gets agentic batch editing, and SynthID goes industry-wide.

๐Ÿ“… Published: May 20, 2026
AI video Gemini Omni Google Flow Google I/O 2026 By Noah Bennett
Veo was text-to-video. Omni is everything-to-video โ€” and you edit by talking. The model understands physics, remembers your characters, and is available today.
What happened

Google launched Gemini Omni โ€” a multimodal world model that generates and edits video from any combination of text, images, audio, and video inputs. Omni Flash is live today.

Why it matters

Video editing becomes conversational. You describe changes in plain language, and the model applies them while keeping characters, physics, and scene continuity intact.

What to watch

Omni Flash starts at 10-second clips. Omni Pro is coming later with longer durations. The real question: can the usage limits support a real creative workflow?

The headline

Gemini Omni: from "text-to-video" to "anything-to-video"

Google already had Veo for video generation. But at I/O 2026, DeepMind CEO Demis Hassabis introduced something fundamentally different: Gemini Omni โ€” a model that combines Gemini's reasoning intelligence with the rendering capabilities of Google's best media models.

The difference from Veo is not just branding. Veo accepted text and images. Omni accepts text, images, audio, and video together in a single prompt. More importantly, it reasons across all of them โ€” understanding physics, gravity, kinetic energy, cultural context, and real-world knowledge to produce video that behaves like the real world.

Google DeepMind director Nicole Brichtova put it clearly: "It's the next step towards combining the intelligence of Gemini with the rendering capabilities of our media models."

"With world models, AI is moving from predicting text to simulating reality. Gemini Omni is the next step in that direction." โ€” Sundar Pichai, Google CEO

Omni is an umbrella that brings together the capabilities of Veo, Nano Banana, and the core Gemini reasoning models under one family. The first release, Gemini Omni Flash, is available today in the Gemini app, Google Flow, and YouTube Shorts.

How it works

Conversational video editing โ€” describe it, and the model applies it

The most striking feature of Omni is not just generation โ€” it is editing through conversation. Upload a video, then describe the changes you want. Change the point of view. Adjust the lighting. Add a new character. Transform the environment into something entirely different.

Every instruction builds on the last. Characters stay consistent, physics holds up, and the scene remembers what came before. Google demoed this by starting with raw footage of a person walking down a hallway, then moving them through different scenes โ€” all without changing the original movement or pacing.

Multimodal input: combine text, images, audio, and video in a single prompt.
Conversational editing: describe changes in plain language โ€” the model applies them iteratively.
Physics understanding: the model understands kinetic energy, gravity, and realistic object interaction.
Character consistency: characters and settings persist across multiple editing turns.
Text rendering in video: accurate text on surfaces โ€” slogans, math notation, product labels.
Digital avatars: create a personal AI avatar from a recording, then use it in generated videos.

In the keynote demo, a prompt for "a claymation explainer of protein folding" produced a stop-motion-style educational video with an AI-generated voice-over that accurately described amino acids and protein structure. It combined scientific knowledge with creative rendering โ€” something previous video models could not do.

Watch: hands-on with Gemini Omni video

Gavin and Kevin from AI Explained tested Omni in person at I/O โ€” including the physics demos, video editing, and the limitations they found.

Omni vs Veo

Where does this leave Veo?

Our earlier analysis of the Omni leak predicted that Veo would be the engine and Omni the experience. That is essentially what happened. The Verge reported it directly: "Unlike Google's Veo model, which is only text to video," Omni accepts multimodal inputs and reasons across them.

Omni is not a Veo rebrand โ€” it is a step above. Veo was a video generation model. Omni is a world model that can generate video as one of its outputs. The distinction matters because Omni's long-term vision extends beyond video: eventually, it will be able to generate images from audio, audio from video, and more.

For creators, the practical split looks like this:

  • Gemini Omni โ€” consumer-facing, chat-native, available in the Gemini app. Best for quick video creation and conversational editing.
  • Google Flow โ€” the professional creative studio. More granular controls, batch processing, multi-agent workflows.
  • YouTube Shorts + YouTube Create โ€” Omni Flash available directly in YouTube for creators making short-form content.
Google Flow

Flow gets agents: one image, sixteen videos, batch editing

Google Flow also got significant upgrades at I/O 2026. The biggest: agentic capabilities. Flow's agent can now run multiple requests simultaneously โ€” take a single image and transform it into 16 different video clips. Then batch-edit all 16 at once, turning them all into nighttime shots or applying a unified style change.

With Omni powering the generation, Flow can now completely change the environment, add visual effects, and insert new characters โ€” all while preserving the original performance. Google demonstrated this by transforming a live-action video of a guitarist into different artistic styles while the music and performance stayed intact.

Flow is also now available as a mobile app (Android in beta, iOS coming soon), and Google Flow Music โ€” the audio companion โ€” can transform a simple piano recording into a professionally mixed, complete musical track.

Batch generation

One input image โ†’ 16 unique video clips. Flow's agent runs them simultaneously instead of one at a time.

Batch editing

Apply a style or environment change across all 16 clips at once. Turn everything to nighttime, change the season, or shift the palette.

Performance preservation

Change the world around the performer without altering their movement. Swap the scene, keep the action.

Flow Music

Record a simple riff or melody, and Flow Music transforms it into a complete, professionally mixed track.

Watch: Agents and Gemini Omni in Google Flow

Google's official walkthrough of Flow's new agentic capabilities โ€” batch generation, style transfer, and Omni-powered editing in action.

Watch: The Verge's I/O 2026 keynote recap

The full keynote condensed to 35 minutes. Gemini Omni starts at 3:34, Google Flow updates at 25:08, and Flow Music at 27:16.

SynthID

AI watermarking goes industry-wide

With AI video becoming this realistic, the verification problem becomes urgent. Google addressed this head-on at I/O: SynthID, its AI watermarking technology, is expanding beyond Google's own products.

OpenAI, Kakao, and ElevenLabs are now adopting SynthID. That is significant โ€” Google's biggest AI rival is voluntarily using Google's watermarking standard. All videos created with Omni are automatically watermarked with SynthID.

On the consumer side, SynthID verification is coming to Google Search and Chrome. You will be able to use Circle to Search or right-click on an image or video to check whether it was generated by AI. Google is also supporting C2PA Content Credentials, which let you verify whether content is an unaltered original from a camera or has been modified by AI tools.

Ask YouTube

YouTube search gets an AI overhaul

A smaller but notable announcement for video creators: Ask YouTube. Instead of keyword-based search, you can now ask YouTube complex, natural-language questions โ€” and it returns relevant video clips, not just video links.

For example, "how do I clean the sensor on a Canon R5?" will surface the exact moment in a video where that specific question is answered, and start playback there. It handles follow-up questions too. Currently rolling out to YouTube Premium subscribers in the US.

Availability

What is available right now

Gemini Omni Flash โ€” live today in the Gemini app, Google Flow, and YouTube Shorts for Google AI Plus, Pro, and Ultra subscribers globally.
Google Flow app โ€” Android beta available now, iOS coming soon.
Google Flow Music app โ€” iOS available now, Android coming soon.
Ask YouTube โ€” rolling out to YouTube Premium subscribers in the US.
SynthID in Search and Chrome โ€” rolling out starting today.
Omni API โ€” coming "soon" for developers; no exact date.
Omni Pro โ€” coming later; Google says it will launch when it delivers "a step change above Flash."
Limits

The constraint creators should know about

Omni Flash currently generates up to 10 seconds of video per clip. Google says this is not a model limitation but a deliberate choice for the initial rollout โ€” longer durations are coming. But for now, if you need 30- or 60-second clips, you will need to stitch multiple generations together.

The usage limit question from the leak is also confirmed. Google is moving from daily prompt limits to a "compute-used" model that factors the complexity of your prompt, the features you use, and the length of your chat. Limits refresh every five hours until you reach your weekly cap. Video generation will eat through those limits fast.

The new AI Ultra plan at $100/month gives 5x higher limits than AI Pro. The previous $250 plan is now $200 with the same capabilities. If you are doing serious video work with Omni, the $100 tier is likely the minimum you will need.

Bottom line

What this means for video creators

  • Omni replaces Veo as the primary video model. It is multimodal, conversational, and understands physics. If you were waiting for Google's answer to Seedance 2, Kling, and Sora โ€” this is it.
  • Conversational editing is the real feature. Generating a video is one step. Being able to say "now change the lighting to golden hour" and having the model apply it while keeping everything else intact โ€” that is the workflow shift.
  • Flow is now the pro layer. Batch generation (1 image โ†’ 16 videos) and batch editing make Flow genuinely useful for content at scale. The mobile app lowers the barrier further.
  • Budget for usage limits. The compute-used model means video generation costs more than text or image generation. Plan your iterations.
  • SynthID matters. Every Omni-generated video is watermarked. As verification tools roll out to Search and Chrome, AI-generated content will become increasingly identifiable. Build that into your content strategy.

The spaghetti test is fully passed. The question is no longer quality โ€” it is cost, speed, and whether chat-native editing can replace a real editing timeline. For short-form creators, it probably can. For long-form, Omni Flash is the starting gun. The race is on.

One gap Omni does not close: format. You still generate in one aspect ratio, but need to publish across YouTube (16:9), Shorts and TikTok (9:16), and Instagram (4:5, 1:1). That last mile โ€” resizing, reframing, and padding without losing the subject โ€” is exactly what OmniFit handles, and it pairs naturally with anything coming out of Omni or Flow.

Key announcements

Model

Gemini Omni Flash โ€” live today

What it does

Generates + edits video from any input via conversation

Clip length

Up to 10 seconds (longer coming)

Available in

Gemini app, Flow, YouTube Shorts

Pricing

AI Plus, Pro, Ultra ($100/mo new tier)

Watermark

SynthID embedded in all output

Also announced for video

Google Flow agents โ€” 1 image โ†’ 16 videos, batch style edits

Flow Music โ€” simple audio โ†’ full produced track

Digital avatars โ€” create your own AI avatar for Shorts

Ask YouTube โ€” AI search with clip-level results

SynthID โ€” adopted by OpenAI, Kakao, ElevenLabs

After you generate

Omni outputs one aspect ratio. Publishing across platforms means resizing every clip for YouTube, Shorts, TikTok, and Instagram.

OmniFit auto-reframes and resizes video for every platform in one step โ€” a natural next step after generating with Omni or Flow.