Summarize this article with
Key takeaways
- Gemini Omni is one model that takes text, images, audio, and video in and generates all four back, instead of chaining separate tools.
- Conversational editing lets you direct it by talking, and each note builds on the last.
- Every image or video it makes carries a SynthID watermark and C2PA credentials that survive cropping and re-export, and cannot be removed.
- It runs on Google AI Plus ($20/mo) or AI Ultra ($100/mo), with Omni Flash capped at 10-second clips for now.
Google announced Gemini Omni at I/O on May 19, and it's a real shift in how these tools work. Omni is a single model that takes in text, images, audio, and video, and generates all four right back. Most setups today chain a handful of separate tools together to pull that off.
Here's what one unified model gets you, and the catch to know before a client job.
One model, no handoffs
This matters more than it sounds. When a single model handles everything, nothing gets lost handing off between an image tool, a video tool, and a voice tool. You give it a photo, a voice clip, and a line of direction all at once, and it works from the whole picture. That leaves you with fewer tools in the chain and fewer places your idea gets watered down along the way.
You direct it by talking
The feature we like most is conversational editing. Omni holds the thread while you work, so you can make a clip, tell it "warmer light, slow the walk down," and watch it adjust just those things. Your last note carries into the next one, and each direction builds on the ones before it. After a few rounds it starts to feel like directing a shot.
It understands physics
Google made a big deal of Omni's grip on physics: gravity, weight, the way fluid and fabric actually move. That's usually where AI video gives itself away, with the floaty, rubbery motion you can spot in a second. If Omni is as good here as the demos look, it gets a lot closer to real.
Avatars of you
Omni can turn your face and voice into an avatar that talks and acts from a typed instruction. You put yourself on camera without ever filming. For anyone who makes a lot of talking-head content, that's a real time saver. Test it carefully before you trust it with your face.
The catch: it's all watermarked
Here's the part Google doesn't lead with. Every image or video Omni makes carries a SynthID watermark and C2PA credentials that tag it as AI, and those survive cropping, compression, and re-export. For most work that's no problem at all.
If you or a client ever need the AI kept quiet, though, know that it's baked in and can't be removed, so plan around it.
How to get it
To use Omni, you'll need one of Google's AI plans:
- AI Plus: $20 a month
- AI Ultra: $100 a month
- Omni Flash: tops out at 10-second clips for now
That clip limit is a temporary launch cap, and Google has said it'll climb, so don't read it as the ceiling.
Our take
Omni is the most interesting release of the season, and the one-model, talk-to-it approach is clearly where this is heading. We're still running it against real work before we call it a daily driver. If you're already on a Google plan, it's the first thing worth opening this month. For the rest of what dropped, see our roundup of what's new in AI.





