Nathan Smith · 23 May 2026
Google Omni: our honest take, and how we'd actually use it for clients
Google's Gemini Omni turns any mix of inputs (a photo, a voice note, a rough script, a clip) into a finished video you refine just by talking to it. We gave it a quick mention in our I/O recap; here's the full version. As a Google Premier Partner we go through what it can actually do, where the cracks are, and exactly how we'd put it to work for clients without flooding the feed with forgettable AI video.
We gave Gemini Omni a few lines in our I/O 2026 recap and called it the announcement with the most immediate commercial pull. A fortnight on, we've had hands on it. This is the longer version: what Omni actually does, where the cracks are, and (the bit that matters to you) how we'd put it to work for clients without drowning the feed in forgettable AI video.
As a Google Premier Partner we'd rather give you the clear-eyed read than the keynote one. So here's each capability, feature by feature, with our honest take and the practical client angle attached.
What Omni actually is
Google's framing is "a model that can create anything from any input, starting with video." In practice: you feed it any mix of a photo, a voice note, a few lines of script and a short clip, and it returns a polished video. The first model in the family is Gemini Omni Flash, available to paid subscribers through the Gemini app and Google Flow, free inside YouTube Shorts and the Create app, with developer and enterprise APIs following in the coming weeks.
Our take: the headline isn't "AI makes video," we've had that for a while. It's the any input part. The gap between a client emailing us a voice memo and a rough brand photo, and a first-cut video existing, has basically collapsed. That's the genuinely new thing, and it's the thing worth getting excited and careful about in equal measure.
Conversational, cumulative editing
You don't re-prompt from scratch. You talk to the cut: "make the opening punchier," "drop the music under the voiceover," "lose the second shot," and edits stack across the conversation rather than resetting each time.
Our take: this is the feature that actually changes how a creative team works, more than the raw generation does. "Prompt to clip" is a party trick; "talk to the edit and have it remember" is a workflow. It maps onto how a producer and an editor already talk to each other, which means the learning curve is close to zero. For clients: the painful middle of social production, the round-trips of "can we just try it with a different intro," gets compressed from days to minutes. We can sit in a review call and reshape a cut live instead of booking another edit pass.
Physics and world-knowledge grounding
Omni leans on Gemini's understanding of how the world behaves (gravity, momentum, fluids) and on its general knowledge of history, science and culture, so generated motion holds together and explanatory video can actually be about something real.
Our take: this is where it pulls ahead of the "pretty but physically nonsense" generation of AI video. Liquids pour roughly right, things fall roughly right, and that's the difference between footage you can put a brand name next to and footage that screams "AI slop." We're cautious, though. "Roughly right" is not "reliably right," and the failures are unpredictable rather than consistent, which makes them harder to QA out. For clients: genuinely useful for explainer and product-mechanism content (how a thing works, why a process matters) where a grounded, physically plausible visual beats a stock animation. We'll still watch every frame before it goes near a client logo.
Avatars: your likeness, your voice
You can generate video featuring your own digital avatar, speaking in your own voice, with visual and audio consistency maintained across the clip.
Our take: commercially this is the spicy one. A founder who films one good session could, in principle, "appear" in a month of content without sitting in front of a camera again. The upside is obvious; so is the risk. An avatar that can say anything in your voice is a brand-safety and consent question before it's a production one: who signs off on the words, who holds the likeness rights, what happens if a clip is taken out of context. For clients: we'd use this only with explicit, documented consent and a human approving every script, and we'd disclose it. Done that way it's a brilliant scale lever for personal-brand and thought-leadership clients. Done carelessly it's a reputational landmine, and we'll say no to the careless version.
Multi-input style and motion references
You can hand it reference material (images, a clip, audio) and have it carry the style, motion or feel across into the output, rather than describing everything in words.
Our take: this is the quietly important one for agencies, because it's how brand consistency survives contact with generative tools. Being able to say "match the look of this" instead of trying to spell out a brand's visual language in a prompt is the difference between on-brand and off-brand output. For clients: we can feed it a brand's existing hero film or palette as the reference and keep new social cuts looking like them, not like the generic Omni house style. That's the single biggest lever for making AI video usable in real brand work.
SynthID watermarking and provenance
Everything Omni produces carries an imperceptible SynthID watermark, verifiable through the Gemini app, Chrome and Search.
Our take: we flagged this in the recap and it matters more here than anywhere. Omni is exactly the kind of tool that makes provenance non-negotiable. We like that it's baked in by default rather than opt-in. The honest tension is that Google is both the biggest generator of this media and the one selling the detector, which is a position worth watching. For clients: we treat disclosure as a feature, not an admission. Signed, credentialed, openly-AI-or-human content is going to read as more trustworthy as the feed fills with the undisclosed kind, which ladders straight into the E-E-A-T argument we keep making.
What it can't do yet
Worth being straight about the limits. Video is the only output for now; audio and image outputs are coming later. Audio input is initially restricted to voice references, with other audio rolling out over time. And standalone speech and audio editing is being held back deliberately while Google works out responsible deployment.
Our take: the held-back audio editing tells you Google knows exactly how dangerous voice cloning is, which is reassuring and also a reminder of where the real risk sits. The "video-only output" limit is fine for most social work but means it's not yet the one-tool-for-everything the framing implies. For clients: we'd brief it as a powerful video tool today, not a finished creative suite. Plan around what it does now, not the roadmap.
So how would we actually use it for clients?
Strip away the demo gloss and Omni earns its place in three concrete ways. First, speed on social-first brands: turning a founder's voice note or a single shoot into a week of on-brand cuts, reshaped live in review. Second, explainer and product content: using the physics and knowledge grounding for "how it works" video that used to need motion design. Third, scaling a personal brand with avatars, under strict consent and human sign-off, so a busy founder shows up consistently without living in front of a camera.
And the honest counterweight, the same one from the recap: putting Omni free into Shorts means everyone has it. Your beautifully made clip now competes in a feed flooded with the same tool's defaults. So the bar for ideas goes up, not down. Our job stops being "can we make a video," anyone can now, and becomes "is this video worth making, on-brand, true, and disclosed." The strategy, the story and the judgement are the scarce things. Omni just does the legwork faster, exactly like the Gemini features we've already wired into our own site.
If you want to talk through where Omni, or any of the I/O announcements, genuinely fits your business rather than just your feed, that's the conversation we enjoy. Tell us about your project and let's get ahead of it together.