Try the New ImagineArt! 🎉 Smarter, Faster, Better!

Try now
HomeBlogsGemini-omni-overview
Gemini Omni: what Google's new AI video model actually does

Gemini Omni: what Google's new AI video model actually does

Gemini Omni replaces Veo in the Gemini app. Here is what it generates, what it edits, what it costs, and the eight things it still cannot do.

Zahida Misher

Zahida Misher

August 5, 2026 • Updated August 5, 2026

14 mins Read

On this page

I spent about five years of my life on the receiving end of phrases like "can we just try it without the blue bowl."

Not a big ask, on paper. In practice it meant back to the edit bay, back to the colourist, and back into a queue behind three other clients. One prop. Two days. The version we shipped was the fourth cut of a fifteen second spot, and by then nobody remembered what was wrong with the first one. That loop, the one where a small change costs a full round trip, is the single most expensive thing in creative production and almost nobody puts it in the budget.

Gemini Omni is Google's attempt to delete that loop. You generate a video, then you tell it "make the violin invisible," and it makes the violin invisible while leaving everything else exactly where it was. No re-render from scratch, no re-describing the whole scene. Google's own framing is the cleanest summary I have read: think of Gemini Omni as Nano Banana, but for video.

It is also, quietly, the end of Veo inside the Gemini app. That part matters more than the demos, and we will get to it.

What is Gemini Omni?

Gemini Omni is Google DeepMind's multimodal video generation and editing model. It takes text, images, video and audio as input, in any combination, and returns video with sound. The tagline Google uses is "create anything from any input, starting with video," and the "starting with" is doing real work in that sentence. Video is the first output type, not the only one planned.

The model doing the work is called Gemini Omni Flash. In the Gemini API it is gemini-omni-flash-preview, and as of the last docs update on 30 June 2026 it is still labelled preview. If you have seen the name Omni Gemini floating around in forums, it is the same thing with the words in the wrong order.

Three things separate it from the video models that came before it, and they are worth naming precisely rather than admiring in the abstract.

It is natively multimodal, which means it processes text, image, audio and video at the same time rather than converting everything into one format first. It does conversational editing, which means each instruction builds on the last result instead of starting a new generation. And it carries Gemini's world knowledge, so it has an understanding of gravity and fluid dynamics alongside an understanding of history, biology and cultural context.

That third one sounds like marketing until you see what it produces. One of Google's own demo prompts is "claymation explainer of protein folding, everything is made out of clay, no hands, stop motion, accurate." The word doing the heavy lifting there is accurate. A model that only knows how clay looks will give you clay. A model that also knows how proteins fold will give you a clay video that a biochemist would not wince at.

What happened to Veo

Here is the part most coverage has buried, so I will put it near the top.

Gemini Omni replaces Veo in the Gemini app. That is not a rumour or a reading of the tea leaves, it is Google's answer in its own FAQ on the Gemini video generation page. If you have been opening the Gemini app to generate video with Veo 3.1, that flow is now Omni.

Veo itself has not been deleted. It still exists as a specialised model on DeepMind's model list, and Veo 3.1 is still one of the engines inside Google Flow. What changed is which model answers when you ask the Gemini app for a video. Omni is now the default, and it brings video to video editing and multi turn editing that Veo did not have.

If you are trying to work out whether your existing workflow just changed, the test is simple. If you were prompting video inside the Gemini app, it changed. If you were working in Flow, or hitting the Veo API directly, it did not, at least not yet.

What Gemini Omni actually does

Conversational editing, which is the whole point

Every other feature on this list is nice. This one is the reason the model exists.

You generate a video. Then you type "make the violin invisible." The model returns the same scene, same violinist, same field, same camera position, with no violin. Then you type "change the camera angle to be over the violinist's shoulder," and it does that, to the version with no violin. Each turn stacks on the one before it, and the parts you did not mention stay put.

Technically this runs on Google's Interactions API, and each turn references the previous one by ID so you are not re-uploading video between edits. Practically, it means the revision round stops being a production task and becomes a sentence.

I want to be honest about why this lands so hard for me specifically. The blue bowl problem was never a technical problem. It was a queue problem. Every small change had to re-enter a process built for big changes, and the process did not care that the ask was tiny. A model that can hold a scene in memory and change one element of it is not a faster edit bay. It is a different shape of edit bay.

Reference almost anything

Omni takes references and combines them into one output. Images for style, images for subject, video for motion, text for everything else.

The photo to video path accepts up to five photos in the Gemini app. In the API you can tag each input so the model knows its job: <FIRST_FRAME> makes an image the opening frame, and <IMAGE_REF_0> through <IMAGE_REF_N> mark images as references rather than frames. One of Google's example prompts chains six reference images across a ten second fashion sequence, assigning two people and four products to specific time ranges.

Motion transfer works the same way. Give it a video of a whale swimming and an image of a reflective material, and it will move the material the way the whale moved without ever showing you the whale. That is a genuinely strange capability and I mean that as a compliment.

It renders text that means something

Text in AI video has been embarrassing for two years. Omni handles it, and more usefully, it syncs text to what is happening on screen.

Google's demo prompt for this is a full alphabet sequence: twenty six unusual objects, one per letter, each with a hand written lower third naming it, roughly nine frames per item at 24fps, calm music underneath. Every part of that prompt is a constraint that older models would have quietly ignored. The output holds all of them.

For anyone who has ever had to get a legal line or a price point to appear correctly in a generated video, this is the feature that moves it from toy to tool.

Avatars, if you want to be in your own content

You can register a digital version of yourself, once, and then cast it into any generation without uploading a photo each time. Google is clear that it is optional and that only you can use your own avatar. It is also one of the features most likely to be restricted depending on where you live, which brings us to the least fun section of this post and the most useful one.

The benchmark numbers, read honestly

Google published head to head results, and I am going to report them the way I would want them reported to me, including the one where Omni did not win.

On video editing, human raters ran side by side comparisons across 504 examples and Omni led on both overall preference and instruction following. This was an internal benchmark, which is worth noting.

On text to video, raters looked at 1,003 prompts from MovieGenBench, which is a public dataset released by Meta rather than by Google. Omni came out best on overall preference and instruction following. Using a competitor's benchmark and winning on it is a stronger claim than winning on your own.

On fast motion specifically, the evaluation set was 500 detailed prompts covering sports and athletic performance, the category where AI video has historically fallen apart into soup.

On image to video, using 355 image and text pairs from VBench, Omni Flash tied with Grok Imagine Video and Kling, and the three of them led the rest of the field together. A tie, not a win. If you are already happy with Kling for image to video, this benchmark is not the reason to switch.

On reference to video, 468 examples, Omni led on overall preference and on speech adherence.

So the shape of it: strongest on editing and on following instructions, competitive but not dominant on image to video. That is a coherent picture of a model built around control rather than raw spectacle.

What Gemini Omni cannot do yet

Nobody publishes this section, which is exactly why I am publishing it. Google lists these limitations in its own developer docs, and knowing them now saves you a wasted afternoon later.

1 - Video extension is not supported. Neither is interpolation, meaning you cannot give it a first frame and a last frame and ask it to generate the middle. If your storyboard depends on either, Omni is not the tool for that shot.

2- Editing uploaded videos is not available in the European Economic Area, Switzerland or the United Kingdom. Videos the model generated itself can still be edited in those regions, but your own footage cannot be uploaded for editing. Uploading or editing images containing minors is also unsupported in those same three territories, and certain recognisable people are blocked everywhere.

3- You cannot reference more than one video in a prompt. Google's wording is that attempting it may produce degraded or unexpected output, which in my experience is engineering for "it will look wrong and you will not know why." Audio reference uploads are not supported at all in the current API version. Video references up to three seconds are accepted by the schema but are not processed correctly yet, which is the specific kind of half working that costs you a morning.

4- Voice editing is not supported. YouTube videos cannot be used as a source. System instructions, temperature, top_p, stop sequences and negative prompts are all unavailable, so if you want to exclude something you put it in the prompt as plain language, along the lines of "no dialogue" or "do not show the drawing."

5- Aspect ratios are 16:9 and 9:16 only, with landscape as the default. Clips run to ten seconds. English is fully supported and other languages have not been evaluated, so they may work and may not.

6- None of that makes it a weak model. It makes it a preview model with a clear specialism, and planning around a known limitation is a great deal cheaper than discovering it at 11pm the night before a launch.

Getting access and what it costs

Gemini Omni requires a paid Google AI plan. It is available on Google AI Plus, Pro and Ultra, to users 18 and over, in every language and market where the Gemini app is available. There is no free tier for Omni, which is a real change in posture given that Nano Banana 2 shipped free to everyone.

You can reach it in three places. The Gemini app at gemini.google.com is the main one, and it is where Omni has replaced Veo. Google Flow runs it as one of its models, alongside Veo 3.1 and the Nano Banana family, so if you already work in Flow you have it. YouTube Shorts is the third surface, which tells you something about who Google thinks the volume user is.

Developers get it through the Gemini API as gemini-omni-flash-preview via the Interactions API. Two practical notes from the docs that will save you time. Videos over 4MB should be retrieved with delivery="uri" rather than inline base64, and if you set store=false for faster synchronous generation you lose the ability to edit that video in later turns. Speed or editability, pick one per job.

Every video Omni generates or edits inside the Gemini app, Flow or YouTube carries a SynthID watermark and C2PA Content Credentials. The watermark is invisible and machine detectable, and you can verify a file by uploading it to the Gemini app and asking whether Google AI made it.

How to prompt it without burning credits

Two habits from Google's own prompt guidance are worth more than any prompt template you will find, because both of them fix mistakes people make on their first day.

By default Omni gives you multiple shots. It assumes you want a small narrative and it will cut between angles unless told otherwise. If you need one continuous take, say so explicitly with a phrase like "in a single unbroken shot" or "no scene cuts." I lost two generations before I worked this out.

For editing, short prompts beat detailed ones, which is the opposite of the instinct everyone brings from text to video work. Google's example is worth copying exactly. Instead of a long paragraph describing a black cat entering from the right, jumping onto a man's lap and being stroked, the prompt that works is "add a cat that jumps onto his lap, he begins to pet it. Keep everything else the same." That last sentence is the important one. Add "keep everything else the same" to any targeted edit and the model stops helpfully changing things you liked.

You can also time events in plain language, either conversationally with "after 3 seconds, a woman enters the scene," or with a timecode block that reads [0-3s], [3-6s], [6-10s]. And you should describe the audio you want, because the model will invent a soundtrack if you do not. "No dialogue" and "include calm background music" are both valid instructions.

If you want the long version, we have a full Gemini Omni Flash prompts guide with tested prompts, and a step by step guide to Gemini Omni Flash video generation that walks through the conversational editing flow properly.

Where this changes a real workflow

I am wary of use case sections that list every industry on earth, so here are the four where the specific capability actually maps to a specific cost.

Product and packaging variants. This is the one I have lived. You shoot one hero, then generate the flavour variants, the seasonal pack, the regional label. The conversational edit is what makes it viable, because "change the label to the mango version, keep everything else the same" is one sentence rather than one reshoot. Photographing packaging in a grocery aisle is still a good way to spend a Saturday, but it is no longer the only way to get a second angle.

Social cutdowns. Native 9:16 plus a ten second cap is not a limitation for Shorts, Reels and TikTok, it is the exact spec. The YouTube Shorts integration is Google being unsubtle about this.

Explainers with real content in them. The protein folding and hippocampus demos are the tell. If your video has to be correct as well as attractive, world knowledge stops being a spec sheet line and starts being the reason you picked the model.

Motion matching to existing footage. If you already own a clip whose movement works, motion transfer lets you keep the movement and change everything else. That is a strange and useful way to get consistency across a campaign without reshooting the campaign.

Where I would not use it yet: anything needing a clip longer than ten seconds without stitching, anything depending on first and last frame interpolation, and anything where your source footage has to be uploaded from inside the UK or the EEA.

If you cannot get Gemini Omni

Two reasons you might be stuck. Either the paid Google AI plan is not something you want to add, or you are in a region where the editing features you need are switched off.

ImagineArt runs Google's models alongside about forty others, so you can generate with Veo 3.1 and Veo 4, edit images with Nano Banana 2, and switch to Kling 3.0 or Seedance 2.0 for the shots those models handle better, without holding three separate subscriptions. Given that Omni tied rather than led on image to video, having Kling in the same tab is more useful than it sounds.

Frequently asked questions

What is Gemini Omni? Gemini Omni is Google DeepMind's multimodal video generation and editing model. It accepts text, images, video and audio as input and returns video with audio, and it lets you refine that video through conversation rather than by re-generating it. The underlying model is called Gemini Omni Flash.

Is Gemini Omni free? No. It requires a paid Google AI subscription on the Plus, Pro or Ultra tier, and you must be 18 or over. This is a departure from Nano Banana 2, which Google released free across its products.

How do I get Gemini Omni? Subscribe to Google AI Plus, Pro or Ultra, then open the Gemini app. It is also available inside Google Flow and YouTube Shorts, and through the Gemini API as gemini-omni-flash-preview. Availability of specific features like avatars and video to video editing varies by country.

What happened to Veo? Gemini Omni has replaced Veo as the video model inside the Gemini app. Veo has not been retired more broadly. It remains a specialised DeepMind model and Veo 3.1 is still available inside Google Flow and through the API.

How long can Gemini Omni videos be? Ten seconds, in either 16:9 or 9:16. Video extension is not supported, so longer sequences have to be built by stitching separate generations.

Is Gemini Omni better than Veo 3.1 or Kling 3.0? On editing and instruction following, Google's human rater tests put it ahead of the field. On image to video it tied with Kling and Grok Imagine Video on the VBench benchmark rather than beating them. Pick Omni when the job is iterative editing and control. The gap narrows considerably when the job is a single image to video render.

What's the difference between Google Flow and Gemini Omni? Google Flow is the actual product, Google's AI filmmaking studio where you generate, edit, and stitch together video, image, and audio content, built on Veo, Imagen, and Gemini working together. Gemini Omni (and its faster Omni Flash version) is the underlying model, not a separate app: it's what powers Flow's newer conversational editing features, like refining a clip through natural language, keeping characters consistent across scenes, and blending real footage with generated content. In short, Flow is the studio you work in; Gemini Omni is the model running under the hood that makes the newer editing magic possible.

Zahida Misher

Zahida Misher

Exploring how AI transforms marketing, I help businesses unlock growth and future-proof their presence.

Endless Possibilities. Just Imagine.

Product

  • Audio Studio
  • AI Film Studio
  • AI Ad Studio
  • Lipsync Studio
  • AI Workflows
  • Features
  • Enterprise
  • Apps
  • API Docs

Image

  • AI Image Generator
  • ImagineArt 1.5
  • ImagineArt 2.0
  • GPT Image 2
  • Nano Banana 2
  • Image Upscaler
  • Flux 2

Video

  • AI Video Generator
  • AI Video Editor
  • Seedance 2.0
  • Sora 2
  • Veo 3.1
  • Kling 3.0
  • Pixverse v6
  • Seedance 2.5

Resources

  • Blogs
  • Enterprise Resources
  • Community
  • Pricing
  • Creator Program
  • Contact Sales

ImagineArt

  • Privacy Policy
  • Terms & Conditions
  • Help Center
  • About Us

All rights reserved.

Blog
Editing Tools

AI Video Editor

Create and edit videos with AI transitions and effects.

AI Image Editor

Edit, retouch, and transform images with AI tools.

Kling AI Motion Control

Add dynamic motion to static images with AI-powered animation controls.

AI Image Generator
BG Remover
AI Anime Generator
AI Image Combiner
AI Image Face Swap
AI Image Replace
AI Video Generator
AI Heygen Avatar
AI Animation Generator
AI Product Video Maker
AI Video Object Removal
AI Video Recolor
AI Video background Changer
AI Models
Seedance 2.0
Kling 3.0
Seedream 5.0
Recraft V4
Runway Gen 4.5
Seedance 2.5
Explore All
ConnectUnlock the future of creativity with our Generative AI community—where art, video, and images are born from the power of AI imagination!
Discord
Facebook
Instagram
Pinterest
Reddit
Snapchat
Twitter
YouTube
WhatsApp
AffiliateAPICreatorsPricing
Launch App