Gemini Omni 1.1 Flash - Google's Talk-to-Edit AI Video Model

Gemini Omni 1.1 Flash is Google DeepMind's any-to-any AI video model for creators who want to generate and edit short clips by talking to them in plain language instead of fighting a timeline. Released as a GA build on August 27, 2026 (model id gemini-omni-1.1-flash), it extends the Gemini Omni family shown at Google I/O 2026 and now tops the Arena text-to-video leaderboard.

Core Features

  • Conversational editing: describe a change in text or voice and the model rewrites the same shot while preserving everything you did not mention.
  • Scene extension: reads up to 10 seconds of prior context and continues footage in 10-second steps to a cumulative 40 seconds.
  • Multi-input any-to-any: takes text, image, and video in one request, with native synced audio generated in the same pass.
  • Resolution ladder: 720p standard, a 360p draft mode at roughly a third of the cost, and upscaling to 1080p or 4K.
  • First and last frame control plus a 3-second reference-video upload for consistent motion and camera work.

Use Cases / Best For

  • YouTube Shorts creators iterating on hooks and transitions quickly.
  • Marketers cutting social variants from a single product shot.
  • Educators and small studios making explainer clips without editing software.

Pricing

Through the Gemini API, Omni 1.1 Flash costs $0.10 per second at 720p (about $1 for a 10-second clip), $0.15 at 1080p, and $0.30 at 4K; a 360p draft mode runs $0.03 per second. It is free inside YouTube Shorts for users 18 and over and is included in Google AI Plus, Pro, and Ultra subscriptions. The model also reached Adobe Firefly and Figma as a connected engine.

Our Take

Best for fast, cheap social video where speed matters more than cinematic polish; the trade-off is a 40-second ceiling and a non-optional SynthID watermark on every clip. Pair it with our AI Video tools roundup for higher-end alternatives.

FacebookXWhatsAppEmail