Google released Gemini Omni 1.1 Flash on August 27 as a video-generation update aimed at production use. The emphasis is control, not just text-to-video: developers can extend an existing scene, specify its opening and closing frames, preview at low resolution, and then render at 1080p or 4K.
Scene extension expands context from one second to ten
Omni 1.1 scene extension can analyze up to ten seconds of prior video context, then generate in ten-second increments up to a cumulative 40 seconds. More context can help preserve characters, settings and narrative direction, but it remains generative and does not guarantee perfect continuity.
Google also exposes an interaction flow through the Gemini API. Developers can pass the previous video interaction into the next request and describe how the scene should continue, making it easier to embed continuation into editors, storyboarding tools and custom production workflows.
First-and-last-frame control makes shots more designable
The new version lets a creator specify a shot’s first and last frame while the model generates motion between the keyframes. For orbiting cameras, pushes, pulls, transitions and seamless loops, that is closer to traditional storyboard thinking than a description alone.
Developers can also provide up to three seconds of video reference for extra visual context around a character or action. A product could turn this into a drop-reference, set-target-frames and generate workflow, but users still need to review identity, motion and continuity.
Pricing: available on the paid API tier
This is not a free API model. Google’s current pricing page lists gemini-omni-1.1-flash on the paid Gemini API tier: $1.50 per 1 million input tokens for text, image, video or audio; $9 per 1 million output tokens for text; and $17.50 per 1 million output tokens for video. The model is not available on the free tier.
Video output is billed by tokens. Google’s standard conversion uses 5,792 output tokens per second of 720p video, which works out to approximately $0.10 per second. Based on the documented 3–10 second output range, that is roughly $0.30 for 3 seconds or $1.00 for 10 seconds. These are estimates; the final bill depends on output-token usage and the account’s pricing tier. Google’s lower 360p cost is a relative statement, not a separate fixed list price.
Prototype at 360p before rendering in 4K
Google says 360p previews can generate up to 60% faster than 720p at about one-third the cost, making them useful for storyboards and prompt iteration. After the direction is confirmed, teams can switch to 1080p or 4K instead of spending heavily on unproven versions.
Omni 1.1 is rolling out through Google AI Studio and the Gemini Enterprise Agent Platform API, and Google AI Plus, Pro and Ultra subscribers can use it in Flow. Availability, quotas and pricing vary by product and account, so an announcement is not the same as unlimited access everywhere.
What this means for video-tool developers
Gemini Omni 1.1 Flash moves video generation from one-off output toward an orchestrated production workflow. Extension, frame control, video references, low-cost previews and 4K output let developers build products closer to editing and storyboarding, while rights, model limits, continuity and cost still require attention before delivery.
For official details, see Google’s Gemini Omni 1.1 Flash announcement 與 Gemini API documentation。
