Google Gemini Omni Flash 1.1 is the generally available release of Google's natively multimodal video generation model. It reasons about the scene before rendering and generates native audio such as music, ambience, or dialogue along with the picture, turning plain-language descriptions into smooth, natural short clips. Describe the scene, subject motion, camera work, and sound, and receive an mp4 with a download link. The control panel above the input box lets you pick the resolution (720P by default, or 1080P and 4K, both upscaled from 720P) and the aspect ratio (16:9 landscape or 9:16 portrait). You can upload one PNG or JPG image as a reference (image-to-video), for example a product photo with an orbiting camera move described in the prompt. With two images, describe a transition from the first picture to the second (first and last frame interpolation), or explain which subject or character each image shows (subject references); the model decides how to use them from your prompt, and at most two reference images are used per run. You can also upload one MP4 video of up to 10 seconds as the source and describe what to change (video editing) or ask the model to continue the scene (video extension); only one video is used per run, and extension is appended to the end of the clip. It supports continuous conversations (multi-turn editing and extension): after a video is generated, type a follow-up instruction (for example, change the sky to a sunset, slow the pacing, or continue the scene) and the model keeps working on the previous video without re-uploading; uploading a new image or video starts a fresh generation from that material. Generation runs in the cloud and usually takes one to three minutes; higher resolutions take longer. Great for social clips, product showcases, storyboard previews, and creative animation. For precise results, add film terms such as dolly zoom or slow motion, and describe the lighting, pacing, and background music. Generated videos are MP4 with audio; the resolution is 720P by default or 1080P / 4K (upscaled) from the panel, the aspect ratio is 16:9 landscape or 9:16 portrait, and the clip length is decided automatically by the model (officially about 3 to 10 seconds). Maximum total upload size per request: 256MB (all files and message content combined). Accepted image formats: jpeg, jpg, png. Maximum number of images: 2. Maximum size per image: 14MB. Maximum total image size per message: 14MB (text and history share the same request size limit; uploads close to the limit may still be rejected). Image dimensions: at least 300 pixels per side. Maximum image resolution: 36,000,000 pixels in total. Accepted video formats: mp4. Maximum number of videos: 1. Maximum size per video: 14MB. Video duration: 1 to 10 seconds (total at most 10 seconds). Video dimensions: 64 to 4,096 pixels per side. Video resolution: 4,096 to 9,437,184 pixels in total. Video aspect ratio: 0.25 to 4.0. Video frame rate: 1 to 120 fps. Output content formats: Video Video output: one MP4 video with audio; the resolution (720P by default, or 1080P / 4K upscaled) and the aspect ratio (16:9 or 9:16) follow the control panel selection; the length is decided by the model (about 3 to 10 seconds). This module supports continuous conversations. AI can make mistakes. Please verify important information.