Google Gemini Omni Flash 1.0 is the preview release of Google's natively multimodal video generation model. It reasons about the scene before rendering, turning plain-language descriptions into smooth, natural short clips. Describe the scene, subject motion, and camera work, and receive a 16:9 mp4 with a download link. You can also upload one PNG or JPG image as a reference (image-to-video): for example, upload a product photo and describe an orbiting camera move, and the video evolves from that picture; only one reference image is used per run. It supports continuous conversations (multi-turn editing): after a video is generated, simply type a follow-up instruction (for example, change the sky to a sunset, or slow the pacing) and the model keeps editing the previous video without re-uploading; uploading a new image starts a fresh generation from that image. Generation runs in the cloud and usually takes one to three minutes. Great for social clips, product showcases, storyboard previews, and creative animation. For precise results, add film terms such as dolly zoom or slow motion to your description. Generated videos are 720P in 16:9 landscape; the clip length is decided automatically by the model (officially about 3 to 10 seconds). Maximum total upload size per request: 256MB (all files and message content combined). Accepted image formats: jpeg, jpg, png. Maximum size per image: 14MB. Image dimensions: at least 300 pixels per side. Maximum image resolution: 36,000,000 pixels in total. Output content formats: Video Video output: one 720P MP4 video (16:9); the length is decided by the model (about 3 to 10 seconds). This module supports continuous conversations. AI can make mistakes. Please verify important information.