Google Veo 3.1
Use supported reference images to guide a new scene. The form limits the visible input types and aspect ratios to the concrete Veo version rather than applying options from another model family.
No videos generated yet. Upload an image to start!
No videos generated yet. Upload an image to start!
This page lists only model versions configured for reference mode. Each one has its own accepted media counts, aspect ratios, resolutions and durations, so the visible form remains the source of truth for a task.
Use supported reference images to guide a new scene. The form limits the visible input types and aspect ratios to the concrete Veo version rather than applying options from another model family.
MiniMax Hailuo H3 supports reference-mode combinations of images, video and audio under the configured per-type and total limits. At least one image or video is required when audio is attached.
Seedance 2 versions expose multimodal reference inputs for short video generation. Review the exact reference controls, resolution, duration and credit estimate for the selected version before submission.
The page filters out versions that only support text or frame input. You can compare the compatible choices without repeatedly switching away from a model that cannot accept the media combination you prepared.
Decide what each reference should control, choose a compatible model and use the prompt to describe the new motion rather than asking the files to explain themselves.
Open the selector and compare the compatible Veo, MiniMax and Seedance versions. The chosen version determines which reference types and output settings remain available in the form.
Upload images for appearance or composition, video for motion context and audio for supported timing or sound context. Audio cannot be submitted alone; include at least one image or video as the visual foundation.
Write a prompt that names the desired action, camera behavior, environment and relationship to the references. Explain which traits matter instead of assuming every detail in every file should be copied.
Check the media counts, aspect ratio, resolution, duration and calculated credit amount. Submit the task and follow it in Recent Videos without losing older generations from other input modes.
Reference mode stays active from initialization through submission, and the model menu contains only versions with a registered reference contract.
The selected model decides which media rows are enabled and how many files each row accepts. Unsupported types cannot be added through an invisible field.
More files do not automatically create more control. Assign a purpose to each reference, remove contradictions and tell the prompt what the new shot should do.
Before uploading, write down the constraint that matters: a character silhouette, product shape, color palette, camera rhythm, movement pattern or audio cue. Choose the smallest set of files that makes that constraint visible. Several near-duplicate images rarely add as much information as one clean front view and one useful alternate angle. A reference generator interprets the material; it does not guarantee exact identity, logos, typography or object geometry across every frame. Review those details in the completed clip and make sure you have permission to use every source file.
Check lighting, perspective, subject scale and style before combining files. If one image shows a studio product and another shows a different version outdoors, the model must guess which attributes to retain. Video references can communicate motion or camera timing, but they can also introduce a competing subject and background. Audio can add useful context only alongside an image or video in the current workflow. Remove a file whose purpose you cannot explain in one sentence. The configured per-type and total limits are upper bounds, not a recommendation to fill every slot.
Name the subject, action, setting and camera behavior, then state how the references should influence the result. For example, ask to retain the jacket and facial proportions from the images while using the sideways tracking pace from a reference clip. Avoid vague requests to copy everything. If references disagree, explicitly prioritize the important one. Keep the scene short enough for the chosen duration and use chronological language for a transformation. The prompt should explain what changes after the reference material, while the files provide concrete visual, motion or audio evidence.
After a task completes, decide whether the problem came from subject consistency, movement, camera, background, timing or the selected model. Replace or remove one reference, clarify one sentence or adjust one output setting, then regenerate and compare the new task with the original in Recent Videos. A different model may interpret the same material differently because its reference contract and available settings differ. Check the calculated credits before each attempt, especially when video reference surcharges apply. Keep the useful generations so later decisions are based on a visible sequence of experiments.
Choose a compatible model, add only the image, video or audio material that supplies a clear constraint, and review the task before generation.
Reference-capable models only
Per-type and total media limits
Full recent-video history
These answers cover the current multimodal reference workflow and the boundaries visible in the form.
It creates a new short video from a prompt plus one or more supported reference files. Images can provide appearance or composition, video can provide motion context and audio can supplement a visual foundation. The model interprets those inputs rather than copying every detail exactly.
Only versions whose 2.VIDEO configuration includes reference mode are shown. The current menu includes compatible Google Veo, MiniMax Hailuo and ByteDance Seedance versions. Use the live selector as the authoritative list because configured availability can change.
No. The current reference contract requires at least one image or video when audio is present. An audio-only draft is rejected before submission. If a selected model does not expose audio, that upload row remains unavailable.
Limits depend on the model. The form applies separate image, video and audio limits plus a combined total and disables an upload type when its limit is reached. Changing the model can reduce those limits, so review the remaining files before generating.
References can guide recognizable traits, but this page does not promise exact identity or frame-perfect product geometry. Use clean, compatible views, identify the important traits in the prompt and inspect faces, text, logos and small details in the result.
Use the files for concrete evidence and the prompt for direction. Describe the new action, setting and camera, then state which appearance, motion or audio traits should influence the shot. Remove references that conflict or do not have a clear role.
The page calculates credits from the selected model, resolution and duration and includes the configured video-reference surcharge where applicable. The exact amount is displayed before submission. Check it again after changing the model or settings.
All recent video tasks are visible. A compatible reference task can be edited here. Editing a text-only or frame task opens the full AI Video Generator so its saved input mode and media are not silently rewritten. Regeneration creates a separate task.