Omni Reference AI Video Generator

Combine image, video, and audio references in one generation. Choose ByteDance's Seedance 2.5 or MiniMax H3, give every reference a role, and generate consistent video.

Two Models, One Platform

Seedance 2.5Long Takes, Deep Reference Control

ByteDance's model. 4–30s at 480P or 720P, with up to 30 images, 10 video clips, and 10 audio files — 50 references in one generation.

MiniMax H3Higher Resolution, Focused Shots

MiniMax's model. 4–15s at 768P up to 2K, with up to 9 images, 3 video clips, and 3 audio files. Suited for short storyboarded scenes.

One Reference Library, Both Models

Switch models without reuploading. Compare both against the same reference pack and keep the take that fits the shot.

Omni Reference in 3 Simple Steps

  1. STEP 1

    Upload Your References

    Drag, paste, or select images, video clips, and audio from your device or the Assets panel, where anything you made in AI Image Editing already sits.
  2. STEP 2

    Assign Reference Roles

    Name them in the prompt: Image 1 the character, Image 2 the scene, Video 1 the motion, Audio 1 the dialogue. One job per file.
  3. STEP 3

    Configure and Generate

    Pick your model, then set duration, aspect ratio, and resolution. Check the credit cost on the Generate button before you submit — it varies.

What Happens When Every Reference Has a Job

Produce AI Drama with a Consistent Cast

Character consistency across shots and across episodes is the most expensive problem in AI drama production. Omni Reference locks cast identity to image references that carry from scene to scene.

One Cast, Every Shot

Upload character references and assign each an identity role. The same face, hair, and outfit carry through every shot.

Continuity Across Episodes

Reuse the same reference pack across scenes and episodes. Cast, sets, and props stay consistent because the references carry the constraints.

Assign Dialogue by Speaker

Attach audio references carrying your dialogue, then assign each line to a named character in the prompt itself.

Matching Scene Transitions

Set first and last frames to control where a scene opens and lands. Both endpoints stay reference-driven.

Produce AI Drama with a Consistent Cast

Music Video Production with Audio References

An AI music video lives or dies on whether the performer looks the same in every cut. Upload your track as an audio reference and performer images as identity references, then generate.

Audio-Driven Generation

With Seedance 2.5, an audio reference guides soundtrack and dialogue timing. Pair it with image references.

Plan Your Beats First

Structure the shot as timed beats in the prompt: one action and one camera intent per beat.

Performer Stays On-Model

Lock the performer with image references. Change set, lighting, and angle while the face stays fixed.

Cross-Cut Continuity

Generate a wide, a close-up, and a profile from one reference pack — same performer, wardrobe, location.

Music Video Production with Audio References

Built for All Creators

Anime & Webtoon Creators

Anime & Webtoon Creators

Lock a character's face, hair, outfit, and proportions across a whole series. Upload character sheets as image references and generate anime character video that keeps the same cast episode after episode.

Short-Form Filmmakers

Short-Form Filmmakers

Cast, location, and wardrobe stay fixed across cuts. Upload identity, set, and costume references once, then generate any shot in the sequence from the same pack.

Brand & Ad Teams

Brand & Ad Teams

Lock the product, spokesperson, and setting in every shot. Image references control silhouette, color, and packaging. Add audio for dialogue or soundtrack input.

Music Video Producers

Music Video Producers

Upload a track as an audio reference and performer images as identity references. Generate scenes where the performer stays on-model across every cut.

Frequently Asked Questions

What is Omni Reference on DomoAI?

Omni Reference is DomoAI's multi-reference video generation feature. Upload images, video clips, and audio files as references, assign each one a role in your prompt, and generate video guided by all of them at once.

What models are available in Omni Reference?

Two. Seedance 2.5, ByteDance's model, generates 4 to 30 seconds at 480P or 720P and accepts up to 30 image, 10 video, and 10 audio references. MiniMax H3, MiniMax's model, generates 4 to 15 seconds at 768P up to 2K and accepts up to 9 image, 3 video, and 3 audio references.

How many references can I upload to Omni Reference?

It depends on the model. Seedance 2.5 accepts up to 30 images, 10 video clips, and 10 audio files — 50 references in total. MiniMax H3 accepts up to 9 images, 3 video clips, and 3 audio files, with a combined limit of 12 references. Each video and audio reference has per-clip and combined duration limits that vary by model.

What resolution does Omni Reference support?

Seedance 2.5 outputs at 480P or 720P. MiniMax H3 outputs at 768P up to 2K. Neither model supports 4K on this surface. For higher resolution, run the finished clip through DomoAI's Video Upscaler after generating.

Do I need to assign each reference a specific role?

You do not have to, but you should. Naming references in your prompt — Image 1 as the character, Image 2 as the scene, Audio 1 as the dialogue — tells the model what each file controls. Without role assignments, the model decides on its own, and the result is less predictable.

How is Omni Reference different from Image to Video?

Image to Video takes one image and adds motion to it. Omni Reference takes multiple images, video clips, and audio files together, each with an assigned role, and generates video guided by all of them. Use Image to Video when one strong still is the starting point. Use Omni Reference when you need character consistency, scene references, motion references, or audio combined in one generation.

Can I generate from audio references alone?

With Seedance 2.5, yes. A generation can run from audio references with no image or video reference attached. Most workflows still pair audio with at least one image reference so the model has an identity or scene to hold onto.

Generate, stylize, and upscale in one place

Create stunning videos from text, images, or footage. Generate, style, and upscale—all in one platform.