Skip to main content

Choosing a video model

What the main video models are good at, which inputs each one takes, why some inputs only work on the visual canvas, and how to pick without wasting credits.

Written by Sebastian Mourra

Visuals runs many models and adds new ones often. You pick one per card, and the card shows what it will cost before it runs. This article is about how to choose.

For the current list, check the release notes. What follows is the shape of the choice, not a fixed catalog.

Start with the default

Gemini Omni 1.1 Flash is the default video model on a new video card. It takes a prompt, up to 3 reference clips to carry a character or a motion style, and a first and last frame to pin where a clip starts and ends. It has 4 resolution tiers: 360p for drafts, 720p, 1080p, and 4K.

If you do not have a reason to pick something else, start here.

The Kling family

5 models, each for a different job. All are billed per second of video produced, so a longer clip costs more.

  • Kling V3 Video. Text to video or image to video, with sound. Up to 15 seconds.

  • Kling V3 Omni. The all-rounder: text or image to video, reference guided generation, and editing or restyling an existing video.

  • Kling V3 Motion Control. Takes the motion out of one video and applies it to a character image. Length follows the reference video.

  • Kling O1 Edit. Edit or restyle an existing video by describing the change. Keeps the original motion and timing.

  • Kling Avatar V2. Turns a portrait into a talking avatar, lip synced to an audio track. Length matches your audio.

On Kling V3, multi-shot playback plays your shots back to back, and the shot lengths must add up to the card duration. Visuals shows the shot total and offers a one click fix if they do not match, before any credits are spent.

Picking by what you have

  • A prompt only. Any text to video model. Start with the default.

  • A still you want to move. Image to video. Feed it as a start frame.

  • A character you need to keep consistent. Reference images, or a model that takes a character input.

  • Motion you want to copy. Kling V3 Motion Control with a reference video.

  • A face and an audio track. Kling Avatar V2.

  • An existing clip you want changed. Kling O1 Edit or Kling V3 Omni.

Where inputs work

This trips people up, so it is worth stating plainly:

  • Prompts and image references work on the image and video generator pages and on the visual canvas.

  • Video references and audio inputs work on the visual canvas only.

So a motion reference video or an avatar's audio track has to be supplied on the canvas.

The app stops you before you pay

  • If a model requires an input you have not given it, the run is stopped before credits are spent and the message names what is missing.

  • The duration slider hides on models that take their length from the input, instead of showing a control that does nothing.

  • Switching to a model that cannot make your chosen aspect ratio resets the ratio immediately, rather than failing at run time.

  • If a model refuses a request, the error is in plain English.

A note on prompt length

Some models reject prompts over a certain length. Video generation forms show how long your prompt is as you write, so you can trim before running rather than after.

Reliability

Some models run across more than one provider. If one provider has a problem, the job retries elsewhere instead of failing your card, and a reference photo blocked in one place is retried where it is accepted.

Did this answer your question?