Video
Add Node → Generate → Video
This node makes video — from a text prompt on its own, from one or two still images, from a set of reference images, or from a clip you feed it.
It arrives already wired to an empty video slot, ready to fill once you generate.
What’s on the node
| Area | What goes in it |
|---|---|
| Prompt box | Where you describe the shot. Type into it directly, or wire something into the prompt handle and it turns read-only, showing the upstream text live. |
| + Presets | Under the prompt: reusable scraps of text — a lens, a film stock, a background — applied with a tick. See Presets. |
| Model and settings | The model picker, and under it whatever that model exposes — resolution, clip length, audio. The controls change with the model, and switching resets them. |
| P | One only. Takes a Prompt node, a Text node, a Prompt Split, or a finished result’s own prompt handle. |
| I | Reference images, numbered in the order they’re sent to the model. Each takes an image, a Media Group or a Media Placeholder. |
| FL | First and last frame, on the models that support it: a starting image and an ending one, with the motion generated between them. |
| V | A finished clip, on the models that restyle or continue existing footage. |
| A | A reference audio track, on Seedance 2.5 — the only model here that takes one. |
Which of these you actually see depends on the model — the slots change shape when you switch. A ring is hollow until something is connected and fills with colour once it is, so you can see at a glance what the node still needs. There’s more on all of this in How nodes connect.
Text-to-video, or image-to-video?
Type a prompt and generate, and the model invents the whole shot from words alone.
Connect an image instead, and it animates that instead of guessing: composition, subject and framing are already decided, and the model only has to supply the motion. This is almost always the better way in — with nothing to anchor it, a model has very little to go on, and you can burn through a lot of generations before landing near what you had in mind. An image made with the Image node works fine as a starting point.
Pick the model before you wire anything up. The node’s input slots change shape depending on which one you choose.
Which model should you use?
| Type | Model | Resolution | Duration | Price range | What it’s good at | |
|---|---|---|---|---|---|---|
| Pro | text-to-video image-to-video video-to-video |
Seedance 2.5 | 480p · 720p · 1080p · 4K | 4–30s | $0.72–54 | The most capable card here, and the longest clips: up to nine reference images, a reference clip and a reference audio track in any combination. Reach for it first. |
| Pro | text-to-video image-to-video video-to-video |
Seedance 2.0 | 480p · 720p · 1080p | 5 · 10 · 15s | $0.50–9.00 | The previous generation, still excellent and cheaper. A Fast tier knocks about a sixth off the price. |
| Pro | text-to-video image-to-video video-to-video |
Kling O3 | 720p · 1080p · 4K | 4–15s | $0.34–6.30 | The one for reference images, and for restyling a clip you already have. No first and last frame — that’s V3’s job. |
| Pro | text-to-video image-to-video |
Kling V3 | 720p · 1080p · 4K | 4–15s | $0.34–6.30 | The first-and-last-frame one: give it a start and an end, get the motion between. Every bit as good as O3, just for a different job. |
| Good | text-to-video image-to-video video-to-video |
FLUX 3 | 720p · 1080p | 5–20s | $0.85–5.80 | Black Forest Labs’ video model, with native audio. The only one here that continues a connected clip rather than restyling it. |
| Good | text-to-video image-to-video |
Veo 3.1 | 720p · 1080p · 4K | 4 · 6 · 8s | $0.40–1.60 | A solid general-purpose pick, and the one to reach for when the clip needs audio. Fixed lengths, and Fast and Lite tiers if the full price is more than the shot deserves. |
| Good | text-to-video image-to-video video-to-video |
Gemini Omni | 720p | 3–10s | $0.39–1.30 | Works from up to seven reference images, or edits a connected clip from a prompt. No first and last frame, and no resolution to choose. |
| Creative | text-to-video image-to-video |
LTX 2.3 | 480p · 720p · 1080p | 5 · 10s | $0.10–0.40 | An open model: cheap by a wide margin, good results, and synchronised audio. The one to test an idea on before spending. |
Prices are indicative and set by the providers, who revise them: see wavespeed.ai or runware.ai for current rates.
Prices are per clip, charged straight to whichever provider you’ve connected. The low end is the shortest clip at the smallest size; the high end is the longest clip at the largest size. Both settings move the figure a lot — on Seedance 2.5 the same model goes from well under a dollar at 480p for four seconds to around $54 at 4K for thirty — so a cheap model run long and large can cost more than an expensive one kept short. Where a model has tiers, the low end is its cheapest tier and the high end its dearest. Rates are WaveSpeed’s. The card quotes the real price before you press Generate: check it there.
Tiers
Some models put a second row under the picker: Type or Resolution, a row of buttons that pick a tier within the same family.
Seedance 2.0 offers Normal and Fast; Veo 3.1 offers Normal, Fast and — on Runware — Lite. Those trade quality for price. Kling O3 and Kling V3 use theirs for the output size instead: 720p, 1080p or 4K, which on those two families is what the tier decides rather than a resolution setting.
Switching tier keeps every compatible setting, so it’s a cheap thing to try: run a test on Fast, then flip to Normal for the take you keep.
Which models and which tiers you see depends on your provider — the two catalogues overlap but aren’t identical, and Veo 3.1 Lite is Runware-only.
First and last frame, or reference images?
These are two different ways of feeding a model stills, and they don’t do the same thing.
With first and last frame you give a starting image and an ending image, and the model generates the motion that connects the two. Use it when you need a precise, controlled transition between two exact shots. Seedance 2.5, Seedance 2.0, Kling V3, Veo 3.1 and FLUX 3 support it; Kling O3, Gemini Omni and LTX 2.3 don’t — O3 works from reference images instead.
With reference images you drop in one or more stills as inspiration — a subject, a location, a style — and the model draws on them without locking any exact frame. More creative freedom, less control over where the shot ends up.
On most models the two are exclusive: a first and last frame, or references, not both. Wire up a mix and the node says so before it spends anything, rather than quietly dropping half of it.
Careful — this one depends on your provider. Seedance 2.0 changes shape around its video input: with no clip connected you get a first and last frame, and connecting one swaps those for up to nine reference images. Which combinations a model will accept differs between WaveSpeed and Runware for several of these, so a graph that runs on one key may want rewiring on the other.
Some models take a whole video as input instead. Most of them restyle it — the original motion stays, the look changes — while FLUX 3 does something different and continues the clip from where it ends.
To connect anything, drag it from the Media Library onto the canvas, then link it to the matching input on the Video node.
Switching to a different model resets your settings, so nothing carries over by mistake.