Upload a reference clip, style frames, or audio — then prompt the change you want. Follow motion, lock a look, or extend an edit without inventing everything from text alone.
Start from a source you trust
Choose Reference (video to video) when an existing clip, style set, or audio bed should steer the output. Text invents without uploads. Image animates a still you already locked. Reference follows sources you attach — video, images, or audio depending on the model. The patterns below are the jobs teams run once they have something to follow.
You have a 2–15s take with camera path or body action you like. Prompt the new subject or setting while asking the model to keep timing and movement energy. Keep the reference short and clean — busy multi-scene masters confuse guidance more than they help.
Some models accept image packs as look references (palette, costume, rendering). Upload 1–N frames that already show the brand look, then prompt the action. Prefer consistent lighting across refs; mixed art styles in one batch usually muddy the result.
When a model supports audio references, attach a short WAV/MP3 bed so lip motion, cuts, or energy can track the track. Keep duration within the model’s limits. Do not expect perfect lip-sync marketing claims — treat audio as guidance, then cut in an NLE if needed.
Certain models treat the first video as the edit or extend target and extra clips as supporting refs. Use that when you need a continuation, a restyle pass, or a controlled rewrite of an existing MP4 — not when you still lack any source footage.
Video to video fails when the prompt contradicts the upload or when you stack too many conflicting refs. Treat sources as the motion/style law; use text for the delta. Recipes below are written for ToVideo operators.
Differentiation is workflow, not another slogan: Reference mode with model-aware slots, visible credits, and the same account as Text and Image.
Open the video generator and choose Reference. Image, video, and audio slots appear only when the selected model supports them — including limits on count, format, and length.
Cost updates with model, duration, quality, audio, and whether video references are attached. Failed finals refund so you can refine guidance without silent balance burn.
Text invents with no upload. Image animates stills. Reference follows clips, style frames, or audio. Switch modes in one form; unavailable options stay hidden.
One login and credit pool across AI Image and Video. Draft in Text, lock a still in Image, or guide with Reference — without exporting between unrelated tools.
Close the tab; the run continues. Download the MP4 from the preview panel or AI Tasks when status succeeds.
Successful generations expose download in the result area. Archive outside ToVideo for ads, clients, or your editor.
Operator answers for reference-guided video to video.
Attach one clear source, prompt the delta, check credits, and generate.