Home › Guides › Reference-to-video

Using Seedance reference-to-video over API

Updated 2026-10-02

Image-to-video animates one frame you supply. Reference-to-video is different: you hand the model a pile of reference material (a character sheet, a motion clip, an audio bed) and ask for a new clip guided by it. The references are never edited or returned. This matters for anyone who needs the same character, product or voice across many shots.

The three reference fields

On POST /videos, references go in three optional arrays that share one shape: {"type": "<kind>_url", "<kind>_url": {"url": ...}}.

FieldEntry typeDocumented cap
input_referencesimage_urlUp to 9 images (model-dependent)
input_video_referencesvideo_urlUp to 3 on supporting models
input_audio_referencesaudio_urlUp to 3 on supporting models

The Fal-hosted Seedance 2.0 reference endpoint (and the Fast and Mini variants built on the same schema) caps the combined total at 12 references. At least one reference across the three arrays is required; a reference-mode call with none is rejected with a 400. Exceeding any per-kind cap or the combined cap is also a 400, not a silent truncation.

A request that uses all three

The mode is chosen by which fields you send; there is no separate "-reference" model id to look up. Pinning a host with the /fal suffix keeps the call on a host with documented reference support:

curl https://videorouter.sh/api/v1/videos \
  -H "Authorization: Bearer llmr_sk_live_..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "bytedance/seedance-2.0/fal",
    "prompt": "the cat from the reference image walks across the scene, camera follows",
    "input_references": [
      {"type": "image_url", "image_url": {"url": "https://example.com/cat.jpg"}}
    ],
    "input_video_references": [
      {"type": "video_url", "video_url": {"url": "https://example.com/motion-ref.mp4"}}
    ],
    "input_audio_references": [
      {"type": "audio_url", "audio_url": {"url": "https://example.com/ambience.mp3"}}
    ],
    "duration_secs": 5
  }'

The response is an async job. Poll GET /videos/{id} until completed or failed; polling is free, and the job is billed once at creation by the requested duration.

A consistent-character workflow

  1. Build the reference set once. Two to four images of the same character from different angles (front, three-quarter, full body) usually give the model more to hold onto than nine near-duplicates. Host them at stable public URLs, or send data: URIs.
  2. Refer to references in the prompt by role. "The woman from the reference images" beats a name the model has never seen. If you attach a video reference, say what it is for: motion, camera move or framing.
  3. Keep the reference set identical across shots and vary only the prompt and duration. That is the single biggest lever for continuity across a sequence.
  4. Draft cheap, finish expensive. The same reference arrays work on the Mini and Fast variants, so you can test shot compositions there and re-run the approved ones on a full model. See the variant guide.

Pitfalls worth knowing before you ship

Reference vs start image: which to use

If you have exactly one frame and want it to be the literal opening of the clip, use start_image_url (covered in the image-to-video guide). If you want the character or style to appear in a scene the model composes itself, or you need motion and audio guidance, use references. Many pipelines use both in different stages: references to establish the look, start-frame animation for shots that need an exact composition.

Writing prompts that use the references

Reference calls fail quietly more often than they fail loudly: the job completes, but the character drifted because the prompt never told the model which reference mattered. A few habits help. Name each role once and keep the wording stable across shots, for example "the courier from the reference images" every time. If the video reference supplies a camera move, say "follow the camera movement of the reference clip" rather than re-describing it. If the audio reference is ambience, say that, so the model does not treat it as dialogue to lip-sync. Keep the prompt short and concrete; ten sentences of scene description compete with the references instead of supporting them.

When a result is wrong, change one thing per run: the prompt, the order of reference images, or the number of images. Changing all three at once leaves you unable to tell what helped. Log the full request body next to each job id so you can reproduce a shot that worked.

Comparing hosts for reference calls

Seedance is sold by many providers at very different per-second rates, and reference-capable hosts are a subset. The live table below covers every Seedance row in the catalog; confirm reference support on the specific model page before pinning.

ModelCheapest hostPriciest hostCheapest isHosts
bytedance/seedance-2.5 (480p)OpenSand
$0.0525 / second
Fal-US
$0.2646 / second
80% lower9
bytedance/seedance-2.0 (2160p)MachGen
$0.59 / second
Fal
$1.5552 / second
62% lower9
bytedance/seedance-2.0-fast (480p)Atlas Cloud
$0.027 / second
Fal
$0.2419 / second
89% lower9
bytedance/seedance-2.0-mini (480p)OpenSand
$0.0104 / second
Fal
$0.0721 / second
86% lower8
seedance-2-mini-unrestricted (480p)OpenSand
$0.0114 / second
SandBase
$0.0721 / second
84% lower3
seedance-2-5-unrestricted (1080p)OpenSand
$0.3482 / second
TOAPIS
$0.5881 / second
41% lower2
seedance-2.0-fast-unrestricted (480p)OpenSand
$0.0344 / second
SandBase
$0.0448 / second
23% lower2
seedance-2-unrestricted (2160p)OpenSand
$0.661 / second
TOAPIS
$0.7966 / second
17% lower2

Per second, before VideoRouter's 2% platform fee. For tiered models each row compares the resolution tier with the widest host-to-host gap. Built 2026-10-02 from the live catalog.

Ready to try it? Create a key and start from the quickstart.

Frequently asked questions

How many reference images can I send to Seedance?

Up to 9 images through input_references on models that support it, plus up to 3 video and 3 audio references on supporting models, with a combined cap of 12 on the documented Fal endpoint.

Can I combine a start image with reference images?

Not in the standard flow. start_image_url is image-to-video, while input_references and its video and audio siblings are reference-to-video, and the two modes are separate requests.

Do I need a separate model id for reference-to-video?

No. The mode is selected by which fields you include. Pin a host that supports the arrays you send, for example by suffixing the model id with /fal.

What happens if I send zero references in reference mode?

A reference-mode call must include at least one reference across the image, video and audio arrays. Send a plain prompt for text-to-video instead.

Keep reading

Using Seedance is one part of the job.

VideoRouter puts it next to dozens of other video and image models behind one API key, so you can compare providers, prices and fail over automatically. Compare providers on VideoRouter →