Lip Sync

Sync audio to a character’s face to create a talking head video.

Overview

The Lip Sync node takes a portrait image and an audio track (speech/voiceover) and generates a video where the character’s lips move in sync with the audio. Supports optional motion prompts for head and expression movements.

It can also dub an existing video: HeyGen Lipsync Precision, Sync Lipsync 2 Pro, Sync Lipsync v3, and Volcengine Lip Sync take a source video plus a replacement audio track and re-animate the lips to match the new speech, billed per second of output. Volcengine is the cheapest modern dubbing option and the only one with multi-speaker scene detection + speaker ID (basic mode).

Configuration

Field Type Default Description
Provider Select kling-avatar AI model for lip sync
Resolution Select 720p Output resolution: 480p/720p (most KIE providers); OmniHuman 1.5 is 720p/1080p (default 1080p); the full seedance-2 adds 1080p / 4K and seedance-2-5 adds 1080p (seedance-2-fast and seedance-2-mini are 480p/720p only)
Motion Prompt Textarea Optional: describe head/expression motions (KIE providers only)
promptPrefix / promptSuffix text Optional pre/post text wrapped around the prompt at run time (settings panel → Pre & post text; hidden from app users; captured by presets). See Prompt pre & post text.

OmniHuman 1.5 options

omnihuman-1-5 is ByteDance’s premium prompt-directed talking avatar (image + audio → performing avatar). It animates people, pets, or anime at any aspect ratio, and uses the motion prompt to direct the performance.

Field Type Default Description
Motion Prompt Textarea Directs the performance (e.g. “sing confidently into a microphone”)
Resolution Select 1080p 720p or 1080p (no 480p)
Fast Mode Toggle Off Trade some quality for faster generation (pe_fast_mode)
Seed Number -1 Reproducibility — same seed + inputs → near-identical result. -1 = random

Seedance 2 (image + audio avatar)

seedance-2, seedance-2-fast, seedance-2-mini, and seedance-2-5 are also offered on the lip-sync surface. They are ByteDance’s multimodal video models, which do native phoneme-level lip sync in 8+ languages when fed an audio track alongside a portrait — routed through the image-to-video provider with the audio passed as reference_audio_urls (not the dedicated lip-sync flow). The full seedance-2 accepts 1080p and 4K output; seedance-2-5 accepts 1080p (no 4K); seedance-2-fast and seedance-2-mini are 480p/720p only (no 1080p SKU). Full controls, durations, and per-second pricing live on the Generate Video page (e.g. 4K 8s = 4160 cr, 1080p 8s = 2040 cr). seedance-2-fast requires each reference audio clip to be ≤ 15.2 s; seedance-2-5 allows up to 30 s.

minimax-h3 (MiniMax Hailuo 3) rides the same audio-as-reference mechanism: portrait + voice line route through its reference-to-video endpoint with the audio as reference_audio_urls. Output renders at H3’s 2K default — the lip-sync surface’s Resolution select never names H3’s own 768P tier, so any value collapses onto the 2K-rate duration composite — audio is always on, reference audio caps at 15 s per clip, and the reservation tier is the 8s duration composite (minimax-h3:8s = 730 cr). See Generate Video for the full pricing formula (768P applies on the generation surfaces, not here).

HeyGen Lipsync Precision options

Field Type Default Description
Dynamic Duration Toggle On Adjust the output length to match the new audio
Remove Music Track Toggle Off Strip background music from the source video
Speech Enhancement Toggle Off Improve speech clarity in the output

Sync Lipsync 2 Pro options

Field Type Default Description
Sync Mode Select loop Behavior when audio/video durations differ: loop, bounce, cut_off, silence, remap
Temperature Slider 0–1 0.5 How expressive the lip sync can be
Active Speaker Detection Toggle Off Lip-sync whoever is speaking in the clip

Sync Lipsync v3 options

sync-lipsync-v3 is the fal.ai-hosted sync.so sync-3 model. Expressiveness, obstruction handling and frame reasoning are managed natively by the model (its temperature is ignored), so it exposes two levers:

Field Type Default Description
Sync Mode Select cut_off Behavior when audio/video durations differ: loop, bounce, cut_off, silence, remap
Active Speaker Detection Toggle Off Lip-sync whoever is speaking in the clip (sync.so’s auto-detect, v3 detector). Turn it on for videos with more than one person on screen.

Volcengine Lip Sync options

volcengine-lipsync is the KIE-hosted Volcengine video-to-video dubbing model. It re-syncs an existing clip’s lips to a new vocal track and is the cheapest modern dubbing option (20 CR/s). Output length follows the audio (the source video is trimmed if longer, looped if shorter).

Field Type Default Description
Mode Select lite lite = single-person frontal, faster · basic = complex scenes, enables multi-speaker scene detection + speaker ID
Separate vocals Toggle Off Strip background noise from the driving audio
Scene detection + speaker ID Toggle Off Basic mode only — segment scene cuts and identify who is speaking (multi-speaker clips)
Loop video if audio is longer Toggle On Lite mode only — loop the source video when the audio runs longer than it
Reverse loop Toggle Off Lite mode only — ping-pong the loop (requires looping to be on)
Template start time Number (s) 0 Where in the source video to start driving the lips (advanced)

Inputs & Outputs

Inputs:

Outputs:

HeyGen Lipsync Precision, Sync Lipsync 2 Pro, Sync Lipsync v3, and Volcengine Lip Sync replace the audio on a video input (not a portrait image).

Audio Length Limits

Provider Max input audio
kling-avatar (Standard) 5 minutes
kling-avatar-pro (Pro) 5 minutes
infinitalk 15 seconds
omnihuman-1-5 60 seconds
heygen-lipsync-precision 5 minutes*
lipsync-2-pro 5 minutes*
sync-lipsync-v3 5 minutes*
volcengine-lipsync 5 minutes

For the KIE providers (kling-avatar(-pro), infinitalk, volcengine-lipsync), audio longer than the cap is auto-trimmed before the upstream call. Long-audio runs on kling-avatar(-pro) and volcengine-lipsync can take tens of minutes — the editor polls for up to ~1 hour before timing out.

* For HeyGen Lipsync Precision, Sync Lipsync 2 Pro, and Sync Lipsync v3 the 5-minute figure is the per-second credit-reservation ceiling, not a hard trim — longer clips reserve at the 5-minute tier.

Credit Cost

Kling AI Avatar and OmniHuman 1.5 bill per-second. Credit reservation buckets to the next supported tier; any reserved credits beyond the actual cost are refunded once the job completes.

Provider Per-second 15s 30s 1min 2min 5min
kling-avatar (720p) 20 CR/s 300 600 1,200 2,400 6,000
kling-avatar-pro (1080p) 40 CR/s 600 1,200 2,400 4,800 12,000
omnihuman-1-5 (720p/1080p) ~67.5 CR/s 1,020 2,030 4,050

omnihuman-1-5 is capped at 60s of audio (longer is auto-trimmed), so only the 15s / 30s / 60s tiers apply. Resolution (720p vs 1080p) does not change the price.

HeyGen Lipsync Precision, Sync Lipsync 2 Pro, and Sync Lipsync v3 also bill per second of output, bucketed to the same 15s / 30s / 1min / 2min / 5min tiers:

Provider 15s 30s 1min 2min 5min
volcengine-lipsync 300 600 1,200 2,400 6,000
heygen-lipsync-precision 510 1,010 2,010 4,010 10,010
lipsync-2-pro 630 1,250 2,500 5,000 12,490
sync-lipsync-v3 1,000 2,000 4,000 8,000 20,000

The credit chip updates once the audio is wired (the node probes its duration). When the duration is unknown, the 5-minute tier is reserved.

Supply audioDurationSec for accurate pricing. When calling sync-lipsync-v3 or volcengine-lipsync via API/SDK, pass the output duration in seconds so the reservation buckets to the correct tier. If it is absent, the request is billed at the 5-minute ceiling (sync-lipsync-v3: 20,000 CR; volcengine-lipsync: 6,000 CR) with no refund — these per-second models commit the reserved bucket verbatim. The editor probes the duration automatically once the audio is wired.

InfiniTalk and Hailuo Avatar use flat per-call pricing (see admin → Models).

Best Practices

Common Use Cases

Tips

Provider cost reporting

For Lipsync 2 Pro, the recorded provider-cost estimate uses the finished video’s measured duration and the provider’s output-second rate. It does not use GPU runtime. If the output duration cannot be read, provider cost remains unknown. Credits still use the duration tier displayed before generation.