AI Avatar

Generate a talking-avatar video from a HeyGen avatar (or a raw image) + voice + script, or wired audio.

Overview

The AI Avatar node creates a talking-head video using HeyGen. You supply either a text script (HeyGen’s built-in TTS delivers the voice) or a pre-recorded audio track. Both paths produce a video of the chosen source speaking the content.

Two source modes control where the visual comes from:

Source What you provide Notes
Avatar (avatarSource: avatar, default) A HeyGen avatar look picked from the in-node avatar picker Animated by the Avatar IV / Avatar V engine
Image (avatarSource: image) A raw image — wired into the node’s Image input, pasted as a URL, or uploaded Animates your own photo/character directly. No avatar creation or training needed, so it works without a higher HeyGen tier. The engine selector is hidden (image mode uses its own engine)

Both source modes support the same speech modes, voice tuning, background, captions, and motion controls. Image mode is billed identically to Avatar IV (it is IV-class).

Two speech modes:

Mode What you provide Voice
Script + Voice (speechMode: text) A text script (up to 5,000 characters) and a voice ID picked from the in-node voice picker HeyGen TTS, driven by the chosen voice and optional voice speed
Wired Audio (speechMode: audio) An audio file wired to the audio input handle Exactly as recorded — no TTS. Audio inputs are capped at 10 minutes (600s); longer audio is automatically trimmed to 600s (you’ll see a notice on the result)

On the canvas

The node card itself walks you through the setup — you rarely need the settings panel for a first run. It has three looks:

State What the card shows
Empty (no avatar yet) Start with an avatar — a row of featured looks (one per presenter) you can pick right on the card, a search box that filters the whole catalog in place (the same search as the settings panel; results scroll, five per row), a Browse all N › button that opens the full catalog in the Choose an avatar window (below), and Use an image instead to animate your own portrait. In image mode the same spot becomes Start with an image — an upload zone (or wire an image into the Image input / paste a URL in settings) plus Choose an avatar to go back to the catalog.
Configured (avatar or image set, no video yet) Left: the portrait with an engine badge (AVATAR IV / AVATAR V, or SOURCE IMAGE), the look’s name and gender, and Change avatar (reopens the featured row with the current look highlighted; ✕ keeps it) or Replace image. Right: the voice strip (name, language · gender, a preview play button — click the voice to open the full voice picker right there, with search, language / gender filters and previews) and the script, editable in place with a live chars · ~duration estimate. When the Script input is wired, the card shows the incoming text read-only and says which node it comes from. In Wired Audio mode the right side shows the audio connection status instead of voice + script.
Generated (a video exists) The video result with the usual hover controls (versions, fullscreen, download, copy URL, NodarCut, settings). Every action lives in the bottom strip: Run makes another version (it lands on top of the earlier ones — the versions badge counts them), and New run hides the results and brings the setup card back so you can start fresh (pick another look, change the voice or the text) — nothing runs; a second click on New run restores the results exactly as they were. Starting a run (from the strip) always returns to the result view.
Failed Nothing earlier to show: the configured card stays editable and its status bar turns red with the error — retry with the strip’s Run. An earlier version on show: it stays, a red banner over the video names the failure, and New run / Run in the strip work as above. Earlier versions are never touched by a retry.

A status bar at the bottom of the card always says whether the node can run and what is still missing — Ready to run · avatar, voice and script are set vs Needs a voice before it can run — plus the engine and resolution that will render (HeyGen Avatar IV · 720p, or Image animation · 720p in image mode). The rules are the same ones the Run button enforces: text mode needs a script (typed or wired) and a voice; audio mode needs wired audio; avatar mode needs an avatar; image mode needs an image (uploaded, pasted or wired). A wired input counts as satisfied even before it has produced anything.

Picking a look on the card writes exactly what the settings-panel picker writes (avatar id, name, preview, Avatar V support, the look’s default voice when you have not chosen one, and the aspect ratio that matches its orientation). On an install without a HeyGen key, the featured row shows the same “Add a HeyGen key or connect nodaro.ai…” notice as the pickers; a workflow authored elsewhere still shows its configured avatar and voice from the node’s own data.

Choose an avatar (the catalog window)

Browse all N › on the card (and Browse all with filters › when a card search finds nothing — your search text carries over) opens a full-screen catalog window built around people, not the flat list of looks:

Area What it does
Header Title with a live count (12 people · 42 looks in view), a search box (name, look or scene; ⌘K / Ctrl+K focuses it; Escape clears an active search first, a second Escape closes), a Grid / List switch and ✕.
Left rail Libraries — All avatars, Your own looks (only when the account has custom looks), Recently used (the last looks you picked in this browser, most recent first) — then Gender and Scene chips (scenes come from the look names: Office, Livingroom, Studio…, spelling variants merged) and a Supports ✦ Avatar V switch.
Cards One card per presenter — a HeyGen avatar group, so two different presenters who share a first name stay two cards and a group whose looks carry descriptive names (“Charismatic Professional 3”) stays one — with the portrait, a N looks count, the violet ✦ tag when any look supports Avatar V, name and Gender · most common scene; a pink check marks the selected person. Load 24 more pages the list (fewer on the last page). The cards form one keyboard group: Tab reaches it once, the arrow keys move the selection.
Detail column The selected look large (with the ✦ AVATAR V badge when it applies), Person — Look, the person’s LOOKS as chips (pick one to switch; with a search or filter active only the looks that match are listed), DETAILS (engines, orientation, default voice, number of looks — and, on Cloud, the node’s run cost), Preview voice (plays the look’s default voice sample) and Use this avatar. When a filter hides the person on show, the column follows the view (the node’s current avatar if it is still in view, else the first person).

Use this avatar writes exactly what a card tile writes (avatar id, name, preview, Avatar V support, the look’s default voice when you have not chosen one, and the aspect ratio that matches its orientation) and closes the window. A look HeyGen is still building shows its status and cannot be used until it is ready.

Selecting an Avatar and Voice

The config panel includes two rich pickers:

Both pickers return empty when HeyGen is not configured for the deployment; an “HeyGen API not configured” notice appears in that case.

The catalogs load progressively. HeyGen serves its ≈7,000 photo-avatar looks in ≈140 pages (plus ≈2,500 voices in one slow call), and a cold fill takes about a minute and a half — you never wait for that. The server starts filling at boot, answers with the pages it already has, the pickers (and the on-node quick pick) render them at once, and the counter reads N avatars · loading more… until the list is whole. Pick as soon as you see what you want; the rest keeps arriving behind it — and only the new pages cross the wire on each poll (the browser asks for what it does not have yet), so a big catalog never re-downloads. A whole list arrives in chunks of 1,500 too, so the first tiles paint before the ≈4 MB list has finished streaming.

The catalog is shared and survives restarts. Once a fill completes, the server publishes the whole list to Redis (heygen:catalog:v1:avatars / :voices); a freshly deployed or restarted instance adopts that snapshot at boot instead of refetching HeyGen, and when several API instances run, one of them refreshes the list for all of them (a Redis lock keeps the others from duplicating the ≈140-call fill). The snapshot is refreshed in the background every HEYGEN_CATALOG_REFRESH_HOURS (default 24) — a stale list is served while the refresh runs, and the browsers switch to the new list on their next poll. Every instance also re-checks the snapshot’s stamp about twice a minute (a few bytes), so when one instance publishes a new list — a manual refresh, or two instances that filled on their own after a cold start — the others pick it up within seconds rather than at their own next daily cycle. Without Redis (or while it is unreachable) each instance falls back to its own in-memory copy exactly as before.

Your own looks are loaded separately. The account’s private photo-avatar looks come from their own small endpoint (GET /v1/heygen/avatars/private), are never cached in the shared snapshot, and are re-read every couple of minutes while a picker is open — so a look you just created in HeyGen shows up (first in the list, and under the Custom filter) without waiting for the daily refresh. A look HeyGen is still building is shown with a Processing… badge (a failed one with Failed) and cannot be picked until it is ready; the on-node featured row skips such looks. On a nodaro.ai-connected install this list is the connected account’s own looks.

Refresh on demand. Integrations → HeyGen avatar catalog → Refresh now refetches every list from HeyGen at once instead of waiting for the daily cycle (the fill takes about two minutes; pickers show the new lists the next time they open). Any signed-in user may do this on Community; where the edition has admins, admins only. The same action is POST /v1/heygen/catalog/refresh with a first-party session (never an API or app token) — it answers at once with what each catalog did (started, already-running, locked-elsewhere when another server of the environment is mid-refresh, adopted when a newer shared list already existed, unconfigured without a HeyGen key); on a nodaro.ai-connected install it forgets the relayed copies so the next pick re-pulls the cloud’s lists.

Configuration

Field Type Default Description
Source Select avatar avatar = HeyGen avatar look; image = animate a raw image
Speech Mode Select text text = script + voice; audio = wired audio input
Avatar Picker HeyGen avatar ID (required in avatar source mode)
Source Image Image input / URL / Upload Source image (required in image source mode) — wire an image node into the Image input, paste a URL, or upload
Script Textarea Spoken text (required in text mode, max 5,000 chars)
Voice Picker HeyGen voice ID (required in text mode)
Voice Speed Slider 1.0 Speaking rate, range 0.5–1.5 (text mode only)
Engine Select avatar-iv avatar-iv = Avatar IV; avatar-v = Avatar V (premium). Avatar source mode only — hidden in image source mode
Resolution Select 720p Output resolution: 720p, 1080p, 4k
Aspect Ratio Select 16:9 16:9 (landscape) or 9:16 (portrait / vertical)
Captions Toggle off Burn auto-generated captions into the video

Inputs & Outputs

Inputs:

Outputs:

Credit Pricing

Credits are metered by the actual length of the generated video. A hold is placed when the job starts; any unused amount is refunded automatically when the job completes — so you only pay for the seconds you get.

Approximate cost

Engine 720p 1080p 4K
Avatar IV ~3.8 credits/sec ~5 credits/sec ~10 credits/sec
Avatar V ~5 credits/sec ~6.3 credits/sec ~12.5 credits/sec

Examples (Avatar IV, 720p): a 30-second clip ≈ 113 credits; a 1-minute clip ≈ 225 credits. Higher resolutions and Avatar V cost proportionally more. Captions add no extra cost.

The exact credit cost is shown in the editor before you run, and the final charge always reflects the real clip length.

Reserve & refund

Graceful Degradation

If HeyGen is not configured for the deployment:

Best Practices

Common Use Cases

Tips