Step 2 — Generate the 3 cinematic moving shots (Seedance)
Step 3 — Design an avatar "look" for intro/outro
Step 4 — Assemble all 5 scenes on the timeline & render
Step 5 — Translate to any of 175+ languages, matched lip-sync
Step 0 — Create Your Avatar Clone
Skip if your avatar clone is already made.
0.1 — Create Avatar. From the dashboard, go to the Avatars tab in the sidebar and click Create Avatar (top right).
0.2 — Clone a real person. On the "Create a new avatar" screen, choose Clone a real person — this makes an avatar that looks, moves, and sounds like you.
0.3 — Choose language. Click the language button (defaults to English) and pick the language you'll speak in.
0.4 — Get in frame & record. Make sure you're fully visible inside the frame, then click Start Recording.
Great-recording tips: speak with high energy (expressive voice = expressive avatar) · keep your position stable in frame · good lighting + a quiet space matter a lot.
0.5 — Read the script clearly. Read out loud, including the number at the end. Natural and confident — this captures your voice, expressions, and movement.
0.6 — Name & create. Give the avatar a name (your own works great), click Create Avatar. Not happy with the take? Hit Record again before moving on.
0.7 — Use this voice. When prompted to set up a voice, choose Use this voice to keep the voice from your footage.
1.3 — Replace the bracketed variables (see the guide below the prompt).
1.4 — Review: no [INSERT HERE] placeholders left.
1.5 — Hit send. ChatGPT returns a 5-scene script ready to drop into HeyGen.
ChatGPT Prompt Template
You are helping me adapt a 5-scene video template to my own video idea. I will paste the template below. The template has a specific structure I must preserve:
2 static shots (avatar stays still, camera locked — used for intros, outros, or key messages). These have Dialogue only — no Scene Description.
3 moving shots (camera moves — orbit, push-in, tracking — avatar performs an action in a dynamic environment). These have both Scene Description and Dialogue.
My video idea:
Avatar gender: [ CHOOSE ONE - male / female / non binary ]
Service / Topic: [INSERT HERE - e.g. "3 things my digital twin can do better than me"]
Target audience: [INSERT HERE - e.g. "Creators, founders, and marketers (25–45) curious about AI avatars"]
Tone / vibe: [INSERT HERE - e.g. "Playful and cinematic with a professional edge — fun visuals, confident delivery"]
Key message: [INSERT HERE - e.g. "My twin goes places I can't — so my content can too"]
Avatar actions for the 3 moving shots:
Scene (moving shot 1) action: [INSERT HERE — e.g. "Flying through the clouds in a superhero pose"]
Scene (moving shot 2) action: [INSERT HERE — e.g. "Petting a friendly T-Rex in a prehistoric jungle"]
Scene (moving shot 3) action: [INSERT HERE — e.g. "Walking on the surface of Mars in a sleek spacesuit"]
Your task:
Rewrite all 5 scenes to support my video idea. Keep the exact structure for each scene type:
Static shots: Scene number · Scene Title · Script only. Do NOT add a Scene Description.
Moving shots: Scene number · Scene Title · Scene Description · Script.
Preserve which scenes are static vs moving — do not switch them.
For each moving shot, build the Scene Description around the avatar action I specified, following the Seedance Core Formula in order. Include:
- Avatar gender reflected in the Subject ("a man" or "a woman") and any gendered clothing terms.
- Avatar clothing that fits the action and environment (spacesuit for Mars, wetsuit for underwater, tailored suit for a boardroom — never generic).
- One camera movement + one lens matched to the action's energy and scale.
- Vivid environment with lighting, atmosphere, and time of day.
- Avatar must face the camera with face clearly visible — no back-of-head shots, no profiles obscuring the face. Camera movement should always end on or pass through a moment where the avatar's face is toward the lens.
- End with positive-phrased constraints and preserve composition and colors.
Script rules:
- Keep each Script line of action shots under 10 words, and each static shot under 15. Natural, matched to the tone, and connected to what the avatar is doing in that scene.
- Do not add or remove scenes. Return only the rewritten template.
Tone consistency rule (CRITICAL):
The Tone / vibe I set above is binding for every scene — both Scene Descriptions and Scripts. All 5 scenes must feel like they belong to the same video.
- Funny / playful → bright vibrant lighting, whimsical/upbeat environments, light expression (smiling, laughing, winking), witty casual script. Even an intense action must be staged playfully — sunny, exaggerated, fun.
- Cinematic / epic → sweeping camera moves, dramatic golden-hour or moody lighting, rich color grading, confident delivery.
- Professional / corporate → clean lighting, polished wardrobe, restrained camera moves, clear direct script.
- Dreamy / aspirational → soft light, slow movement, warm or pastel palette, gentle delivery.
Before writing each scene, check: does this lighting, environment, and script line match the tone I set? If a scene's natural mood would clash, bend the environment and script toward the tone rather than defaulting to the action's stereotypical mood.
Template to adapt:
Scene 1 — Avatar Intro (Static Shot)
Script: "Now I'm going to show you what my AI video clone can do… better than me."
Scene 2 — Paris
Scene Description: A woman in a stylish beige trench coat, black turtleneck, and sunglasses pushed up on her head, carrying a baguette under one arm. She walks briskly and confidently toward the camera, face fully visible. Eiffel Tower behind her during golden hour, warm Paris street atmosphere, glowing café lights, soft breeze moving hair and coat naturally. Medium tracking shot on a 50mm lens, smooth cinematic motion. Cinematic, shallow depth of field, elegant European aesthetic. Smooth motion, stable framing, anatomically correct, natural proportions, consistent lighting, preserve composition and colors.
Script: "My avatar can speak French… Bonjour."
Scene 3 — Skydiving
Scene Description: A woman in a fitted black-and-orange skydiving jumpsuit, harness secured, goggles pushed up on her forehead. She glides smoothly through the air with parachute fully open, face turned clearly toward the camera. High above bright blue clouds at midday, sunlight beaming overhead, wind moving naturally through hair and fabric. Wide aerial tracking shot on a 24mm lens gliding alongside her. Cinematic aerial perspective, realistic parachute physics, film grain. Smooth motion, stable framing, anatomically correct, natural proportions, detailed natural hands, consistent lighting, preserve composition and colors.
Script: "My avatar can jump with a parachute… and honestly, it looks way cooler than I imagined."
Scene 4 — Candy Fantasy World
Scene Description: A woman in a playful pastel outfit — soft pink sweater and white pants — laughs while biting into an oversized lollipop, face fully visible as the camera completes its arc. Surrounded by a vibrant fantasy candy universe of giant donuts, chocolate rivers, gummy bear trees, cotton candy clouds, and glowing sweets under bright whimsical daylight. Wide orbit shot on a 35mm lens, smooth steady arc around her. Vivid cinematic colors, surreal fantasy atmosphere, shallow depth of field. Smooth motion, stable framing, anatomically correct, natural proportions, detailed natural hands, consistent lighting, preserve composition and colors.
Script: "My avatar can eat whatever it wants… without counting calories."
Scene 5 — Avatar Outro (Static Shot)
Script: "So… what should my avatar do next?"
How to fill the variables
Avatar gender — male / female / non-binary. Drives wardrobe & physical description in every scene.
Video topic — what it's about. Stuck? Keep the example.
Target audience — be specific.
Tone & vibe — a few words for the feel.
Key message — the one takeaway.
3 avatar actions — three things your avatar does on screen. Get playful — fun actions make the video pop even on a serious topic.
For a channel/business: make the 3 action shots showcase what you do. Realtor → walking a luxury villa, handing over keys, sunlit kitchen. Fitness coach → demoing a workout, meal-prepping, celebrating a client win. Just for fun: pet a dinosaur, time-travel, jump with a parachute.
The 5-Scene Structure (memorize this)
Scene 1 — IntroStatic Dialogue only. Avatar talks to camera. Built later in AI Studio.
Scene 2 — MovingSeedance Scene Description + Dialogue. Generated in Avatar Shots.
Scene 3 — MovingSeedance Scene Description + Dialogue. Generated in Avatar Shots.
Scene 4 — MovingSeedance Scene Description + Dialogue. Generated in Avatar Shots.
Scene 5 — OutroStatic Dialogue only. Avatar talks to camera. Built later in AI Studio.
Step 2 — Generate Your Cinematic Shots
2.1 — Open Avatar Shots. Sidebar → Avatar Shots. Up top, switch to the Cinematic tab.
2.2 — Select Cinematic Shots. Make sure Cinematic (Seedance 2) is selected — short cinematic clips with real camera movement, lighting, full-body performance.
2.3 — Open the avatar selector. Click the + (or avatar thumbnail) in the shot box and select your avatar.
⚠️ Consent required: only avatar groups with verified consent appear in Avatar Shots. If yours isn't in the picker, finish the consent step on that avatar first.
Consistency tip: keep the same avatar look (same outfit, same setting style) across all scenes so the final video feels like one cohesive piece.
2.4 — Shot settings. Duration → 5s · Resolution → 720p · Aspect ratio → Portrait (TikTok/Reels/Shorts) or Landscape (YouTube/web/ads).
Why 5s + 720p? Shorter, lower-res = fewer credits + faster renders, so you can try more prompt variations. You can upscale the FINAL edited video to 4K later in HeyGen.
2.5 — Copy Scene 2 from ChatGPT — the entire block (Scene Description + Dialogue). Both matter: description = what to show, dialogue = what the avatar says.
⚠️ One scene at a time. Don't paste Scenes 2, 3, 4 together — HeyGen will make one confused shot instead of three clean ones.
2.6 — Paste & generate. Drop the scene into the Describe Your Shot box → hit Generate.
2.7 — Repeat for Scene 3 & Scene 4, one at a time. No need to wait for renders — kick all three off, then move to the next step while they process in the background.
Why skip Scenes 1 & 5 here? Those are your static intro/outro — no camera movement needed. You build them in AI Studio, avatar talking straight to camera.
Step 3 — Create Your Avatar Look
Design a custom look that matches the world of your cinematic shots — the "face" of your intro/outro.
3.1 — Open AI Studio from the sidebar.
3.2 — Open the Avatar panel. With your first scene open, click the Avatar icon (right side) to open Avatar & Voice. Click your avatar box → your look gallery opens → hover "Add a look" → choose Design with AI (recommended for beginners).
3.3 — Pick your path: Path A — Remix a Style (easiest): browse pre-made looks (office, café, studio, outdoor, beachside), hover the one you like → Remix look.
Path B — Custom Prompt (more control): type a description covering 3 things — pose, environment, clothing. Example: "In a modern office with beige blazer, sitting confidently, soft natural light."
⚠️ Match aspect ratios. Your look's aspect ratio must match your Seedance shots' ratio.
3.4 — Click View to see your looks, pick the one to use. ✅ Done.
Step 4 — Assemble Your Video
4.1 — Build the timeline. Your timeline starts with 1 scene. Click + Add scene (bottom-left) 4 more times → 5 scene thumbnails in a row.
4.2 — Add Scene 1 & 5 scripts. From ChatGPT, copy ONLY the quoted line after "Script:" for Scene 1 → paste into Script box 1. Same for Scene 5 → Script box 5.
⚠️ Dialogue only for scenes 1 & 5 (static, avatar speaks to camera). Boxes 2, 3, 4 stay empty — your Seedance shots already have dialogue baked in.
4.3 — Add Seedance shots. Click Scene 2 in the timeline → right sidebar → Media icon → My Media tab → click your Scene 2 shot (it drops into the canvas) → resize to fill → select "Fit to scene" under Playback. Repeat for Scene 3 and Scene 4.
Don't see your shots in My Media? (1) Still rendering — check the Avatar Shots tab. (2) Finished but not showing — refresh AI Studio (Cmd+R / Ctrl+R).
Turn your finished English video into 175+ language versions — same avatar, same voice, perfect lip-sync. HeyGen's biggest unlock.
5.1 — From Projects, click the green Translate icon (あA symbol) in the far-left sidebar.
5.2 — Drag your finished video onto the "Drop your videos here to translate" zone.
5.3 — In Translate to, type a language name (English, Spanish, Portuguese, Japanese, Hindi, Mandarin…). You can pick multiple at once.
5.4 — Hit Translate. HeyGen transcribes → translates → re-voices → re-syncs lips. Progress bar ticks up in the sidebar.
5.5 — At 100%, open the video → Download. Two options: Video (no captions) or Captions (translated captions burned in).
Captions vs no captions: No captions → clean final video, or add custom captions later. With captions → social (Reels/TikTok/Shorts) where ~85% watch on mute. You can download the same video both ways — no need to re-translate.
Where This Fits Your Channels
Ringside Row — avatar host doing fight breakdowns; the Seedance moving shots handle "in the arena" energy.
Life Max 101 / Loan Insider — an avatar face adds a presenter without you filming daily; translate for reach.
Translation = free distribution. One English video → Spanish, Portuguese, Hindi versions multiplies a single script across audiences.
Batch the clone once. Record the Step 0 avatar a single time; every future video skips straight to Step 1.
Saved 7/21/26 from the HeyGen workshop guide · creation tool (input-diet exempt) · for non-MLO channels · NOT compliance-cleared for loan or notary marketing — anything MLO/notary needs NMLS #2067609, NEXA disclosures, and adscompliance approval first.