A 90-Second Onboarding Lesson in 5 Shots — Full VO3 AI Workflow

ByteDance's Seedance 2.5 landed earlier today with claims of 30-second single generations. Here's the exact 5-shot workflow for turning that into a finished onboarding lesson — with the verbatim prompts, an identity block that fights face drift, and an honest list of what still breaks.
The most useful explainer shot is the most boring one
Earlier today, ByteDance's Dreamina team pushed Seedance 2.5 globally — and the headline claim buried in the announcement matters more for training video than for anything cinematic: single generations up to 30 seconds, with long-take consistency stretched toward three minutes.
That's not a filmmaking feature. That's an onboarding feature. Because the shot an explainer video actually lives or dies on is the least glamorous one in the catalogue: a person, centered, talking straight into the lens, not moving. Every second you don't have to cut is a second your instructor's face can't change on you.
So this is a build, not a review. Five shots, five prompts you can paste, one 90-second lesson at the end. If you want to skip the reading and start assembling, the AI explainer video maker has the beat structure pre-loaded.
▶️ Try in VO3 AI
Build your first explainer beat → — one shot, ~8 seconds, native audio.
One caveat before we spend anything: the 30-second and 3-minute numbers above come from the launch post, not from a spec sheet I've independently timed. Treat them as the vendor's claim until your own render clock says otherwise.
🧰 The Workflow Block
Model: Seedance 2.5 — picked specifically for the long single take. Use the Seedance AI video generator when a beat runs past 10 seconds and the speaker's face has to survive the whole thing. For 5–8 second cutaways, a faster model is cheaper and just as good.
Prompt (verbatim — this is Beat 1):
Priya Raghavan, 34, South Indian woman, shoulder-length black hair tied back,
thin gold hoop earrings, charcoal crewneck sweater, warm brown eyes, small
mole above the left eyebrow. Neutral grey studio wall, soft key light from
camera left, 50mm lens, eye-level, shallow depth of field.
Beat 1 of 5. Medium shot, waist up, centered. She speaks directly to camera
with a small welcoming smile, hands relaxed. She says: "Before you touch the
dashboard, there are exactly three things you need to set up. This takes four
minutes."
Static tripod. No camera move. 8 seconds, 16:9, native audio, clean room tone.
Credits / cost: VO3 AI's starter pack is $2.99. I'm quoting the published pack price rather than inventing a per-second receipt — per-model credit costs move, and Seedance 2.5 is one day old. Pull the live per-model rate before you commit a budget. Practically: a 90-second lesson is 5 finished beats plus re-rolls, and re-rolls are the line item that surprises people, not the beats.
Output: 5 beats × 8–20s → ~90 seconds assembled. Generation runs roughly 1–3 minutes per beat depending on length and queue.
What the demo clips are (and aren't)
Both videos below are talking-head and macro work from other projects — a locksmith and a pho shop, not a software lesson. I'm showing them anyway because shot grammar transfers and subject matter doesn't. The locksmith clip is the exact grammar of Beat 1: one human, eye-level, direct address, no cutting. If this holds up, your instructor beat holds up.
Generated with VO3 AI — direct-to-camera talking head. Different subject, identical shot grammar to an explainer host beat.
Step 1 — Break the lesson into beats before you write a single prompt
An explainer isn't a script, it's a sequence of claims. Write five one-sentence claims first. If a claim needs more than one sentence, it's two beats.
Run this in any chat model to get there fast:
Turn this onboarding doc into exactly 5 video beats. Each beat = one claim,
spoken in ONE sentence of at most 22 spoken words, plus a note on whether it
should be (a) instructor on camera or (b) a demonstration cutaway.
No intros, no outros, no "in this video we'll cover." Start at the first
useful instruction.
[paste your doc]
The "no intros" line matters. AI writes a 15-second preamble by default, and preamble is where course video loses viewers.
Step 2 — Freeze the instructor into an identity block
This is the single technique that separates a lesson from five unrelated clips. Write one paragraph describing your instructor down to an asymmetric detail — a mole, a chipped nail polish, one earring style — and paste it verbatim, unedited, as the first paragraph of every beat prompt.
The same discipline showed up in a workflow doing the rounds this week, using a fixed character prompt to hold one face across outfit and style changes:
Better than prompt text alone: generate one still of your instructor, then feed it as a reference. Reference-to-video gives the model a pixel target instead of an adjective, which is a much harder thing to drift away from.
Step 3 — Generate the on-camera beats, long and static
For beats 1, 3, and 5 (the ones where a human explains), the rules are narrow:
- Static tripod, always. Camera moves are where faces morph.
- Quote the spoken line in the prompt. Don't describe what she says — write it.
- One idea per beat. Two claims in one generation = mismatched lip timing on the second.
[PASTE IDENTITY BLOCK VERBATIM]
Beat 3 of 5. Same studio, same lighting, same wardrobe. Medium shot, waist up,
centered, matching Beat 1 framing exactly. She glances briefly down and to her
right as if at a screen, then back to camera. She says: "Set your notification
window before you invite anyone — changing it later resets every teammate's
preferences."
Static tripod. No camera move. 12 seconds, 16:9, native audio.
Run the consistency check yourself. Export a still from Beat 1 and a still from Beat 5, put them side by side at 100%, and look at three things: hairline, earrings, sweater neckline. I'm not going to tell you the identity block makes them identical — it doesn't. It makes drift small enough to survive a cut, most of the time. The check is how you find the one beat that needs re-rolling.
Step 4 — Generate demonstration cutaways for the beats that show, not tell
Beats 2 and 4 are where the lesson proves something. Two hard-won rules:
Never ask AI to render your actual UI. It will produce a plausible-looking dashboard with nonsense labels, and nonsense labels in a training video destroy trust instantly. Screen-record the real product for anything with text.
Do use AI for the physical-world analogue — hands, objects, tactile close-ups that make an abstract idea land. That's what the macro grammar below is for:
Generated with VO3 AI — extreme macro, handheld, slight natural shake. The cutaway grammar you want for demonstration beats.
Extreme macro close-up, handheld with slight natural shake, shallow depth of
field, warm desk lamp light from the left. Two hands slide a single index card
out of a dense stack of paper and set it apart on a wooden desk. The isolated
card is in sharp focus; the stack falls out of focus behind it.
No text, no writing, no logos. 6 seconds, 16:9, no dialogue, quiet room tone.
The no text, no logos line is not optional. Left out, models hallucinate garbled lettering into every surface.
▶️ Try in VO3 AI
Generate your instructor beat and cutaway → — same identity block, two shots, one sitting.
Step 5 — Assemble, then fix the audio last
Stitch in beat order and cut on the last frame of speech, not on silence — trailing silence is where AI faces do their strangest micro-movements. Multi-beat stitching lives in the long video editor.
Then do one audio pass. Native model audio is good enough for a first cut but inconsistent in level across beats; normalize everything to one loudness target. If your lesson runs past four minutes, record narration separately and treat the video as illustration — that's closer to how the AI podcast video maker workflow handles long-form.
Pablo Stanley's breakdown of his own pipeline lands on the same shape — generate the anchor frame, generate video, then finish in a real editor:
What still breaks
No hedging here — these are the failure modes you will hit:
- Hands. Any beat where the instructor gestures near her face is a coin flip. Keep hands low or out of frame.
- Lip-sync past ~15 seconds. Drift accumulates. If a beat needs 20 seconds of speech, split it.
- Re-roll rate. Budget for roughly one re-roll per finished beat. On a 5-beat lesson that's ~8–10 generations, not 5. This is the number people get wrong when they price a project.
- Wardrobe micro-drift. Necklines and earrings shift between beats more than faces do. Nobody notices at 1× playback; everybody notices in a side-by-side still.
- Anything with readable text. Screen-record it. Always.
Worth watching, too: digital-avatar tooling is compressing this whole first step. Mirage's Avatar X claims a clone from 10 seconds of source footage — if that holds, the identity block stops being a prompt problem and becomes an upload.
The five steps, for skimmers
- Beat it out — 5 one-sentence claims, no intro.
- Freeze the instructor — one identity block, pasted verbatim into every prompt.
- On-camera beats — static tripod, quoted dialogue, one idea each.
- Cutaways — physical-world analogues only; screen-record real UI.
- Assemble, normalize audio last — cut on the last frame of speech.
Try It Yourself
Start with Beat 1 only. One prompt, one 8-second render — if the face and the read hold up, the other four are mechanical. The AI explainer video maker is the fastest entry point; if you're producing this to sell a course rather than teach one, the AI course promo video template uses a punchier beat structure.
Everything runs on the same model catalogue at VO3 AI — Seedance 2.5, Veo, Kling, and the rest, one credit balance across all of them. Open the AI video tool and paste the Beat 1 prompt above verbatim; swap Priya's identity block for your own instructor.
▶️ Try in VO3 AI
Start your explainer video → — paste the identity block, run Beat 1, see if it holds.
Ready to Create Your First AI Video?
Join thousands of creators worldwide using VO3 AI Video Generator to transform their ideas into stunning videos.
📚 Related Posts:
What is VO3 AI Video Generator: The Ultimate AI-Powered Video Creation Platform
Discover VO3 AI Video Generator - the revolutionary AI video creation platform
Read More →VO3 AI vs. Veo3 — What's the Difference?
Understand the key differences between VO3 AI and Google's Veo3
Read More →How to Use VO3 AI Video Generator: Complete Guide
Master VO3 AI Video Generator with our comprehensive tutorial
Read More →VO3 AI Video Generator - Where imagination meets innovation
Built on top of multiple AI video models including Veo3. Start your creative journey today and join the future of video creation.