The Complete Seedance 2.5 Prompting Guide
Ori Silver
·
Co-founder, Maxfusion AI
·

Seedance 2.5 is not a text-to-video model you describe things to. It is a production system you brief.
It takes up to 50 reference materials in a single generation. Images, videos, audio, all at once. It generates 30-second videos, edits existing footage, extends clips forward and backward, and bridges 2 videos with a seamless transition. Every one of those capabilities is unlocked or wasted by how the prompt is written.
This guide covers the full prompting system: the core formula, how to bind reference materials to roles, multi-reference workflows, long videos, editing, extension, keyframes, storyboards, blockouts, transitions, emotional direction, and camera language. It is long because the model can do a lot. Everything here runs on Seedance 2.5 inside Maxfusion AI.
One thing to be clear on before we start. Prompt writing affects instruction following, material consistency, and controllability. It does not change what the model is fundamentally capable of. Visual quality, human realism, physics, and complex camera execution are still shaped by model capability, your input materials, and generation randomness. A good prompt raises the odds. It does not guarantee the shot.
The core formula
Every Seedance 2.5 prompt is built from the same parts, combined as needed:
Subject + Action or Event + Scene and Environment + Visual Style + Camera + Audio
Subject and action are the foundation. Who or what is doing what. Scene covers location, time, weather, spatial relationships, and background state. Visual style covers lighting, color, materials, texture, and mood. Camera covers shot size, angle, movement, focus, and cuts. Audio covers dialogue, voice, ambience, sound effects, and music.
The basic template:
<Subject> performs <primary action or event> in <scene and environment>.
The visuals feature <visual style>.
Use <shot size, camera angle, camera movement, or cuts>.
Audio includes <dialogue, ambience, sound effects, or music>.
A working example:
A ceramic artist finishes a pale blue cup in a studio at dawn, lifts it from the wheel, and places it in the center of a wooden shelf.
Soft morning light enters through the window. The wet clay has a delicate sheen, and the workbench remains tidy.
Begin with a medium shot of the wheel-throwing process, slowly push in toward the cup’s surface texture, then cut to a frontal view of the shelf.
Retain the low hum of the pottery wheel, the friction of clay, and subtle indoor ambience.
Drop any component you do not need. Generation parameters like resolution, aspect ratio, and duration do not belong in the prompt. Those are set on the generation side.
Reference materials: quantity and limits
Seedance 2.5 can combine up to 50 reference materials in 1 generation. Each type has hard input limits and a recommended range. The recommended ranges exist for stability, not because the model refuses more.
Images: up to 30, each no larger than 4K. Prefer 1 to 8 distinct subjects across your subject-reference images.
Videos: up to 10, combined duration no more than 30 seconds. Prefer 1 to 5 distinct subjects and 5 to 10 seconds per subject video.
Audio: up to 10 clips, combined duration no more than 30 seconds. Keep only dialogue, voice characteristics, ambience, or music that is directly relevant to the task.
Video editing: a source video plus reference images. Prefer a source video under 20 seconds and 1 to 5 reference images.
You can push past the recommendations. 9 to 12 subjects across images, 6 to 10 subjects in audio or video, 6 to 8 reference images for an edit. Stability drops as material count grows, so expect more regenerations up there.
One structural rule that saves a lot of failed runs: if more than 5 subjects each need multiple views, put each view in its own image. Independent view images are more stable than collaging several views into 1 grid.
Define every material’s role
This is the single most important habit in Seedance 2.5 prompting. After uploading references, state exactly what each one contributes, and state what it must not contribute.
The mapping has to live in the prompt text. Do not rely on text labels inside the images. Do not make the model guess which person, prop, or scene a material represents. It will guess wrong at the worst time.
The role template:
@Image 1 defines <subject>’s <appearance, clothing, structure, or material>.
@Video 1 defines <motion, camera movement, or pacing>.
@Audio 1 defines <character or sound type>’s <voice, dialogue, ambience, or music>.
<Subject> completes <primary action or event> in <scene>.
The visuals feature <visual style>, with <camera treatment>.
And in practice:
@Image 1 defines the ceramic artist’s facial features, hairstyle, and dark green apron. Do not use the image background.
@Image 2 defines the wooden workbench, window placement, and morning light of the pottery studio. Do not use the people in the image.
@Video 1 defines the pacing of throwing clay with both hands, lifting the cup, and placing it down. Do not use the person’s identity, clothing, or scene from the video.
The ceramic artist finishes a pale blue cup in the pottery studio at dawn, lifts it from the wheel, and places it in the center of a wooden shelf.
Begin with a medium shot of the wheel-throwing process, then slowly push in toward the cup’s surface texture. Retain the sound of the wheel, the friction of clay, and indoor ambience.
Notice the exclusions. “Do not use the image background.” “Do not use the people in the image.” Every reference carries baggage: backgrounds, bystanders, compositions. If you do not fence it off, it can bleed into the output.
When several images show different views of the same person or product, say so explicitly:
@Image 1 defines the front view of the same folding desk lamp.
@Image 2 defines the left-side structure of the same folding desk lamp.
@Image 3 defines the right-side structure of the same folding desk lamp.
@Image 4 defines the rear structure of the same folding desk lamp.
All four images define one folding desk lamp. The output must contain only one lamp throughout.
That last line matters for product work. Without it, 4 reference angles can become 4 lamps in the shot.
If a reference video already nails the motion, camera, and sequence, state only which attributes to inherit. Do not restate every action, because a rewritten motion description can conflict with the video it is supposed to describe. And if the reference is a blockout, remember it only carries motion and spatial structure. The prompt still has to define the actual subjects, scene, action, and style.
The audio and text syntax
Prompts can be pure natural language. When you need the model to cleanly separate music, sound effects, dialogue, and on-screen text, Seedance 2.5 has dedicated syntax:
Music goes in parentheses: (Soft, rhythmic piano music plays in the background)
Sound effects go in angle brackets: <A bell rings in the distance>
Dialogue goes in curly braces: {Hello, welcome back.}
Subtitles go in corner brackets: 【Chapter One: Departure】
For dialogue in a language other than Chinese, specify the language before the line:
The girl says softly in Japanese: {もう大丈夫です}
If your dialogue text is English but the model keeps delivering it in the wrong language, or you need a specific regional variety, reinforce it with this formula: dialogue language + regional variety or accent + delivery style + speaker + the line.
Dialogue language: American English. The girl says in natural, conversational American English: {I thought you weren’t coming.}
Dialogue language: authentic Los Angeles English. The young man says in natural Los Angeles vernacular: {No way, you actually made it.}
This is the fix for the most common lip-sync complaint in ad work. The words were right, the language came out wrong, and the prompt never pinned it down.
Multi-reference creation: the material selection system
50 references is not an invitation to cram everything into 1 sentence. The point of multi-reference prompting is to define the relationships among characters, props, scenes, actions, and audio, so the model can select the right materials for each scene.
The workflow runs in a fixed order: define each material’s role, map subjects, group by type, create subject profiles, then select references per scene.
Step 1: name and map each subject individually
Bind each person, product, and prop to its reference separately:
<Character A> corresponds to @Image 1. Use only the appearance, hairstyle, and clothing.
<Character B> corresponds to @Image 2. Use only the appearance, hairstyle, and clothing.
<Prop A> corresponds to @Image 3. Use only the structure, material, and color.
<Scene A> references @Image 4. Use only the spatial layout, architecture, and lighting. Do not use the people in the image.
Do not write “@Images 1 through 4 define four characters respectively.” That sentence contains zero information about which image is which character, and the model will assign them however it likes.
Step 2: group materials by type
[Characters]
<Conservator> corresponds to @Image 1. Use only the appearance, hairstyle, and clothing.
<Registrar> corresponds to @Image 2. Use only the appearance, hairstyle, and clothing.
<Exhibition Installer> corresponds to @Image 3. Use only the appearance, hairstyle, and clothing.
<Guide> corresponds to @Image 4. Use only the appearance, hairstyle, and clothing.
Do not interchange the four characters’ appearances, clothing, actions, positions, or dialogue.
[Props]
<Sample Case> corresponds to @Image 5 and belongs only to <Conservator>.
<Record Board> corresponds to @Image 6 and belongs only to <Registrar>.
[Scenes]
<Conservation Lab> references @Image 7. Use only the space, materials, and lighting.
<Gallery> references @Image 8. Use only the space, materials, and lighting.
[Motion and Audio]
@Video 1 defines the motion of <Conservator> opening <Sample Case>. Do not use the person or scene from the video.
@Audio 1 defines <Guide>’s voice and specified dialogue.
Note the ownership lines on props. “Belongs only to” prevents the model from handing your hero product to the wrong character mid-video.
Step 3: create a profile for important subjects
When 1 character uses several references across multiple scenes, centralize everything into a profile:
[Subject Profile: Conservator]
Appearance and clothing: @Image 1.
Fixed prop: <Sample Case> from @Image 5.
Locations: <Conservation Lab> and <Gallery>.
Motion references: the case-opening motion from @Video 1 and the sample-placement motion from @Video 2.
Do not use: other characters’ clothing. Do not give this character <Record Board> or guide equipment.
Step 4: select references by scene
Scene 1 | Inspection in the Conservation Lab
Use: <Conservator>, <Sample Case>, <Conservation Lab>, and the case-opening motion from @Video 1.
Event: <Conservator> opens <Sample Case> at the workbench and inspects the sample inside.
End state: <Conservator> remains on the inner side of the workbench. <Sample Case> stays beside the conservator’s right hand, which is on the left side of the frame.
Scene 2 | Registration in the Gallery
Use: <Registrar>, <Record Board>, and <Gallery>.
Event: <Registrar> checks the number on <Record Board> beside the display case.
End state: <Registrar> still holds <Record Board> with both hands. No other character enters the display-case area.
The goal of multi-reference creation is to help the model pick the correct materials for the current scene. It is not to make every material appear at the same time. Load 15 references without scene selection and the model will try to be polite and use all of them at once. That is where the chaos comes from.
30-second videos: stages and end states
When a video contains several events, split the story into consecutive stages. Each stage gets exactly 1 primary state change, plus an explicit description of what should be directly visible when the stage ends.
The end state is the trick. Most long-video failures are not bad generation. They are the model finishing a beat in a state you never specified, then building the next beat on top of the wrong state.
The template:
[Generation Goal]
Generate a <video type>. The central subject is <subject>, and the primary event is <story summary>.
[Stage 1]
Initial state: <initial state of characters, props, and scene>.
Primary event: <one primary action or event>.
End state: <character positions, prop ownership, or visible scene state>.
[Stage 2]
Continue from the previous stage: <state that must remain unchanged>.
Primary event: <one primary action or event>.
End state: <observable state>.
[Stage 3]
Primary event: <closing event>.
End state: <final visible state>.
[Maintain Consistency]
Keep <character identity, number of characters, clothing, prop ownership, spatial direction, and audio relationships> consistent.
A full example:
[Generation Goal]
Generate an instructional video showing a flower shop’s order-packing process. <Florist> and <Store Assistant> arrange, wrap, and hand off a bouquet together.
[Stage 1]
Initial state: <Florist> stands behind the workbench. Loose flower stems, scissors, and wrapping paper lie on the tabletop.
Primary event: <Florist> arranges the stems and trims them to length.
End state: <Florist> holds the bouquet in the left hand, and the scissors are back on the right side of the workbench.
[Stage 2]
Continue from the previous stage: both characters retain the same identities and clothing, and <Florist> still holds the bouquet.
Primary event: <Store Assistant> unfolds the wrapping paper. <Florist> places the bouquet inside and ties it with a green ribbon.
End state: the wrapped bouquet lies flat in the center of the workbench, with the ribbon bow facing the camera.
[Stage 3]
Primary event: <Store Assistant> picks up the bouquet and places it on the pickup shelf.
End state: the bouquet is centered on the pickup shelf, and both characters stand behind the workbench inspecting the finished order.
[Maintain Consistency]
Keep <Florist> and <Store Assistant>’s identities, clothing, workbench orientation, scissors position, and bouquet ownership consistent.
Timestamps and pacing
Stages are the default. Reach for 1-second precision only when you need to control a critical handoff, an entrance or exit, a transition, or a specific beat.
3 patterns are available. Time ranges allocate pacing: “0-3 seconds… 3-7 seconds… 7-12 seconds…” Exact time points mark a single key event: “At 5 seconds, the camera whip-pans rapidly to the left and completes the transition.” Relative timing describes a delay: “Three seconds after the character presses the button, the room lights gradually turn off.”
A timestamped sequence in practice:
0-5 seconds: Show an empty wooden display table. A hand places a white ceramic plate on it. End state: the hand has left the frame, and only the white plate remains in the center of the table.
5-10 seconds: Remove the white plate, then place a clear glass on the table. End state: only the clear glass remains in the center of the table.
10-15 seconds: Remove the clear glass, then place a green ceramic vase on the table. End state: only the green vase remains in the center of the table.
Keep time ranges consecutive and non-overlapping. And understand what they are: a time budget for an event, not a frame-accurate edit point. Actions can land slightly before or after a boundary. Too little content in a range gives the model freedom to improvise. Too much causes excessive cutting or dropped events. Do not use timestamps to demand frequencies, like completing 3 actions in 1 second. That is not what they do.
Parameter locks on editing, first/last frame, and extension
3 task types automatically lock generation parameters based on the input, and you cannot override the locks from the generation page or the API.
Video editing preserves the input video’s aspect ratio and approximately its duration. Neither can be set separately. Input-frame processing can introduce a difference of up to about 0.3 seconds, usually from transition-frame handling.
First-frame and first-and-last-frame generation locks the aspect ratio to the first image. Duration stays settable. Make sure the first and last images share the same aspect ratio, or the last frame gets stretched.
Video extension locks the input video’s aspect ratio. Extension duration stays settable.
Know these before you build the workflow, not after the render.
Video editing: master, scope, preserve
Seedance 2.5 edits existing video. The discipline is the same every time: declare the source video as the sole editing master, then define the edit target, the edit scope, the target material, and the content to preserve.
The general pattern:
[Edit Goal]
Edit @Video 1. Within <the entire video or a specific time range>, <add, remove, replace, or adjust> <visual object, region, or audio category>.
[Source Video Role]
@Video 1 is the sole editing master. It defines <characters, scene, actions, composition, camera movement, occlusion relationships, audio, and event order>.
[Target Material Role]
@Image 1 or @Audio 1 defines <specified attributes of the target object or sound>.
[Edit Scope]
Modify only <object, region, time range, or audio category>.
[Content to Preserve]
Keep <visual content, motion, audio, and timing relationships that must not change> from @Video 1.
A scoped lighting change:
[Edit Goal]
Edit @Video 1. Only from 4-7 seconds, change the cool blue light on the right wall to warm orange light.
[Source Video Role]
@Video 1 is the sole editing master. It defines the character, room layout, actions, composition, camera movement, audio, and event order.
[Edit Scope]
Change only the light color on the right wall and the area it illuminates. Allow the character’s skin tone to respond naturally to the environmental light.
[Content to Preserve]
Keep the character’s identity, clothing, expression, position, motion, room structure, camera movement, dialogue, and ambience from @Video 1.
Subject replacement
This is the edit that matters most for ad production. Take a working video and swap the product. The key addition is timeline inheritance: the new object has to inherit every appearance, motion, occlusion, and exit of the original, including timing, path, and speed changes.
[Edit Goal]
Edit @Video 1. Replace only the yellow folding desk lamp with the white folding desk lamp in @Image 1.
[Source Video Role]
@Video 1 is the sole editing master. It defines the desk, books, hand movements, camera position, camera movement, occlusion relationships, and event order.
[Target Reference Role]
@Image 1 defines only the white folding desk lamp’s appearance, structure, and material. Do not use the image’s background, composition, or other objects.
[Edit Scope]
Keep exactly one white folding desk lamp throughout the video. Replace only the original yellow folding desk lamp. Do not modify the books, desk, hands, or background.
[Timeline Inheritance]
The white folding desk lamp inherits every appearance, lamp-arm rotation, hand occlusion, and exit of the original yellow folding desk lamp, including timing, path, and speed changes.
Except for the object or area explicitly modified above, keep all other people, props, scene content, camera movements, cuts, and event order from @Video 1 unchanged.
Note the quantity line. “Keep exactly one white folding desk lamp throughout the video.” State the count, always.
Background replacement
Same discipline, aimed at everything outside the subject’s silhouette:
@Video 1 is the sole editing master. It defines the people, actions, composition, camera treatment, and event order.
@Image 1 provides only the spatial layout, depth of field, ambient color, and lighting direction of a daylit glass greenhouse. Do not use the people in the image.
Replace only the light gray background outside the person’s silhouette in @Video 1 with the daylit glass greenhouse from @Image 1.
Keep the person’s identity, facial features, hairstyle, clothing, expression, position, size, and arm-raising motion from @Video 1.
Except for the object or area explicitly modified above, keep all other people, props, scene content, camera movements, cuts, and event order from @Video 1 unchanged.
Audio editing
Dialogue, language, voice, music, and sound effects can each be edited separately. Name the speaker or sound category, the change, and everything that must stay untouched:
Edit @Video 1. Remove only the original background music. Keep the character dialogue, lip sync, ambience, and action sound effects; preserve the visuals, camera treatment, and editing rhythm from @Video 1.
Edit @Video 1. Change <Presenter>’s spoken language to natural American English while preserving the dialogue content and speaking times. Keep all other character voices, background music, ambience, and visuals from @Video 1.
That second one is localization in a single prompt. Same video, same presenter, same timing, new language.
Video extension: align the boundary frame first
Extension generates content beyond the edge of a source video, forward after the last frame or backward before the first frame. The whole game is the boundary frame. Describe its state before you describe anything new.
Forward extension
Describe the continuous state of the source video’s last frame, then what happens next:
@Video 1 is the source video to extend forward.
Extend @Video 1 forward. The first frame of the extended segment directly continues from the last frame of @Video 1. Maintain the same locked-off medium shot, the orange paper airplane’s position and orientation, the classroom-window background, the afternoon lighting, and its movement toward the right side of the frame.
Then, the orange paper airplane continues gliding toward the right and exits the frame while the white curtain beside the window sways slightly. Keep the camera and classroom background in the state established by the source video’s last frame.
Extensions can take additional references for characters, props, or audio. Define every added material’s role first, and make it explicit that new materials must not override the source video’s last-frame control over the opening image:
@Image 1 defines <Gardener>’s facial features.
@Image 2 defines <Gardener>’s light green work apron.
@Image 3 defines <Wicker Flower Basket>’s structure and material.
@Video 1 is the source video to extend forward.
Extend @Video 1 forward. The first frame of the extended segment directly continues from the last frame of @Video 1. Maintain the greenhouse workbench, <Gardener>’s position, and <Wicker Flower Basket>’s position.
Then, <Gardener> lifts <Wicker Flower Basket> with both hands and places it on the middle shelf of the wooden rack behind them.
Throughout the extension, maintain continuity in <Gardener>’s face, apron, greenhouse layout, and camera direction.
Add the instance line to every extension: keep each subject as the same continuous instance throughout, do not duplicate or split it, and keep the person’s appearance or the object’s part count stable.
Backward extension
Backward is trickier. Describe what happens before the source video begins, then define the source video’s first frame as the explicit end state of the extended segment.
Do not write only “then connect to the source video.” That phrasing lets characters or effects that belong to the later footage leak in too early, or lets the image keep changing after it reaches the target state.
@Video 1 is the source video to extend backward.
Extend @Video 1 backward. Before the source video begins, show an empty establishing shot of the same glass greenhouse. Morning mist drifts slowly near the floor, the overhead shade rises gradually, and no people are present yet.
The last frame of the extended segment naturally connects to the first frame of @Video 1. Match the greenhouse’s central aisle, planting tables on both sides, glass frame, soft morning light, and locked-off wide composition. At the end, the shade is fully raised, the aisle is empty, and the leaves still sway slightly.
With additional references, also state which materials are used in the backward segment and which should appear only after the source video begins:
@Image 1 defines <Curator>’s facial features.
@Image 2 defines <Curator>’s dark blue work jacket.
@Image 3 defines <Wooden Display Case>’s structure and material.
@Image 4 defines the gray workwear of two <Exhibition Assistants>.
@Image 5 defines <Exhibition Preparation Room>’s space and lighting.
@Video 1 is the source video to extend backward.
Extend @Video 1 backward. Before the source video begins, <Curator> walks to the workbench, picks up the closed <Wooden Display Case>, and opens its lid.
The last frame of the extended segment naturally connects to the first frame of @Video 1. <Curator> stands in the center of the frame, holding the open <Wooden Display Case> with both hands. The two <Exhibition Assistants> stand behind the curator, one on each side. Match the vertical frontal medium shot, workbench position, preparation-room background, and morning light from the left established by @Video 1’s first frame.
Boundary frames connect naturally at a visual level. They will not be pixel-identical. When reviewing, inspect both sides of the boundary and the full extended segment, not just the seam.
First and last frames, plus extra references
In multimodal reference mode you can declare frames directly. State in the first line that @Image 1 is the first frame and @Image 2 is the last frame. No separate first/last-frame mode needed.
The system locks the output aspect ratio to the first image. Duration is set on the generation side. Keep the first and last images at the same aspect ratio or the last frame stretches. Additional images can still define characters, props, scenes, and materials on top of the frame anchors.
@Image 1 is the first frame. It defines the opening composition, character positions, poses, tabletop prop states, perfume-workshop scene, and camera direction.
@Image 2 is the last frame. It defines the ending composition, character positions, poses, tabletop prop states, perfume-workshop scene, and camera direction.
@Image 3 defines <Perfumer>’s face, hairstyle, and dark green apron. Do not change the first-frame composition defined by @Image 1 or the last-frame composition defined by @Image 2.
@Image 4 defines <Glass Perfume Bottle>’s shape, material, and label position. Do not change the first-frame composition defined by @Image 1 or the last-frame composition defined by @Image 2.
Starting from the first-frame pose, <Perfumer> picks up a dropper and <Glass Perfume Bottle>, drips amber fragrance oil into the bottle, swirls it gently, closes the stopper, places the finished bottle in the center of the table, and naturally reaches the last frame defined by @Image 2.
Between the first and last frames, maintain continuity in <Perfumer>’s identity and clothing, bottle count and structure, wooden-table layout, warm side lighting, and camera direction.
Describe each anchor image separately. Do not compress them into “@Images 1 and 2 are the first and last frames.” And keep the supporting references in their lane: they supplement their specified attributes only, they never replace the frame compositions.
Multi-keyframe sequences
When separate images define different stages of a process, open with “Use @Image 1 through @Image N as keyframes in this order,” then describe the key state each image represents.
Independent keyframe images align more easily than several frames combined into 1 grid. And know what keyframes do: they control stage order and key states. They do not reproduce every frame exactly.
Use @Image 1 through @Image 4 as keyframes in this order.
@Image 1 is the first frame. It shows an orange paper airplane resting on the left side of a classroom desk, pointed toward the right, in a locked-off medium shot.
@Image 2 defines the second keyframe: one hand lifts the same orange paper airplane from the desk without changing its direction.
@Image 3 defines the third keyframe: the same orange paper airplane passes the window while the curtain moves slightly to the right.
@Image 4 is the last frame. It shows the same orange paper airplane resting on the middle shelf of the bookcase on the right, still pointed toward the right.
The video passes through the states defined by @Image 1, @Image 2, @Image 3, and @Image 4 in order. Keep flight direction and speed continuous between stages.
Maintain the paper airplane’s orange material, size, and folds, as well as the classroom layout, afternoon side lighting, and camera axis.
Storyboard grids
A storyboard grid communicates the overall story, shot order, and approximate compositions. It is not a strict reproduction target for every panel detail.
Keep it to 15 panels or fewer, in clean line art or simple diagrams, with minimal text labels. In the prompt, state the reading order, then describe each panel’s subject action, shot size or camera movement, the final visual style, and the audio.
@Image 1 provides a four-panel pottery-making storyboard for shot order and approximate composition. Read it left to right, top to bottom. Do not use the storyboard’s line-art style or text labels.
@Image 2 defines <Ceramic Artist>’s face, short hair, and dark gray apron.
@Image 3 defines <Blue-Glazed Cup>’s proportions, glaze color, and curved handle.
Shot 1: a wide shot establishes a quiet pottery studio with <Ceramic Artist> seated at the wheel.
Shot 2: a side medium shot shows both hands shaping the rotating clay as the cup body takes form.
Shot 3: a close-up shows fingers refining the rim and handle joint while slip moves slowly over the fingertips.
Shot 4: a medium close-up shows the fired <Blue-Glazed Cup> placed on a wooden shelf as <Ceramic Artist> withdraws both hands.
Use a realistic documentary look. Retain the wheel’s rotation, wet-clay friction, and studio ambience.
The exclusion in line 1 does the heavy lifting. Without it, your finished ad comes out in line art.
Blockout references: coarse and fine
Blockouts are pre-visualization inputs, and they come in 2 categories that need different prompts. First identify which one you have.
A coarse blockout is simple geometry that previews temporal information: action, paths, blocking, camera movement, cuts, lighting changes, and sound rhythm. Use it when the shapes have clear relationships and the action sequence is complete.
A fine blockout already contains complete structures. Use it to re-render materials, colors, characters, scenes, and overall style while keeping the structure and motion. Keep the model clean: remove path lines, coordinate axes, controllers, camera frustums, and other production markers before uploading.
Coarse blockouts
Map each geometric object separately to its final subject or prop, and state which temporal information to inherit: paths and blocking, camera position and movement, lighting direction and timing, cut positions with the composition on each side, and whether to inherit any audio. Prefer simple geometry. Arms, wings, and other appendages only help when the action sequence is complete; otherwise they cause stiff motion or structural misreads.
@Video 1 is a coarse blockout reference. It provides only the character’s walking path, cart direction, locked-off camera, one push-in, and two cuts. Do not use its gray geometry or empty scene.
The tall cylinder in @Video 1 corresponds to <Guide>.
The rectangular block in @Video 1 corresponds to <Mobile Display Cart>.
@Image 1 defines <Guide>’s face, blue uniform, and name badge.
@Image 2 defines <Mobile Display Cart>’s white metal frame and clear cover.
@Image 3 defines the technology gallery’s curved walls, gray floor, and overhead strip lights.
<Guide> pushes <Mobile Display Cart> along the curved wall, stops in front of the central display, and opens the clear cover.
Keep the walking path, subject blocking, push-in direction, and cut points from @Video 1.
Use a bright, realistic museum-documentary style. Retain footsteps, wheel sounds, and gallery ambience.
Fine blockouts
Preserve the structure, action, and camera treatment. Define what gets re-rendered:
@Video 1 is a fine blockout reference. Preserve the kinetic sculpture’s complete structure, three-ring rotation relationship, pedestal position, orbiting camera movement, and cuts. Do not use the gray materials or empty background.
@Image 1 defines the outer ring’s brushed-brass material.
@Image 2 defines the inner blades’ translucent blue-glass material.
@Image 3 defines a contemporary gallery with white curved walls, a dark gray floor, and soft overhead lighting.
Re-render the ring structure from @Video 1 as a kinetic sculpture made of brass and blue glass, and re-render the scene as a contemporary art gallery.
Keep the structure, rotation rhythm, orbiting camera movement, and cuts from @Video 1. Retain the sculpture’s low mechanical rotation sound and quiet interior ambience.
One-click video
One-click video organizes multiple images, or images plus a style-reference video, into a complete edited video with consistent pacing and packaging. The failure mode is writing “turn these materials into a video” and nothing else.
The structure to follow: material roles, image order, motion amount, editing style, visual treatment, audio.
[Material Roles]
@Image 1 is used for the night-market entrance and opening environment.
@Image 2 is used for <Traveler> walking along the street.
@Image 3 is used for the lantern stall and craft details.
@Image 4 is used for three friends eating together.
@Image 5 is used for the riverside night view and reflections.
@Image 6 is used for the final group photo by the bridge.
@Video 1 is used only for light editing rhythm, hand-drawn stickers, and transition style. Do not use its character identities or locations.
[Arrangement]
Show @Image 1 through @Image 6 in order to form a complete sequence: arrival, street exploration, dinner, riverside walk, and group photo.
Keep the three friends’ appearances and clothing consistent. Do not mix their identities.
[Image Motion]
Use slow push-ins and subtle parallax for environment images. Add only natural blinking, head turns, glass-raising, and slight clothing movement to character images.
Keep stall structure, table position, and bridge railing stable.
[Final Style]
Use an upbeat travel-video rhythm. Connect scenes with natural occlusion and similar colors. Keep hand-drawn stickers at the frame edges.
[Audio]
Retain night-market chatter, light dish sounds, and riverside wind, with upbeat but unobtrusive instrumental music.
If image order matters, state the exact sequence. If the model can arrange freely, say it may organize by theme. And when several characters or products appear, the binding rule from earlier still applies: name and bind each one separately.
Seamless video transitions
A seamless transition generates continuous bridge content between 2 videos. The prompt identifies the before video and the after video, then walks through the trigger action, camera movement, visual transformation, arrival state, and audio transition.
5 transition methods, each with its own thing to specify. A dive or reverse movement needs camera direction, speed change, and when the next scene begins. A character rotation needs the pose, rotation direction, and how clothing or background changes continuously. Foreground occlusion needs the moment the foreground object fills the frame and the composition that follows. An object morph needs the corresponding shapes, materials, and the transformation process. A push, pull, or focus change needs the camera movement, focus target, and the continuous spatial relationship.
@Video 1 is the before-transition clip. Use its rainy night street, red umbrella, slow push-in, and rain sound.
@Video 2 is the after-transition clip. Use its circular gallery skylight, upward camera movement, and quiet interior reverberation.
Keep the people, street, gallery structure, and primary actions in the two original videos stable.
At the end of @Video 1, the red umbrella approaches the camera and gradually fills the entire frame, triggering the transition.
The camera continues moving forward. The umbrella’s circular edge gradually becomes the skylight’s metal ring, and the red fabric transitions into white daylight passing through the skylight.
The transition ends naturally at @Video 2’s upward-looking opening composition, with the camera movement changing smoothly from forward motion to an upward rise.
The rain gradually fades into footsteps reverberating inside the gallery.
The goal is visual and audio continuity. A prompt can ask to preserve the primary content of both source videos, but the generated bridge is not a pixel-identical edit splice.
Emotional direction: write what the camera can see
Emotion words like “tense,” “warm,” or “oppressive” set an overall direction and leave the performance open to interpretation. For stable acting, add cues that are directly visible or audible: eye movement, brow tension, mouth movement, breathing, gaze direction, hand movement.
You do not need every facial detail. For a single emotional transition, 2 to 4 clear cues are enough. Save event-triggered staging for emotions that change several times.
Single transition, default structure:
The overall emotion shifts from <starting emotion> to <ending emotion>.
After <triggering event>, <subject> first shows <immediate observable reaction>.
Then, <eyes, brows, mouth, breathing, gaze, or hand movement> gradually <changes>.
Finally, <subject> expresses <target emotion> through <restrained or explicit outward behavior>.
Multi-stage emotion, progressed through triggering events:
When <subject> hears or sees <first triggering event>, <first observable reaction>.
When <second triggering event> occurs, <change in expression, gaze, or breathing>.
After confirming <critical information>, the emotion that <subject> tries to restrain or conceal gradually becomes visible through <observable behavior>.
Finally, <subject’s final action, expression, or manner of speaking>.
What it looks like written out:
Applause marking the end of the performance comes from behind the stage. The young actor’s fingers suddenly stop on the program, the gaze turns slowly toward the curtain, and the shoulders remain tense.
After confirming that the curtain call is over, the actor exhales softly. The shoulders gradually relax, a restrained smile appears, and the eyes slowly well with tears, but the actor never turns to leave.
No emotion word appears in that example. Every line is something a camera can record. That is the standard.
Camera language
Basic terms go straight into the prompt. Shot sizes: extreme wide shot, wide shot, medium shot, close-up, extreme close-up. Movements: push in, pull out, pan, lateral move, follow shot, orbit, dive, dolly out, tilt up, handheld shake. Positions and viewpoints: low angle, overhead view, first-person view.
Popular techniques also work directly, but when the frame has several subjects, still state which subject the camera follows, where the movement begins, and where it ends. Each technique has a spec:
A one-take shot needs the subjects, spaces, and events the continuous camera passes through in order. A dolly zoom needs the subject size to preserve and whether the background appears to move closer or farther. An aerial view needs viewing height, movement direction, and the environmental area to reveal. FPV needs the flight or traversal path, speed, and turns. Bullet time needs the action to freeze or slow and the camera’s orbit direction. Handheld needs the subject being followed and the amount of shake. A bounce speed ramp needs where the action accelerates, decelerates, or rebounds, and its final resting state.
For niche terms, terms the industry uses inconsistently, or terms the model may not know, keep the term and translate it into an observable visual change. The formula: cinematography term + target subject + visual change + foreground/background relationship + direction or speed.
Rack focus: shift focus smoothly from the leaves in the foreground to the person in the background. The leaves gradually blur while the person’s face changes from soft to sharp.
More of the same pattern:
Shallow-depth-of-field portrait: keep <Pastry Chef>’s eyes and face sharp while the glass jars and lights in the background become soft, circular bokeh.
Tracking shot: move horizontally at the same speed as <Skateboarder>, keeping the subject sharp while the roadside wall forms horizontal motion blur from right to left.
Golden hour: warm, low-angle sunlight enters from behind and to the left of <Hiker>, casting long shadows across the mountain ridge.
Natural vignette: darken the four corners gradually while keeping the brightness and skin tone of <Pianist> in the center natural, without a black border.
Whip-pan transition: at 5 seconds, move the camera rapidly to the left. Cut when the foreground bookshelf fully covers the frame, then continue moving left at a similar speed in the next scene.
Aperture, focal length, and shutter values can go in the prompt, but the visible result is usually clearer to the model than a number.
The pre-submission checklist
Run this before every generation. It catches the failures before they cost credits.
Does the prompt clearly state the subject and primary action or event?
Does every reference material state what to use and what not to use?
Is every distinct character, product, and prop named and bound to a reference?
Are references selected by scene instead of being required to appear all at once?
Does each stage of a long video contain only 1 primary change and a clear end state?
Do the number of characters, clothing, prop ownership, and spatial relationships stay consistent?
For video editing, does the prompt define the sole editing master, edit scope, target quantity, and content to preserve?
Are abstract emotions and cinematography terms paired with directly visible or audible cues?
Are first/last frames and keyframes assigned 1 role per image, and do the first and last images share an aspect ratio?
Does the storyboard state which structure to inherit? For blockouts, did you identify coarse versus fine and specify the temporal, structural, material, and style information to inherit?
Do editing, first/last-frame, and extension tasks respect their automatically locked aspect-ratio and duration rules?
For extension, did you check the boundary image, motion trend, and audio continuity?
For one-click video, does the prompt define material roles, image order, motion amount, editing style, and audio?
For transitions, does the prompt define both videos’ roles, the trigger action, the transition process, and the arrival state?
What prompting cannot fix
A few limits worth internalizing, because no amount of prompt engineering moves them.
Timestamps allocate time to events. They are not frame-accurate edit points.
Editing prompts raise the probability that critical events align with the source video. They cannot guarantee frame-by-frame overlap.
Multi-reference creation selects and combines the correct materials. It does not make every material appear at once.
For subtitles, formulas, signs, product specifications, or frame-level timing that must be completely accurate, combine prepared reference materials, generation, and post-production. Do not lean on generation alone.
Editing locks the input’s aspect ratio and approximate duration, with up to about 0.3 seconds of drift. First/last-frame generation locks the ratio to the first image. Extension locks the input’s ratio, and the extended segment’s volume may differ slightly from the source.
Seamless transitions aim for continuity, not pixel-identical preservation of both sources.
Bottom line
Seedance 2.5 rewards prompts that read like production briefs. Bind every material to a role, fence off what each reference must not contribute, give every stage 1 change and an explicit end state, and translate emotion and camera language into things a lens can record.
Seedance 2.5 is available now on Maxfusion AI, and you can run everything in this guide from a chat window through the Maxfusion MCP.