1. Summary
Seedance 2.5 (official model id dreamina-seedance-2-5-260628) is ByteDance’s mainline video generation model, launched 31 July 2026 on Jimeng AI and Doubao Pro, with API access live on BytePlus ModelArk. It generates a single video of up to 30 seconds and accepts up to 50 reference assets, 30 images, 10 video clips and 10 audio clips, in one request. It adds professional editing and extension controls, timestamp-level targeted edits, and native generation in more than ten languages. Output runs at 480p, 720p or 1080p with 10-bit colour at 24 fps. On Kittl, the model is exposed with up to 20 reference images, 720p output, and no audio or video reference inputs.
In Kittl hands-on testing the defining trait is fidelity to what you give it: a supplied start frame is held pixel-faithfully, and printed lettering stays legible and stable as the shot moves. The trade is that it reinterprets explicit shot lists rather than executing them, and it drops the 4K tier that Seedance 2.0 still offers.
Top strengths
- Camera Control (4.67/5.0) – locked-off, orbit, dolly and handheld moves execute as written; the highest-scoring dimension in testing.
- Audio Quality (4.6/5.0) – ambience, effects and dialogue arrive locked to frame in the same pass, with lip-sync generated inline rather than post-hoc.
- Temporal Consistency (4.55/5.0) – subjects, outfits and product artwork hold across a shot; the locked/unlocked task model means editing, first-and-last-frame and extension tasks inherit the source’s aspect ratio and duration instead of re-deriving them.
Top gaps
- Speed (3.0/5.0) – the slowest dimension by a wide margin; a short clip takes minutes, not seconds.
- Motion Quality (4.3/5.0) – human motion reads slightly stiff, and secondary objects can vanish or freeze when they should react.
- Resolution ceiling – the model tops out at 1080p while Seedance 2.0 reaches 4K at 10-bit 2. On Kittl the output is 720p; Seedance 2.0 Full HD and Seedance 2.0 4K cover the higher tiers on the same platform.
Best-fit use cases
- Long single-take storytelling – 30 seconds in one pass, with multi-round extension beyond that 1.
- Reference-heavy production – character, set and palette locked from many aligned assets rather than inferred from one hero frame.
- Product and packaging video – printed lettering survives movement and lighting change, and an existing product photo animates without being repainted.
2. How We Tested
Test environment
- Platform: Kittl AI generator, scored on the Kittl model comparison wall
- Date: August 2026
- Model version tested: Seedance 2.5 via Kittl
- Output format: 9:16, six seconds, audio on where the brief called for it
- Both text-to-video and image-to-video were tested; Seedance 2.5 supports text-to-video natively 2
Prompt set
Briefs were written to stress one named capability each, then scored on every dimension the finished clip could actually evidence. Coverage spans prompt adherence, text in video, camera and composition, face persistence, temporal consistency, motion quality, physics realism, audio quality, very short prompts, and image-to-video start-frame fidelity.
Scoring scale
Each capability is scored 1 – 5 against a definition. Scores are absolute – a 5 means the model executes the definition without meaningful failure; a 1 means execution is unusable for that capability.
- 5 = Executes the definition without meaningful failure
- 4 = Executes reliably with minor inconsistencies
- 3 = Execution is inconsistent or partial
- 2 = Execution frequently fails or requires significant workarounds
- 1 = Unusable for this capability
Capability definitions

3. Capability scores and breakdown
Model overall score: 4.3/5.0

3.1 Prompt adherence: Score 4.5
What this measures: Whether the clip matches the described camera, action, subject and mood.
What works:
- Named subjects, wardrobe and setting land accurately from a written brief.
- Deliberately impossible directions are followed rather than corrected back to realism.
- Stronger instruction following is an official claim for this generation 5.
What breaks:
- Explicit shot lists are reinterpreted rather than executed – multi-shot briefs come back merged into fewer cuts.
- Named scene counts are not honoured.
- Very short prompts invite invention: the model adds agents, copy and effects nobody asked for.
Example:
See the prompt here
Create an exactly 10-second photorealistic 9:16 waterproof-makeup commercial in one continuous high-speed beauty shot. A blonde female model faces the camera against a sleek metallic silver-gray studio background. She has striking bright-blue irises and long blonde hair styled in a polished high ponytail with a smooth crown and controlled flyaways. She wears precise black winged eyeliner, vivid cherry-red lipstick, softly defined brows, matte foundation, a silver sleeveless top, and small diamond stud earrings.
Locked throughout the entire video: the model’s identity, facial structure, skin tone, bright-blue irises, blonde hair, high ponytail, cherry-red lipstick, eyeliner shape, foundation, clothing, earrings, silver-gray background, lighting direction, camera position, and framing.
0:00–0:03 Show a clean, dry close-up of her face. She looks directly into the camera, clearly revealing her bright-blue eyes and glossy cherry-red lips. Cool studio lighting creates soft silver highlights around her face. No water appears yet.
0:03–0:05 Unlock only the water action. At exactly 0:03, a controlled splash enters from the left and strikes the left side of her face. Show realistic impact, separated droplets, and water spreading across her cheek. She makes one natural blink without turning away.
0:05–0:07 At exactly 0:05, a second controlled splash enters from the right and strikes the right side of her face. Preserve all earlier droplets as the new water crosses her skin. Her eyeliner, foundation, brows, and cherry-red lipstick remain completely intact.
0:07–0:10 Lock the water action again. No additional splashes appear. The camera slowly pushes closer as droplets slide down her face under gravity. She opens her bright-blue eyes and gives a confident smile, revealing flawless waterproof makeup. Use dramatic slow motion only during each impact. Do not change her iris color, lipstick color, blonde hair, high ponytail, facial features, makeup design, outfit, background, or lighting. Do not add extra splashes, reset existing droplets, introduce text, or switch camera angles.
3.2 Text rendering accuracy: Score 4.5
What this measures: Legibility of static signs, neon, printed lettering and scrolling text across frames.
What works:
- Printed lettering and wordmarks keep correct spelling with no letter drift as the shot moves.
- Type stays anchored through liquid distortion, rotation and lighting change.
- Backdrop typography holds position and spelling while the product moves in front of it.
- Native generation in more than ten languages is an official capability of this generation 5.
What breaks:
- Light-coloured lettering can lose contrast when the scene brightens behind it, effectively un-revealing itself.
- The model will invent copy that was never in the brief, adding a tagline of its own.
- Type can pick up an unexplained animation.
Example:
See the prompt here
Create an exactly 12-second photorealistic 9:16 cinematic video in one continuous macro shot.A vintage leather-bound book lies open on an antique wooden desk. Its yellowed pages have worn edges, delicate creases, and visible paper fibres. A handwritten love poem appears in dark sepia fountain-pen ink:
“I found your memory
between these pages.
The years moved forward,
but my heart stayed.”
Behind the open book, three vintage book spines display the titles “Old Letters,” “Lost Summers,” and “Remember Me” in sharp gold lettering. Preserve the exact spelling, letterforms, spacing, and placement throughout the video.
0:00–0:04 Begin with a locked overhead view. Warm afternoon light falls across the page while the entire poem remains clear and readable. A fingertip enters the frame and pauses beneath the first line.
0:04–0:08 The camera slowly dollies closer as the fingertip gently traces beneath each handwritten line. The touch should feel careful and emotional, like revisiting a memory. The ink remains firmly anchored to the textured paper without drifting, changing, or moving independently.
0:08–0:12. The fingertip softly strokes the final words, “my heart stayed.” A lace curtain casts a moving shadow across the page while the camera shifts into a shallow three-quarter angle. Preserve strong contrast and keep every word readable through the lighting change and natural page curvature. The gold book-spine titles remain stationary and correctly spelled in the background.
Mood: intimate, nostalgic, tender, and slightly melancholic. Use warm film grain, soft candle ambience, and slow, deliberate movement. Show only the provided poem and three book titles. Do not invent an author name, date, tagline, signature, additional writing, subtitles, glowing letters, disappearing text, animated handwriting, or changing words.
3.3 Camera & composition: Score 4.6
What this measures: Pans, dollies, orbits and tracking shots executing the move that was asked for.
What works:
- Locked-off tripod framing is genuinely locked – no drift, no creeping zoom.
- Orbits and dolly push-ins run smoothly at a constant rate without jitter.
- Highest-scoring dimension in testing, and consistent across both text-to-video and image-to-video.
What breaks:
- Combined moves (pull back while orbiting) can resolve without following the stated path.
- Lighting direction can flip relative to the brief, with the key arriving from the opposite side.
Example:
See the prompt here
Create an exactly 12-second photorealistic high-fashion beauty shoot in one continuous take. A female model poses against a seamless burgundy studio backdrop. She wears a structured black satin dress, sculptural gold earrings, glossy burgundy lipstick, soft smoky eye makeup, and a sleek low bun. A large softbox positioned on her left is the only key light. Keep highlights on the left side of her face and shadows on her right throughout the video.
0:00–0:03. Locked frontal shot. Begin with a symmetrical chest-up portrait. The model looks directly into the lens with a confident, neutral expression. The tripod framing remains completely locked, with no drift, shake, pan, tilt, or creeping zoom.
0:03–0:06. Smooth pan. Pan smoothly from left to right at a constant speed as she slowly turns into a clean side profile. Maintain the same camera distance, height, and focal length. Do not push in or orbit during the pan.
0:06–0:09. Dolly close-up. Stop panning. Move the camera directly forward in a straight line toward her face, revealing realistic skin texture, precise makeup, and reflections on the gold earrings. Do not zoom, orbit, tilt, or drift sideways.
0:09–0:12. Controlled orbit. Stop moving forward. Perform a smooth 60-degree clockwise orbit around the model at a fixed distance while she holds her pose. Reveal her face from profile to a polished three-quarter angle. Do not pull back, push in, or change camera height during the orbit. Keep each camera movement separate and execute the requested path at a constant speed. Preserve the model’s identity, makeup, hairstyle, dress, earrings, body proportions, and backdrop across every angle. The key light must remain physically positioned on her left and must never flip sides. Do not add handheld shake, random zooms, extra models, outfit changes, or unrequested camera motion.
3.4 Generation speed: Score 3.0
What this measures: Time from prompt to finished clip.
What works:
- Throughput is predictable, and failed generations are not billed 3.
What breaks:
- A short clip takes minutes, not seconds – the weakest dimension by a clear margin.
- Per-second billing means iteration cost scales with clip length as well as resolution 3.
3.5 First-try quality: Score 4.3
What this measures: How often output is usable without regenerating.
What works:
- Single-shot product and camera briefs land usable on the first pass.
- Start-frame fidelity removes a whole class of re-rolls: the supplied frame is never repainted.
What breaks:
- Multi-shot and dialogue-heavy briefs are the most likely to need another pass.
- Very short prompts are the least predictable, since the model fills the gaps itself.
- Image-to-video output is consistently rated lower overall than text-to-video, so animating an existing asset needs more supervision than the start-frame fidelity alone would suggest.
Example:
See the prompt here
12-second photorealistic 9:16 UGC-style video using the supplied image as the exact first frame. Preserve the mother and son’s faces, ages, skin tones, hairstyles, clothing, body proportions, kitchen, lighting, table arrangement, milk packaging, and all background details. Do not repaint, redesign, replace, or beautify anything from the source image. Use one continuous handheld smartphone shot with no cuts, transitions, angle changes, or additional scenes. Keep the camera at the mother’s eye level with subtle natural hand movement and stable framing.
0:00–0:03 Continue naturally from the supplied frame. The mother pours pure white milk from the original container into her son’s glass. The milk volume inside the container decreases consistently as the glass fills. She smiles and asks: Mother: “Why do you always choose pure milk?”
0:03–0:08 The son looks at the milk, then at his mother. He lifts the glass naturally and replies with clear lip synchronization: Son: “Because it tastes fresh, and you said it helps me grow strong.”
0:08–0:12 The mother laughs softly and responds: Mother: “That is my smart boy.” The son takes one small sip while the mother gently touches his hair. End with both smiling naturally, without looking directly at the camera. Keep the performance warm, casual, and slightly imperfect like a real family video. Use realistic pauses, eye contact, facial expressions, hand movement, milk flow, and synchronized dialogue. Preserve the exact milk container and its readable label throughout. Do not add subtitles, music, narration, new people, extra products, camera cuts, dramatic effects, or unrequested movement.
3.6 Face persistence: Score 4.4
What this measures: Identity staying stable through head turns, expressions and lighting changes.
What works:
- Identity survives head turns and re-framing without morphing.
- Supplying a face as a start frame measurably improves stability over generating one from text.
- Skin texture holds without the plastic smoothing common to the category.
What breaks:
- Hand and arm positions can end in awkward poses.
- Delivery to camera still carries a detectable synthetic quality.
Example:
See the prompt here
Create an exactly 6-second photorealistic image-to-video portrait using the supplied image as the exact start frame and identity reference.Preserve the woman’s facial structure, age, skin tone, green-gray eyes, nose, lips, jawline, ears, natural facial asymmetry, fine lines, pores, brown hair, gold earrings, cream shirt, blue denim apron, and white clay mark on her right cheek. Do not beautify, de-age, smooth, or reshape her face. 0:00–0:02 She looks directly into the camera with the same calm expression as the reference. The camera slowly pushes closer while warm studio light illuminates the left side of her face. 0:02–0:04 She naturally turns her head toward the pottery shelves, moving from a frontal view into a clear three-quarter profile. Her eyes follow first, then her head. Her facial proportions and distinctive features remain identical throughout the turn. 0:04–0:06 A cooler daylight reflection briefly crosses her face as she turns back toward the camera and develops a gentle, genuine smile. Preserve realistic cheek movement, eye creases, pores, fine lines, and subtle skin texture as her expression changes. Use one continuous shot with natural head movement and stable re-framing. Keep the same woman as the only subject. Do not morph her identity, alter the clay mark, change her age, replace facial features, add makeup, create plastic skin, or modify her clothing and background.
3.7 Temporal consistency: Score 4.5
What this measures: Faces, outfits, objects and details staying identical across all frames.
What works:
- Product artwork, garment prints and packaging lettering hold their design across a full shot.
- The locked/unlocked task model keeps editing, first-and-last-frame and extension outputs aligned to the source rather than re-derived.
- Reference capacity of 50 assets is designed specifically to hold identity across a long take 1 – on Kittl this is 20 reference images, still the deepest reference support of any model on the platform.
What breaks:
- Role continuity can slip in multi-scene briefs – the character established as the lead reappeared as a background figure.
- Scene element is not always persistent across a movement.
- Consistency is measurably weaker on image-to-video than on text-to-video, despite the start frame itself being held faithfully.
Example:
See the prompt here
Create an exactly 12-second photorealistic 9:16 morning coffee video in one continuous take with no cuts or scene changes. A young woman with warm brown skin, an oval face, dark-brown eyes, a small beauty mark beneath her left eye, and shoulder-length wavy black hair prepares coffee in a sunlit home kitchen. She wears the same cream ribbed cardigan over a sage-green camisole, small gold hoop earrings, and a thin red bracelet on her right wrist throughout the video. On the counter is one matte burgundy coffee bag labeled “MORRA. MORNING ROAST” in cream serif lettering, beside one ivory ceramic mug decorated with three small blue flowers. Preserve the exact packaging colors, lettering, logo placement, mug shape, and floral pattern in every frame, including when objects are moved, rotated, or partially hidden.
0:00–0:04 The camera begins in a medium shot as she opens the MORRA coffee bag, scoops the grounds, and places them into a glass pour-over dripper. The camera slowly moves closer while keeping her as the clear lead subject.
0:04–0:08 Without cutting, she pours hot water in a controlled circular motion. Coffee drips naturally into the same floral mug. She briefly moves between the camera and the coffee bag, but the bag remains in its original position and reappears unchanged when visible again.
0:08–0:12 She removes the dripper, lifts the same floral mug with her right hand, and smells the fresh coffee while smiling softly. The camera gently arcs around her, ending with her face, outfit, bracelet, coffee bag, and mug clearly visible together. Maintain one identical woman, outfit, hairstyle, accessories, coffee bag, label design, mug, dripper, and kitchen arrangement throughout. Do not swap roles, duplicate the woman, move her into the background, change garment details, distort text, alter product artwork, replace objects, or make scene elements disappear during camera movement.
3.8 Movement quality: Score 4.3
What this measures: Human locomotion, animal motion, fluid dynamics, multi-character choreography.
What works:
- Fluid behaviour is convincing – pouring, splashing and viscous flow read with the right weight.
- Mechanical and object motion is clean, including grab-and-lift actions and rotation on a fixed axis.
- Official framing positions realistic performance and camera movement as a headline improvement of this generation 5.
What breaks:
- Choreographed movement looks odd when the brief calls for expressive dance.
- Secondary objects can drop out mid-shot – a chair disappeared as the subjects walked away from it.
- Viscous liquid can flow at the wrong rate for its material.
Example :
See the prompt here
Create an exactly 12-second photorealistic 9:16 macro product video featuring a blush-pink rose soap bar embossed with a detailed rose emblem. Set it beside a brushed-brass faucet in a cream ceramic sink, surrounded by loose rose petals and a softly draped linen towel.
0:00–0:03 A wet hand lifts the soap from the linen towel. As its weight is removed, the unsupported fabric collapses naturally into softer folds. Water droplets fall from the soap under gravity, striking the sink and creating small splashes that move every nearby rose petal consistently.
0:03–0:08 The hand places the soap beneath the running faucet and slowly rotates it. The water stream hits the bar with believable force, dividing around its curved surface and splashing into the sink. Foam begins as a thin creamy layer, then progressively accumulates into dense bubbles around the fingers and soap. The soap’s edges gradually soften under subtle time compression rather than changing instantly.
0:08–0:10 The faucet is turned off. The water stream narrows, breaks into droplets, and stops naturally. Residual water continues sliding down the soap and hand into the sink. The accumulated foam, wet traces, displaced petals, and droplets remain visible instead of resetting.
0:10–0:12 The hand releases the soap onto a ceramic dish. It lands with believable weight, slightly compressing the foam and shifting every petal touching the dish. Several bubbles wobble and pop on impact while the linen towel finishes settling under gravity. Use realistic water volume, liquid flow, surface tension, foam growth, splash direction, and material weight. Water must originate from the faucet and leave through the drain without appearing or disappearing. Keep background rose stems gently moving from the nearby disturbance. Preserve the same soap shape, blush-pink color, rose emblem, hand, faucet, dish, petals, and towel throughout.
3.9 Physics Realism: Score 4.4
What this measures: Gravity, cloth drape, water splash, hair inertia, cause-and-effect chains.
What works:
- Impacts land with believable weight, and debris scatters on contact.
- Objects respond to a change in support, settling and shifting under gravity.
- Time-compression is understood – melting and accumulation progress rather than merely wobbling.
What breaks:
- Cause-and-effect is incomplete: one object in a group reacts to an impact while identical neighbours stay inert. Residue and traces are not always carried through.
- Background elements that should be in motion can stay frozen from the first frame.
- Liquid volume is not always conserved relative to its container.
Example:
See the prompt here
Create 12-second photorealistic 9:16 cinematic product video in one continuous shot.
A sealed vintage perfume bottle rests on a slightly raised bed of blush roses, peonies, jasmine, and soft green leaves. The faceted glass bottle contains pale amber perfume and has an antique brass collar, a firmly secured crystal stopper, and an ivory label reading “FLEUR No. 8.”
0:00–0:03 A cream silk ribbon supporting one side of the bottle slowly slips away. Without that support, the bottle naturally tilts and begins falling from a low height toward the flowers.
0:03–0:07 The bottle lands sideways on the flower bed with believable weight. The flowers directly beneath it compress first, followed by the surrounding blooms and stems. Nearby petals scatter in different directions, while more distant flowers sway gently from the transferred force. The bottle makes one small, cushioned bounce before rolling only a few centimetres. Inside the sealed bottle, the amber perfume sloshes in the opposite direction of the fall. Keep the liquid volume unchanged.
0:07–0:12 The bottle gradually settles into the compressed flowers. Its movement becomes smaller until it stops completely. The perfume inside continues moving briefly, then slowly becomes still. Bent stems rebound partially, loose petals fall under gravity, and the silk ribbon collapses into natural folds beside the bottle. Keep the bottle completely intact and sealed. Preserve every displaced petal, compressed flower, bent stem, and fabric fold through the final frame. Use realistic gravity, weight transfer, flower compression, liquid inertia, and gradual settling. Do not break the glass, detach the stopper, spill perfume, refill the bottle, freeze surrounding flowers, reverse motion, or make the bottle float.
3.10 Audio quality: Score 4
What this measures: Dialogue lip-sync, sound-effect timing, ambient sound and music sync.
What works:
- Audio and video are generated jointly, so effects and dialogue arrive locked to the frames they belong to 5.
- Lip-sync is produced in the same pass, with no separate voice or sync step.
- Tagged sound design in the prompt lands on the intended beat.
What breaks:
- Spoken delivery retains a mild synthetic quality.
- Sound can be materially wrong for what is on screen.
- Stray noise appears in otherwise clean ambience.
Example:
See the prompt here
UGC bed-jump product intro, vertical 9:16, filmed on a phone. Natural bedroom daylight, phone-propped imperfections, no color grading. This must NOT look like a commercial, it must look like a real person messing around in her own bedroom and filming it for fun.
Reference: @image1 shows the same single woman in three views — front full body, back full body, and a close-up portrait. It is one person, not three. Use only her facial features, hairstyle and jewellery. Completely ignore the clothing and the white studio background from @image1. Only one woman appears in the video.
@image2 defines the oat milk carton only — its packaging shape, proportions, colour blocking, oat illustration and white screw cap. Do not use its background, and do not let it define any other object in the room. There is only ever one carton in the video, and it stays sealed with the cap on throughout.
Character: Deep brown skin with a natural sheen, very short buzzed hair bleached platinum blonde, high cheekbones, full lips, small silver hoop earrings, a thin delicate silver chain. She wears an oversized cream ribbed knit sweater and soft grey shorts, barefoot. She holds the carton in her right hand for the entire video.
Performance: Playful and light at the start, happy and genuinely enthusiastic once she is talking, but never exaggerated and never performing for an audience — this is the energy of someone filming a small everyday moment for herself. Quick grins, eyebrows lifting, loose shoulders, a real small laugh. No mugging, no big gestures, no wide theatrical eyes.
Setting: Her own bedroom in the late afternoon. A low bed with a thick white duvet and four soft pillows piled at the head, placed side-on to the camera, a woven rug, a plant on the windowsill, warm low sun through a half-open curtain. Lived-in, slightly messy.
Camera: The phone is propped on a dresser across the room and never physically moves — the camera position, angle and height are the same for the whole video. The only camera change is a slow digital zoom in the second half, pushing from the wide framing into a medium framing on her upper body and the carton. No panning, no tilting, no handheld drift.
Audio: A warm upbeat indie-pop track with a clear beat, playing under the whole clip. Her landing on the duvet lands exactly on a beat. Retain the real sounds: her bare feet on the rug, the soft heavy thump of her body hitting the duvet on her side, pillows shifting, the liquid sloshing once inside the carton on impact, a small genuine laugh as she lands, and her breath.
Timeline
00:00 – 00:02: Wide framing, the whole bed visible side-on. Only her head and one shoulder lean in from the right edge of frame, the sealed carton held up beside her face in her right hand. She looks straight into the lens, gives a quick grin with her eyebrows lifted, and holds it for a playful beat.
00:02 – 00:03: She pulls back out of frame to the right and disappears completely from view.
00:03 – 00:05: She comes bounding back into frame, plants one bare foot at the near side of the bed and flops sideways onto the duvet in one quick unbraced motion, the oversized sweater rippling. She lands on her left side facing the camera with a bounce, the duvet compressing under her and the pillows shifting, and lets out a small laugh. Her landing hits on the beat. The carton stays gripped in her right hand.
00:05 – 00:06: Still on her left side facing the lens, she settles and props her head up on her left hand, elbow sinking into the duvet, and brings the carton up in front of her chest with her right hand. The slow digital zoom begins here.
00:06 – 00:09: The framing has pushed in to a medium shot on her upper body and the carton. Smiling, slightly out of breath, she turns the front face of the pack squarely toward the lens and says brightly but casually: {Okay, this one. I actually love this one.}
00:09 – 00:11: She rotates the carton a few degrees each way, keeping the front panel facing the lens, and adds with an easy grin: {It’s in my coffee every single morning. Like, every morning.}
00:11 – 00:12: She lets the carton rest down against the duvet in front of her, still propped on her left hand, gives the lens one last small smile, and the shot holds at the medium framing.
Label handling: only the front face of the carton is ever turned toward the lens. The dense paragraph of small body text on the side panel is never rotated toward the camera at any point.
Every object appears only once. Her face, hair, jewellery and outfit are identical throughout. She stays on her left side from the landing to the end — she never rolls onto her back or sits up.
4. Competitive Position
Where Seedance 2.5 leads its own predecessor:
Duration: 4–30 s in a single pass, against 4–15 s on Seedance 2.0 2.
Example:
See the prompt here
Create an exactly 12-second photorealistic 9:16 cinematic video of a professional ballerina performing a dramatic arabesque in a grand, softly lit ballet studio.
She wears a pale blush leotard, a lightweight knee-length chiffon skirt, pink tights, and satin pointe shoes. Her hair is secured in a neat ballet bun. Maintain the same dancer, face, outfit, hairstyle, and body proportions throughout.
0:00–0:03. Begin with a full-body side view. The ballerina stands in fifth position beneath a focused warm spotlight. She takes one controlled step forward and slowly opens her arms, preparing to balance.
0:03–0:09. Shift into elegant cinematic slow motion as she rises onto pointe and gradually raises her back leg into a high classical arabesque. Emphasize the precise moment her foot leaves the floor. Her leg lifts smoothly against gravity while her supporting leg remains straight and stable.
Her front arm reaches forward as the other extends softly behind her. The chiffon skirt floats upward with natural delay, then settles around her body. Tiny dust particles move through the spotlight while the camera performs a slow dolly push-in. Use dramatic orchestral music that gradually rises with her leg.
0:09–0:12 She reaches the complete arabesque and holds it with a long neck, relaxed shoulders, pointed toes, and a calm but powerful expression. The music reaches a soft emotional peak as the camera finishes at a three-quarter angle. Use realistic slow-motion ballet technique, muscle control, balance, cloth movement, hair inertia, and gravity. Keep her entire body, both hands, and both feet visible. Do not exaggerate her flexibility, distort her limbs, change dancers, alter her outfit, or use sudden camera movements.
Reference capacity: up to 50 assets – 30 images, 10 video, 10 audio – in one request 1.
Example:



See the prompt here
Create an exactly 10-second photorealistic 9:16 lifestyle video of a professional woman enjoying coffee in a modern office.
She wears a tailored charcoal-gray trouser suit, an ivory blouse, and polished black closed-toe high heels. Her hair is styled in a sleek half updo with a clean part and controlled flyaways. Maintain the same woman, face, outfit, hairstyle, coffee cup, and office throughout.
0:00–0:03
Begin with a full-body tracking shot as she walks confidently toward a large office window, holding a coffee cup in her right hand. Her heels land naturally with each step while the coffee remains stable inside the cup.
0:03–0:07
She stops beside the window and looks at the city view. The camera slowly moves closer as she raises the cup toward her face. Soft steam rises from the fresh coffee. She closes her eyes briefly and breathes in the aroma with a relaxed expression.
0:07–0:10
She takes one small sip, lowers the cup, and smiles softly as if enjoying a peaceful break before returning to work. End with a polished three-quarter portrait showing her suit, coffee cup, professional hairstyle, and calm expression.
Use elegant natural movement, realistic coffee volume, subtle steam, soft daylight, and clean office ambience. Keep her posture confident and professional. Do not change her outfit, hairstyle, shoes, cup, or identity. Do not spill the coffee, add extra people, or introduce text and logos.
Editing model: timestamp-level targeted editing and a locked/unlocked task distinction that Seedance 2.0 does not have 5.
Example:
See the prompt here
Create an exactly 10-second photorealistic food commercial with exactly four separate shots. Do not merge, skip, or add shots.
Product: a clear glass jar labeled “Sizzle,” filled with an appetizing dry rub made from smoked paprika, garlic, black pepper, dried thyme, rosemary, and chili flakes. Keep the jar, label, spice color, and fill level consistent throughout.
0:00–0:02.5: Macro close-up. The Sizzle jar stands on a dark wooden kitchen counter beside raw chicken. Warm, moody lighting highlights the colorful spice texture through the glass.
0:02.5–0:05: Overhead shot. A hand wearing a clean black cooking glove opens the jar and sprinkles the dry rub evenly over the chicken. Show individual herbs and spices landing clearly on the surface.
0:05–0:07.5: Side-view slow motion. The falling spice particles stop completely in mid-air for one second, then move upward in reverse and return neatly into the jar. Treat this impossible movement as intentional, elegant, and visually convincing.
0:07.5–0:10: Hero shot. Show the sealed Sizzle jar beside richly seasoned grilled chicken with a crisp, smoky surface. Gentle steam rises while the camera slowly pushes toward the jar.
Mood: bold, warm, premium, smoky, and appetizing. Use realistic food textures, controlled cinematic lighting, and clean camera movement. Include exactly one hand and one jar. Do not add people, utensils, flames, extra products, extra text, floating graphics, or unrequested visual effects
Output container: .mp4 and .mov, against .mp4 only on 2.0 2.
Where Seedance 2.5 is outperformed:
Resolution: Seedance 2.0 renders up to 4K at 10-bit colour; 2.5 stops at 1080p 2. For broadcast or large-format delivery the older model is the stronger choice.
Example:
See the prompt here
Create an exactly 12-second photorealistic 9:16 video focused entirely on Japanese fireworks, mastered in native 4K with 10-bit color for broadcast and large-format display.
0:00–0:04
Against a pure black night sky, a single golden firework rises and explodes into a perfectly symmetrical chrysanthemum pattern. Show thousands of fine sparks expanding outward with crisp detail, then slowly fading from bright gold to warm amber.
0:04–0:08
Deep crimson, violet, and sapphire fireworks bloom in layered circular patterns. Preserve smooth gradients between the bright cores and darker outer sparks without color banding. Each explosion produces realistic light, smoke, falling embers, and gradual color decay.
0:08–0:12
A large Japanese-style finale fills the frame with overlapping gold, blue, pink, and silver fireworks. Fine sparkling trails fall naturally under gravity while soft illuminated smoke remains visible between explosions. End with one enormous golden willow firework spreading across the entire screen.
Use a locked camera aimed only at the sky. Deliver crisp 4K detail, rich 10-bit color, smooth tonal transitions, controlled highlights, and deep clean blacks. Do not show people, buildings, landscapes, water, lanterns, text, logos, or scenery. Avoid blown-out centers, crushed shadows, color banding, excessive saturation, blurry sparks, artificial symmetry errors, or frozen smoke.
Shot-list compliance: briefs written as explicit shot lists are executed more literally by 2.0; 2.5 optimises for continuity instead.
Example:
See the prompt here
Create 12-second photorealistic 9:16 time-lapse of brownies baking inside an oven. Execute exactly four separate shots in the stated order. Each shot lasts exactly three seconds. Use clean hard cuts. Do not merge, skip, extend, reorder, or add shots.
Maintain the same square black metal baking pan, oven rack, chocolate batter, chocolate chips, parchment paper, oven interior, and warm oven light throughout.
Shot 1. 0:00–0:03. Wide front view
From inside the oven, show the pan sliding onto the centre rack. The raw brownie batter is dark, glossy, thick, and level, with evenly scattered chocolate chips. The pan stops securely in the centre.
Shot 2. 0:03–0:06. Macro side view
Hard cut to an extreme close-up of the batter’s edge. In accelerated time-lapse, tiny bubbles form and pop, the edges gradually set, and the batter rises slightly against the parchment paper. The progression must move continuously from raw to partially baked.
Shot 3. 0:06–0:09. Direct overhead view
Hard cut to a top-down view inside the oven. The glossy surface slowly becomes matte while a thin, shiny crust develops. Natural cracks spread outward from the centre, and the chocolate chips soften without disappearing.
Shot 4. 0:09–0:12. Low three-quarter hero view
Hard cut to a low angle near the oven rack. Show the fully baked brownies with a delicate crinkled top, defined cracks, firm edges, and a slightly soft centre. Subtle steam rises as the rack moves slowly toward the open oven door.
Keep the baking transformation progressive and physically believable. Do not return the brownies to an earlier stage, replace the pan, change the batter volume, merge camera angles, add ingredients, show people, or introduce extra scenes.
5. Use Case Guide
Best fit: product storytelling and reference-locked production
- Long single-take brand and product films: 30 seconds in one pass, extendable further 1.
- Packaging and label-forward product video: printed type stays legible and anchored through movement and lighting change.
- Animating existing product photography: the supplied start frame is preserved rather than reinterpreted.
- Reference-heavy campaigns: character, set, style and audio locked from many aligned assets in a single request 5.
- Editing and extending existing footage: targeted edits and forward or backward extension avoid a full regeneration 5.
6. Kittl’s Verdict
Seedance 2.5 is the model to reach for when the job is long, reference-heavy, and text-bearing. It holds a supplied start frame without repainting it, keeps printed lettering legible and anchored as the shot moves, executes camera moves as written, and produces audio locked to picture in a single pass 5. Thirty seconds in one take is a real production advantage over anything in the previous generation 1, and on Kittl the twenty reference images it accepts are the deepest reference support on the platform.
What it will not do is follow a shot list. Explicit multi-shot briefs come back merged, named scene counts are not honoured, and very short prompts invite the model to invent copy and agents of its own. It is also slow, and the resolution ceiling is the sharpest limitation: the model stops at 1080p where Seedance 2.0 reaches 4K 2, and on Kittl it delivers 720p, Seedance 2.0 Full HD and Seedance 2.0 4K are the routes to higher resolution on the same platform.
Bottom line: Use Seedance 2.5 for continuity, references and legible type, and Seedance 2.5 HD when the deliverable needs 1080p. Move to Seedance 2.0 4K when it needs more than that, and to Seedance 2.0 when the brief is a shot list that must be executed literally.
References
- ByteDance Seed — Seedance 2.5 launch post: https://seed.bytedance.com/en/blog/one-take-creation-flexible-referencing-introducing-seedance-2-5
- BytePlus ModelArk — Model list (resolution, frame rate, duration, formats, rate limits): https://docs.byteplus.com/en/docs/ModelArk/1330310
- BytePlus ModelArk — Pricing (token rates, billing formula, draft mode): https://docs.byteplus.com/en/docs/ModelArk/1544106
- BytePlus ModelArk — Create a video generation task (API parameters): https://docs.byteplus.com/en/docs/ModelArk/1520757
- BytePlus ModelArk — Dreamina Seedance 2.5 prompt guide (capabilities, task model, languages): https://docs.byteplus.com/en/docs/ModelArk/2607689
- Dreamina, ByteDance — Seedance 2.5 product page: https://dreamina.capcut.com/seedance/seedance-2-5
- Kittl hands-on scoring — Seedance 2.5 on the Kittl model comparison wall, August 2026. Uncited observations throughout this page come from this source.
Last updated: August 2026. Tested on Seedance 2.5 via Kittl. Scores are Kittl hands-on; specifications are cited to ByteDance and BytePlus official documentation. Model capabilities subject to change.

Kittl gives you the power of a world-class creative agency on one platform. Build product imagery, labels, packaging, and campaign assets in one place, with AI tools, a professional editor, and templates built by some of the best designers in their fields. Set your brand once, then create on brand every time.



