Model eval · round r1

Flux 3 Video vs Seedance 2

Seven live proven templates, five fashion and two beauty. Every row reuses the exact inputs that produced the template's example video: the same reference images, the same enhanced prompt, the same duration and aspect ratio. The video model is the only thing that changes.

Flux 3 is not available on our Kie account, so the Flux arms run on fal under blackforestlabs/flux-3. Seedance clips are the existing renders and were not regenerated.

What the 21 clips actually showed

  1. Flux image-to-video is unusable for fashion. It accepts one image, so it never sees the garment reference and invents the clothes. All five fashion rows came back wearing the wrong outfit. For a template whose job is "show my product," that is disqualifying regardless of how good the video looks.
  2. Keyframes is the only Flux path that respects the inputs, and it works at two to four references. At five it collapses into a crossfade of the raw input photos rather than a generated scene.
  3. New failure mode: baked-in supplier watermarks. Because keyframes puts the source photos on screen, the supplier's overlay text ("AWC 13740 Exclusives design") is rendered into the video on three rows. That breaks the no-on-screen-text rule and leaks a third party's watermark into customer content. Seedance never does this, because it treats references as conditioning rather than frames.
  4. Both Flux arms miss the before/after format entirely. Given the product shots, Flux made product b-roll with no transformation and no face. Seedance built the full arc from the same two images.
  5. Flux always opens on the raw input photo as a literal first frame, which reads as a static open into a hard cut.
  6. Where Flux is genuinely competitive: GRWM. Both arms produced clean, usable clips, and the keyframes one is arguably the most polished clip in the whole set.
  7. Several Flux clips end on black or blank frames.
TemplateRefsSeedance 2Flux i2vFlux keyframes
3 Ways to Style One Piece5holds upfailsfails
Mirror Outfit Check3holds upfailsholds up
Honest Fit Review3holds upfailsflawed
Festive Get-Ready4holds upfailsholds up
Fabric & Detail Close-Up3holds upfailsflawed
Before / After Glow2holds upfailsfails
GRWM (No Talking)2holds upholds upholds up

3 Ways to Style One Piece

style-three-ways fashion apparel Anmol Collections 12s · 9:16 5 reference images

Seedance 2 (live)bytedance/seedance-2 · the clip on the live template card
Three distinct styling looks on the correct garment, identity holds.
Flux 3 (image-to-video)blackforestlabs/flux-3/image-to-video · 1 image in
Wrong garment (invented a plain magenta kurta) and only one look, not three.
Flux 3 (keyframes)blackforestlabs/flux-3/keyframes-to-video · all images in
Collapses into a crossfade of the raw input photos, and bakes in the supplier watermark "AWC 13740 Exclusives design". Five keyframes over 12s leaves no room to generate.

The inputs, identical across all three

Prompt and source cell
Voice-over script
3 ways to wear this—Anmol Collections. Dupatta loose, belt karke, jacket add; har look easy aur festive. Link in bio, check karo.
Source
ex-fashion/EX-F-FA-P01-T10-B
Prompt (4372 chars)
Vertical 9:16 "3 ways to style this one piece" fashion UGC video for Instagram Reels, ultra-realistic and shot on an iPhone camera. Total duration: 12 seconds.

Use @Image1 as the character reference. Maintain the same face, hairstyle, body type, skin tone, and overall identity throughout the full video. Character persona: graceful ethnic-wear creator — poised, warm, and traditional in movement.

This format shows the SAME core garment styled in different ways. The hero piece must be visually consistent across every beat — only the accessories, footwear, layering, or styling combinations change between cuts.

Use @Image2 as Way 1, @Image3 as Way 2, and @Image4 as Way 3 — each is the same core piece styled differently. Preserve the core garment exactly across all 3 beats; vary only the accessories, footwear, layering, or styling cues that distinguish each way.

Use @Image5 as the setting reference: a warm Chandni Chowk-style boutique corner — soft warm light from the right, wooden display shelf and brass accents visible, patterned rug on the floor. Keep the setting CONSISTENT across all 3 ways.

Camera: vertical iPhone footage, mostly static, full-body or 3/4-body framing, slight natural handheld feel. SAME camera position and framing across all 3 beats (framing: chest-to-toe, centered, distance unchanged).

Timing: three equal beats, ~4.0s each. Use hard jump cuts between beats.

[0–4.0s] Beat 1 — Way 1
- Action: Creator walks in from left and stops in center, holds one confident pose showing full outfit. Minimal movement (one hand adjusting dupatta).
- Styling: use @Image2 styling exactly. Preservation: preserve the core garment's color (rich maroon embroidered kurta), longline flowy silhouette, and visible dupatta fabric and border — do not alter embroidery, length, or color.

[4.0–8.0s] Beat 2 — Way 2
- Hard jump cut. Action: same framing; creator does a 3/4 turn to camera and places hand on hip to show changed styling.
- Styling: use @Image3 styling exactly. Preservation: preserve the core garment's color (rich maroon embroidered kurta), longline flowy silhouette, and embroidery details — only change accessories/layering as in @Image3 (e.g., add belt or tuck) and footwear; do not alter fabric or pattern.

[8.0–12.0s] Beat 3 — Way 3
- Hard jump cut. Action: same framing; creator faces camera and gives a final confident pose with slight shoulder pop.
- Styling: use @Image4 styling exactly. Preservation: preserve the core garment's color (rich maroon embroidered kurta), longline flowy silhouette, and embroidery details — only add outer layer or jacket and accessories as shown in @Image4; do not alter the garment itself.

Style direction: comparative styling reel, warm festive vibe, natural skin tones, soft golden highlights from right-side light. Keep movements graceful and deliberate; avoid chaotic camera moves or extreme zooms.

VO (spoken, Hinglish) — Dialogue mode ON. Use this exact script as the only spoken dialogue / voice-over, delivered naturally by the creator and timed across the three beats. Do not add extra spoken CTAs.
- Exact script (must be used verbatim): "3 ways to wear this—Anmol Collections. Dupatta loose, belt karke, jacket add; har look easy aur festive. Link in bio, check karo."
- VO timing suggestion: deliver the first clause in [0–2s] as the hook, middle styling callouts across [2–9s], and final CTA line in [9–12s]. If physically cannot fit, trim only minimally but do NOT rewrite brand wording or language.

Music: brisk upbeat snippet with a clean beat-drop at each jump cut (at 4.0s and 8.0s).

Restrictions: no on-screen text overlays, no added people, no clothing morphs — use hard cuts only. Preserve identity and all facial features from @Image1 throughout.

Output-ready notes for production team: use @Image1 for the actor; style each beat to match @Image2, @Image3, @Image4 respectively; use @Image5 as the set dressing reference; keep each beat ~4s, VO exact Hinglish script as above.

@Audio1 is the complete and only spoken performance for this video. The output's voice track must be this exact audio, word for word — same words, same pacing, same pronunciation. If a person is on camera while speech plays, their lips sync to @Audio1 verbatim. Do not generate, add, or substitute any other speech, narration, or voice-over. Background music stays subtle under the voice.

Mirror Outfit Check

mirror-selfie fashion apparel Anmol Collections 12s · 9:16 3 reference images

Seedance 2 (live)bytedance/seedance-2 · the clip on the live template card
Correct crimson-pink sharara, real interior, consistent identity.
Flux 3 (image-to-video)blackforestlabs/flux-3/image-to-video · 1 image in
Invented a mint-green outfit. Never saw the garment reference. Opens on the raw headshot, then hard-cuts.
Flux 3 (keyframes)blackforestlabs/flux-3/keyframes-to-video · all images in
Correct magenta embroidered sharara, identity holds, no bad cut. Best Flux result in the set.

The inputs, identical across all three

Prompt and source cell
Voice-over script
Dekho yeh sharara set. Hand embroidery, halka stretch fit, bahut comfortable. Anmol Collections. Link se check karo.
Source
ex-fashion/EX-F-FA-P01-T1-B
Prompt (3504 chars)
Vertical 9:16 mirror-selfie fashion UGC video for TikTok / Instagram Reels, ultra-realistic and shot on an iPhone camera. Total duration: 12 seconds.

Use @Image1 as the character reference. Maintain the same face, hairstyle, body type, skin tone, and overall identity throughout the full video. Persona: graceful Brand-clinic ethnic creator — poised, warm, traditional energy.

Use @Image2 as the outfit reference. Preserve the exact outfit design, color, sharara silhouette, hand-embroidery detail, fabric drape, slight stretch fit, and visible embellishment from @Image2. Preservation note: keep the sharara set's color, flowy flared silhouette, and hand embroidery detail exactly as shown; do not alter fit, pattern, or embellishment placement.

Use @Image3 as the mirror / dressing-area setting reference. Setting description: a Chandni Chowk-style dressing area with a full-length mirror, warm ambient light from the right, tiled or wooden floor and a visible dressing table surface — realistic reflection of the room as in @Image3.

Camera: mirror-selfie style iPhone capture. The creator (from @Image1) holds the phone naturally while filming herself in the mirror. Use realistic reflection framing, slight natural handheld movement, and clear outfit visibility. No on-screen text overlays or brand text rendered into the video.

Action progression (three beats, timed to 12s total):

[0–4s] Beat 1 — Stand and turn: Creator stands in front of the mirror wearing @Image2, holding the phone as in a mirror-selfie. Small body turn to show the full sharara silhouette; steady framing so the full outfit is visible in the mirror. Spoken dialogue (voice-over / natural lip-sync): "Dekho yeh sharara set. Hand embroidery, halka stretch fit, bahut comfortable. Anmol Collections. Link se check karo." (Begin speaking immediately; natural paced delivery fits within the 12s total.)

[4–8s] Beat 2 — Reframe and side pose: Creator angles the phone slightly to capture the outfit better, shifts weight to a side pose to show flared sharara movement and embroidery close to camera. Maintain soft confident expression and gentle handheld motion. Continue same spoken line if still mid-sentence, natural cadence.

[8–12s] Beat 3 — Final outfit-check: Creator steps slightly closer or back for a clear hero hold of the outfit in the mirror, pauses with a subtle head tilt and soft smile for a clean final shot. End with the last words of the supplied script. Hold final frame cleanly for the last second.

Style: authentic mirror-check content, platform-native, casual yet traditional. Music: soft lo-fi or chill Indian pop bed at low volume under dialogue; no aggressive drops. Avoid distorted reflections, identity or outfit drift, extra limbs, chaotic camera movement, and any inserted on-screen brand text.

VO placement: full supplied Hinglish script must be the only spoken dialogue/voice-over. Trim only if physically cannot fit; do not rewrite wording.

Deliverable: single continuous 12s mirror-selfie clip shot as described, using @Image1, @Image2, @Image3 references and preserving the outfit details exactly.

@Audio1 is the complete and only spoken performance for this video. The output's voice track must be this exact audio, word for word — same words, same pacing, same pronunciation. If a person is on camera while speech plays, their lips sync to @Audio1 verbatim. Do not generate, add, or substitute any other speech, narration, or voice-over. Background music stays subtle under the voice.

Honest Fit Review

fit-review fashion apparel Anmol Collections 12s · 9:16 3 reference images

Seedance 2 (live)bytedance/seedance-2 · the clip on the live template card
Correct garment, shop setting, close-up detail beat.
Flux 3 (image-to-video)blackforestlabs/flux-3/image-to-video · 1 image in
Wrong garment (beige/gold instead of pink).
Flux 3 (keyframes)blackforestlabs/flux-3/keyframes-to-video · all images in
Correct garment and setting, but bakes in the supplier watermark "AWC 13740 Exclusives design".

The inputs, identical across all three

Prompt and source cell
Voice-over script
Anmol Collections — yeh suit perfect fit hai, soft fabric aur detailed embroidery. Ghar baithe order karo, link pe jao.
Source
ex-fashion/EX-F-FA-P01-T6-B
Prompt (3204 chars)
Vertical 9:16 fashion fit-review UGC video for TikTok / Instagram Reels, ultra-realistic and shot on an iPhone camera. Total duration: 12 seconds.

Use @Image1 as the character reference. Maintain the same face, hairstyle, body type, skin tone, and overall identity throughout the full video.

Use @Image2 as the outfit reference. Preserve the exact outfit design, color, silhouette, fabric feel, print, pattern, embroidery, fit, and major visible details from @Image2.

Use @Image3 as the setting reference. Describe setting exactly as in @Image3: a warm indoor Chandni Chowk–style boutique corner with soft overhead warm light, wooden floor, and a decorative mirror or traditional textile backdrop visible — natural practical lighting from the right side.

Camera: vertical iPhone footage, creator talking-to-camera style, natural UGC framing with a mix of medium and full-body shots. Keep the outfit clearly visible.

Action progression (three beats, scaled to 12 seconds total):

[0–4s] Beat 1 — Intro reaction
- Medium shot, creator facing camera wearing the outfit from @Image2, natural honest smile and slight head-nod. She lightly adjusts the neckline or dupatta once with a slow, natural hand movement to show fabric feel. VO (same timing): spoken line in Hinglish begins immediately.

[4–8s] Beat 2 — Demonstrate fit
- Full-body to 3/4 turn: creator turns slightly to right then left to show front and side fit; she takes two small walking steps forward to demonstrate how the outfit falls and moves. Quick close-up of embroidery when she pauses with hand near the bodice.

[8–12s] Beat 3 — Final verdict
- Medium shot, creator gives a satisfied confident pose (hand on hip or gentle dupatta toss), smiling to camera for the final look. Hold final hero frame as VO finishes.

Dialogue / Voice-over (Hinglish, lip-sync on-screen): use ONLY this exact line — do not add or change language: "Anmol Collections — yeh suit perfect fit hai, soft fabric aur detailed embroidery. Ghar baithe order karo, link pe jao." Align VO timing across the three beats so speech is clearly audible and sits above minimal ambient bed.

Style: honest creator review, natural and trustworthy, realistic fashion UGC, warm indoor light, clear outfit visibility.

Sound: minimal — soft ambient room tone with a barely-there acoustic bed; speech sits on top.

Preservation instructions:
- Outfit (@Image2): preserve the exact color, longline ethnic salwar silhouette, flowy fabric drape, and the visible detailed embroidery on the bodice and borders as shown in @Image2; do not alter fit or embellishment placement.

Technical notes / constraints: hard cuts only between beats; no on-screen text overlays or added brand text; single person in frame only; avoid chaotic camera moves or aggressive zooms; natural hand movements only.

@Audio1 is the complete and only spoken performance for this video. The output's voice track must be this exact audio, word for word — same words, same pacing, same pronunciation. If a person is on camera while speech plays, their lips sync to @Audio1 verbatim. Do not generate, add, or substitute any other speech, narration, or voice-over. Background music stays subtle under the voice.

Festive Get-Ready

festive-get-ready fashion apparel Anmol Collections 12s · 9:16 4 reference images

Seedance 2 (live)bytedance/seedance-2 · the clip on the live template card
Full arc: garment, drape, jewellery beat, mirror reveal.
Flux 3 (image-to-video)blackforestlabs/flux-3/image-to-video · 1 image in
Wrong garment (beige anarkali), a baked-in studio watermark, ends on a black frame.
Flux 3 (keyframes)blackforestlabs/flux-3/keyframes-to-video · all images in
Correct garment, diya and jewellery beats, coherent arc.

The inputs, identical across all three

Prompt and source cell
Voice-over script
Diwali ready? Anmol Collections ka hand-embroidered sharara, lightweight comfort, perfect for family puja. Shop the collection.
Source
ex-fashion/EX-F-FA-P01-T4-B
Prompt (4018 chars)
Vertical 9:16 festive get-ready fashion UGC video for Instagram Reels, ultra-realistic and shot on an iPhone camera. Total duration: 12 seconds.

Character persona: graceful ethnic-wear creator — warm, poised, familiar with traditional dressing. Use @Image1 as the character reference. Maintain the same face, hairstyle, body type, skin tone, and overall identity throughout the full video.

Outfits & jewellery preservation:
- Preserve @Image2 exactly: the visible garment is a festive ethnic sharara-style outfit — keep the color, flared silhouette, hand-embroidered embellishment, lightweight fabric drape, and any dupatta or sequin/embroidered details exactly as seen.
- Preserve @Image3 exactly as the jewellery set: keep the design, metal tone, stonework, shape (jhumkas/necklace/bangles if visible) and placement exactly as shown.

Setting: use @Image4 as the getting-ready space reference. Describe the space exactly as in @Image4: a cozy dressing corner with warm ambient light, a wooden surface (vanity or low table) and a visible mirror or soft textured backdrop. Light direction: warm golden-side fill; surfaces: wooden/soft fabric; visual anchors: the mirror and a small tray for jewellery.

Camera: vertical iPhone handheld feel. Medium-framing for ritual moments, 3/4 or full-body framing for final reveal. Use hard jump cuts between outfit stages (no morphing).

Dialogue / voice-over: Dialogue mode ON. Use this exact spoken script in Hinglish, delivered naturally by the on-screen character as voice-over or lip-synced speech: "Diwali ready? Anmol Collections ka hand-embroidered sharara, lightweight comfort, perfect for family puja. Shop the collection." Do not add any other spoken CTA or text. Ensure timing fits the 12s cut.

Beat structure and precise timing (total 12s):
[0–3s] Beat 1 — Casual base, holding the festive piece: Medium shot of @Image1 in a simple cotton base layer near the vanity from @Image4, holding or smoothing the sharara from @Image2 laid on the surface. Small anticipatory smile, gentle fingertips on embroidery. No jewellery yet. Natural warm side light.

[3–7s] Beat 2 — Hard cut to wearing the festive outfit: Jump cut. @Image1 now wearing the exact outfit from @Image2. 3/4 framing as she adjusts the dupatta or smooths the sharara flare with one deliberate motion. Keep movement slow and natural; preserve outfit details and drape.

[7–10s] Beat 3 — Jewellery layer: Close-medium shot as she puts on the jewellery from @Image3 (necklace or jhumkas visible). Show the action of clasping/placing one piece; avoid extreme hand close-ups. Preserve the jewellery design exactly.

[10–12s] Beat 4 — Final festive reveal: Full-body or mirror-facing shot. She turns to camera or mirror, gives a confident soft smile and holds a poised final pose showing the full look (outfit + jewellery). Warm golden framing.

VO beat map and sync: Fit the supplied Hinglish line across the full 12s with natural pacing (approximately 2 words/sec). Start spoken hook in Beat 1 and complete CTA by Beat 4. Ensure lip-sync or voice-over matches the cuts; do not shorten or rewrite wording except only if physically impossible to fit.

Style & mood: warm festive get-ready, sensorial and intimate, realistic UGC. Music: warm modern-traditional bed (soft sitar or subtle tabla) low under VO. Avoid on-screen text, logos, or rendered brand text. Avoid multiple people in frame and any clothing morphing transitions.

Final notes: keep continuity of the same person (@Image1) across all cuts, keep all outfit and jewellery details identical to @Image2 and @Image3, and use the exact setting from @Image4.

@Audio1 is the complete and only spoken performance for this video. The output's voice track must be this exact audio, word for word — same words, same pacing, same pronunciation. If a person is on camera while speech plays, their lips sync to @Audio1 verbatim. Do not generate, add, or substitute any other speech, narration, or voice-over. Background music stays subtle under the voice.

Fabric & Detail Close-Up

fabric-detail fashion apparel Anmol Collections 12s · 9:16 3 reference images

Seedance 2 (live)bytedance/seedance-2 · the clip on the live template card
Clean macro on the real embroidery, mirror-work and tassels.
Flux 3 (image-to-video)blackforestlabs/flux-3/image-to-video · 1 image in
Wrong garment, then degrades to black frames.
Flux 3 (keyframes)blackforestlabs/flux-3/keyframes-to-video · all images in
Good macro detail, but bakes in the supplier watermark and ends on a black frame.

The inputs, identical across all three

Prompt and source cell
Voice-over script
Chandni Chowk ki zardozi aur resham kaam, aaraamdayak fit, shandar kapda. Anmol Collections. Dekho aur mangwao.
Source
ex-fashion/EX-F-FA-P01-T11-B
Prompt (3334 chars)
Vertical 9:16 fabric and detail close-up reel for Instagram Reels, ultra-realistic and shot on an iPhone camera. Total duration: 12 seconds.

Character persona: graceful ethnic-wear creator (soft, traditional, poised). Use @Image1 only as an off-camera presence: a wrist/hand entering frame to touch or lift fabric; match skin tone from @Image1. Never show the face.

Setting: use @Image3 as the surface/backdrop reference — a warm textured backdrop with soft directional warm side-light and a low-reflective surface that lets the garment colors and embroidery pop. Keep surfaces matte and minimal; avoid extra props.

Garment: use @Image2 as the garment hero. Preserve the exact outfit design, color, silhouette, fabric texture, print/pattern, embroidery/beadwork, neckline finish, and visible craft details from @Image2.

Preservation instruction for outfit (@Image2): Preserve the exact color, longline ethnic silhouette, and the resham + zardozi embroidery placement and stitch texture visible on @Image2 — do not alter thread color, motif density, or embellishment finish.

Format notes: product-first, slow and sensorial. Mostly still camera with very slow deliberate pushes/pans. One slow cut at most. No face frames, no jump cuts, no on-screen text overlays, no brand text rendered in-frame. Minimal ambient acoustic bed only.

VO (Hinglish, spoken exactly as written, reverent slow tone; total pacing ~1.5 words/sec): "Chandni Chowk ki zardozi aur resham kaam, aaraamdayak fit, shandar kapda. Anmol Collections. Dekho aur mangwao." Do not add or change words. Sync VO across beats as indicated below.

Camera/action progression and timing (total 12s):
[0–4s] Beat 1 — Wide hero: static or very-slow push-in wide shot of the full garment laid on or draped over the @Image3 surface so the silhouette, color, and overall embroidery layout read clearly. Off-camera hand from @Image1 may gently smooth a fold at frame edge. VO: first phrase "Chandni Chowk ki zardozi aur resham kaam," timed to introduce craft.

[4–8s] Beat 2 — Detail close-up: slow smooth iPhone pan to a tight close-up on a primary craft detail from @Image2 (embroidered neckline or prominent motif). Hold steady, reveal stitch texture and metallic sheen. No shake. Off-camera finger may lightly trace fabric edge (no blocking of embroidery). VO: middle phrase "aaraamdayak fit, shandar kapda."

[8–12s] Beat 3 — Texture & finish: another slow close-up on a secondary detail or the drape — gently lift a corner with the off-camera hand to show fluidity and inner finish, or show beadwork catching light. End on a calm hero hold of the most striking detail. VO final phrase "Anmol Collections. Dekho aur mangwao." Conclude with a soft fade-out in audio and image.

Additional constraints: no fast movements, no aggressive zooms, avoid harsh shadows or specular hotspots on metallic threads, keep motion minimal and reverent. Use only the supplied VO; no extra spoken CTAs.

@Audio1 is the complete and only spoken performance for this video. The output's voice track must be this exact audio, word for word — same words, same pacing, same pronunciation. If a person is on camera while speech plays, their lips sync to @Audio1 verbatim. Do not generate, add, or substitute any other speech, narration, or voice-over. Background music stays subtle under the voice.

Before / After Glow

before-after-glow beauty wellness Vaviya Developers 12s · 9:16 2 reference images

Seedance 2 (live)bytedance/seedance-2 · the clip on the live template card
Complete before/after arc: acne close-up, product, foam, clearer skin.
Flux 3 (image-to-video)blackforestlabs/flux-3/image-to-video · 1 image in
No before/after at all. Static product, blurred hands, blank white frame. Format missed.
Flux 3 (keyframes)blackforestlabs/flux-3/keyframes-to-video · all images in
No before/after either. Product b-roll only. Format missed.

The inputs, identical across all three

Prompt and source cell
Voice-over script
Skin bejaan lag rahi hai? Roz face wash + SPF50 se nikhra aur protection milega. Kaleigh Beauty, ek baar try karo.
Source
preview-01/PREV-BW-P01-T1-B
Prompt (3176 chars)
Format: Vertical 9:16, 12 seconds. Use provided images as exact assets: @Image1 is the hero product/treatment photo; @Image2 is the real vanity/studio setting and must be used for the ritual scene. Preserve the product packaging: the product is a tube (or pump-style bottle if visible) in a muted pastel color with a cylindrical/rounded shape — keep label visible but not legible. Tone / voice: Affordable Gen-Z / problem-solving, real-skin, believable improvement. Dialogue mode ON — spoken audio only (Hinglish script exactly as provided). No on-screen text generated in-scene; any tags or brand text are post only. Natural sound cues (dropper click/lather/soft towel) layered under music. Camera: vertical iPhone UGC framing, close macro and short handheld moves. No full faces; focus on cheek/hand/neck area or product application area shown in @Image1. Hard cuts between beats, realistic skin texture, subtle glow in after shot. Timecodes and beats: [0–3s] Beat 1 — Dull (first ~25%): Start with a muted, slightly desaturated close-up pulled from the aesthetic of @Image1 showing the 'before' skin state (dull, slightly textured cheek or hand) in soft, flat light. Slow 1–2 second micro-pan to show honest texture. Ambient foam/drip sound muted. No talking yet. [3–8s] Beat 2 — The ritual (middle ~40%): Hard cut into the vanity from @Image2 — warm, slightly backlit. Show the hero using the product from @Image1: one deliberate application gesture (squeeze or pump then rub between fingertips, then gentle upward strokes on cheek/hand). Include two close macros: (a) product texture on fingertips, (b) application stroke. Preserve packaging appearance exactly as described. Diegetic sounds: product click/pump and soft rubbing; music opens slightly. Voice-over (start at ~3.2s, intimate tone, only this exact line): "Skin bejaan lag rahi hai? Roz face wash + SPF50 se nikhra aur protection milega. Kaleigh Beauty, ek baar try karo." (Use natural, friendly Hinglish; sync to the ritual so words flow over the application.) [8–12s] Beat 3 — Glow (final ~35%): Hard cut back to the original framing from Beat 1. Brighter warm grade, subtle specular highlights on hydrated skin, realistic pore texture visible. Single small gesture (brush hair back or light exhale/soft smile off-camera) to sell confidence. End on a calm hold: the product from @Image1 resting next to the glowing cheek/hand in the @Image2 vanity environment for the final 0.8–1s. Music swells gently then fades. Audio notes: spoken script is the only VO; no additional spoken copy or CTAs. Keep improvement believable (hydration, subtle glow), no airbrushing or impossible changes. Avoid on-screen text, legible labels, numbers, price/claims visually. Deliver as single 12s vertical cut with hard cuts between beats.

@Audio1 is the complete and only spoken performance for this video. The output's voice track must be this exact audio, word for word — same words, same pacing, same pronunciation. If a person is on camera while speech plays, their lips sync to @Audio1 verbatim. Do not generate, add, or substitute any other speech, narration, or voice-over. Background music stays subtle under the voice.

GRWM (No Talking)

grwm-no-talking beauty wellness Vaviya Developers 12s · 9:16 2 reference images

Seedance 2 (live)bytedance/seedance-2 · the clip on the live template card
Correct vanity GRWM with the product in frame.
Flux 3 (image-to-video)blackforestlabs/flux-3/image-to-video · 1 image in
Genuinely good GRWM, product label legible and correct. Opens on a static product box.
Flux 3 (keyframes)blackforestlabs/flux-3/keyframes-to-video · all images in
Most polished clip Flux produced. Clean vanity arc, correct product.

The inputs, identical across all three

Prompt and source cell
Voice-over script
Aaj ka look ready—salicylic-glycolic face wash plus SPF 50 gives skin protection. Kaleigh Beauty, try it.
Source
preview-01/PREV-BW-P01-T3-B
Prompt (2931 chars)
Concept: A 12s vertical GRWM, no-talking-on-camera UGC focused on a single hero product @Image1 with vanity setting from @Image2. Realistic Indian creator (young adult, fresh-faced skincare user) — hair clipped back, natural skin texture, subtle imperfections allowed. Preserve packaging: the product is shown as a white squeeze tube with a matte finish and rounded cap (keep label visible but not legible). Keep results believable: subtle brighter, hydrated skin.

Beats (strict timecodes):
[0–2.4s] Beat 1 — Sit-down: Creator enters frame at the vanity visible in @Image2, hair clipped back, sits facing the mirror/phone tripod. Static vertical framing, natural daylight or soft vanity light. No mouth movement. Natural diegetic sounds (mirror tap, seat shift). Spoken VO starts immediately: "Aaj ka look ready—salicylic-glycolic face wash plus SPF 50 gives skin protection. Kaleigh Beauty, try it." (VO only; face remains silent.)

[2.4–9.0s] Beat 2 — Steps run (jump-cut rhythm ~115–120 BPM): Quick jump-cut 1: small prep dab (splash of water or towel gesture) [2.4–3.3s]. Jump-cut 2 (longest close-up) [3.3–6.6s]: deliberate application with the white squeeze tube @Image1 — clear hand-to-face gesture, slow gentle lather/texture push-in macro, keep packaging visible in hand. Jump-cut 3 [6.6–7.8s]: blend/rinse-off or pat motion at the mirror, natural small adjustments (hair tuck). Keep framing consistent for snap cuts. VO continues over these cuts within the same delivered script line.

[9.0–12.0s] Beat 3 — The look (final): Hard cut to final mirror check — creator turns slightly to camera with a quiet confident half-smile (mouth closed), then looks to mirror. Place the white tube (@Image1) on the vanity (from @Image2) in-frame lower third. Hold final pose ~2.5–3s. End on natural ambient vanity sound; VO concludes within the 12s.

Camera & style notes: Vertical iPhone UGC framing, mix of medium vanity shots and a macro texture push-in for the product application. Lighting soft and flattering, no heavy beauty filters. Keep skin pores/texture visible. No on-screen text, no labels rendered legible, no additional products, no impossible skin transforms, and no model lip-syncing. Music: muted trending GRWM pop at low volume behind VO; add small diegetic taps and water sounds.

Audio requirement: Use only the supplied Hinglish spoken script exactly: "Aaj ka look ready—salicylic-glycolic face wash plus SPF 50 gives skin protection. Kaleigh Beauty, try it." Timing must fit within 12s; trim only if absolutely necessary.

@Audio1 is the complete and only spoken performance for this video. The output's voice track must be this exact audio, word for word — same words, same pacing, same pronunciation. If a person is on camera while speech plays, their lips sync to @Audio1 verbatim. Do not generate, add, or substitute any other speech, narration, or voice-over. Background music stays subtle under the voice.