Video creation & character continuity

Imagine Video 1.5

Video creation & character continuity

Every practical way to prompt scenarios, lock a character across clips, keep voice consistent, and build progressive multi-part stories.

Imagine Video 1.5 with References is the leap: text-to-video, multi image references, voice consistency, and native 1080p. SuperGrok Heavy / Plus get priority access; features continue to roll out by tier and region. This chapter is the production playbook.

Generation modes (know which dial you are turning)

ModeInputsBest for
Text-to-video (T2V)Prompt onlyNew shots, establishing scenes, concepting
Image-to-video (I2V)Still + motion promptExact first frame control; animate a designed shot
Reference-to-video (R2V)1–7 refs + promptSame face/product/location in a new scene
Image + voice referenceCharacter image + voice sampleDialogue / singing with locked timbre
Note

Durations are short by design (commonly 6s; longer tiers where available). Professional workflows are many short shots, not one endless take. Assemble with an editor or ffmpeg.

Prompt craft for different scenarios

Write present tense, one clear action, one camera move. Overloading motion causes mush.

Cinematic establishing

text
Dawn over a rain-slick neo-Tokyo alley, neon reflecting in puddles.
Slow aerial push-in toward a lone figure under a red umbrella.
Cinematic anamorphic bokeh, grounded physics, subtle ambient city hum.

Product hero

text
A matte-black wireless earbud rotates slowly on a glass pedestal.
Soft studio key light, gentle camera orbit, crisp reflections,
premium commercial look, no text overlays.

Action beat

text
The courier vaults a rooftop gap at night, coat flaring.
Handheld medium shot tracking beside her, light motion blur,
rain streaks, tense electronic underscore. Single jump only.

Dialogue / performance

text
Close-up: the captain speaks calmly to camera on a starship bridge.
Subtle head movement, soft blinks, stable framing, shallow depth of field.
Keep lips natural; no rapid cuts.

Music / singing

text
The same woman sings a single sustained chorus line on a dim stage.
Microphone in hand, soft backlight haze, locked face and voice reference,
camera slow dolly-in, no choreography changes.

Game-style / Minecraft-ish adventure

text
Blocky explorer and a wolf crest a snowy hill at golden hour.
Gentle side-tracking shot, coherent character proportions,
charming adventure mood, stable silhouette.

Step 1 — Build a character bible (do this once)

Continuity fails when you re-describe a person from memory every time. Create a master reference pack:

  1. Master portrait — neutral light, clear face, signature outfit, plain background.
  2. Turnarounds — front / 3⁄4 / side (edit from the master, do not re-roll from scratch).
  3. Expression set — calm, smile, determined (same lighting).
  4. Outfit card — bullet list of colors, materials, props.
  5. Voice sample — clean line reading for voice lock (when available on your tier).
Master portrait prompt
Studio portrait of Mira Chen, 28, sharp emerald eyes, defined jaw,
messy auburn hair in a loose bun, small scar above left eyebrow,
tailored navy trench over white silk blouse, thin silver pendant.
Neutral gray background, Rembrandt lighting, 85mm, photoreal,
identity-defining details sharp.

Step 2 — Lock identity across videos

Method A — Multi image references (primary)

Upload up to seven reference images. Each ref should lock one thing: face, full body outfit, location, key prop. Prompt the change (action/scene), not a full redescription of the face.

R2V continuity prompt
Using the character references for face and outfit:
Mira walks through a sunlit archive hall, dust in light shafts.
Medium shot, steady steadicam follow, calm confidence.
Keep exact face geometry, hair, trench coat, and pendant.
Change only location and walk cycle.

Method B — Image-to-video from a composed still

For maximum first-frame control: generate/edit a perfect still (with character refs), then animate that still. Best for complex composition.

I2V motion prompt
Animate: slow push-in, she turns her head slightly toward camera,
coat fabric settles, dust motes drift. No new characters, no outfit change.

Method C — Voice + face

Pass character image + voice reference so dialogue scenes keep both likeness and timbre. Reuse the same voice pack for the whole series.

Method D — Edit chain for stills before video

Never re-generate the character from text alone for episode 2. Always image_edit / reference from the master: “Keep this exact character — change only the background to a night market.”

Freeze list

In every continuity prompt, list what must not change first (face, hair, scars, outfit colors, body proportions), then the single change. Models drift toward average faces without a freeze list.

Progressive multi-video series (episode pipeline)

Think like an episodic director. Each clip is one beat; the series is assembled later.

  1. Series bible — logline, 3-act beats, recurring locations, cast list with ref packs.
  2. Shot list — rows: shot id, duration, refs used, camera, action, dialogue.
  3. Still pass — keyframes for hard shots (fight, reveal, product insert).
  4. Animate pass — T2V/I2V/R2V per shot with the same refs.
  5. Select — keep only takes that hold identity; regenerate failures with tighter freeze lists.
  6. Assemble — editor or ffmpeg concat; color match lightly.
  7. Progressive continuity — episode N+1 reuses end-state wardrobe/injury/location from episode N’s final still as a new ref.
Episode progression example
Ep1 refs: mira-master, mira-trench, archive-hall
Ep1 end still: mira with torn sleeve, holding brass key

Ep2 refs: mira-master, mira-torn-sleeve-still, brass-key, night-market
Prompt: Mira sprints through a night market clutching the brass key.
Neon bokeh, handheld chase energy, same face and hair, torn sleeve visible.
Do not heal the sleeve; do not change pendant.

Carrying state forward (wardrobe, wounds, props)

StateHow to carry it
Outfit damageExport last frame → add as ref; freeze torn regions in prompt
New propStill of prop alone + character ref; show grip contact
Time of dayLocation ref + lighting words; keep character refs
Aging / arcSubtle stepwise edits to master between arcs, not jumps
Second characterSeparate ref pack; multi-ref with both faces

Multi-reference strategy (up to 7)

  • Ref 1: face close-up (identity)
  • Ref 2: full-body outfit
  • Ref 3: location plate
  • Ref 4: critical prop
  • Ref 5: secondary character face (if any)
  • Ref 6: style frame (film stock / grade) if needed
  • Ref 7: previous episode end state

Do not upload seven near-duplicate faces — you waste slots. Diversity of what is locked beats quantity of the same portrait.

Resolution, aspect, duration

  • 1080p when available for finals; 720p for drafts.
  • Match aspect to platform: 16:9 YouTube, 9:16 stories/reels, 1:1 feeds.
  • Prefer 6s shots; chain many. Longer clips only when motion stays simple.

API sketch (automation)

Conceptual API usage
# Pseudocode — check current xAI API docs for exact fields
client.video.generate(
    model="grok-imagine-video-1.5",
    prompt="...",
    reference_image_urls=[face, outfit, location],
    duration=6,
    aspect_ratio="16:9",
    resolution="1080p",
)

CLI vs web for video

  • Web / grok.com/imagine — fastest interactive refs + voice tools.
  • Grok Build CLI — great when video is part of a larger repo (marketing site embeds, game trailers, ffmpeg assembly scripts). Use /imagine-video when exposed; otherwise agent tools.
  • Build Mode apps — generate trailers that ship inside the product you are building.

Failure modes & fixes

SymptomFix
Face morphs mid-seriesStronger face ref; shorter freeze list; less outfit change in same shot
Outfit driftsFull-body outfit ref + name colors/materials every time
Busy warpingSimplify source still; move camera only; shorten action
Voice mismatchReuse one voice ref; avoid noisy samples
Lip sync softCloser shot, slower speech, shorter line
Style jumpsAdd a grade/style ref; keep same lighting words
Tip

Pair this with game design for mascot trailers, or websites for hero background loops (keep files light on Hostinger — poster frame + short muted loop).