Imagine Video 1.5
Video creation & character continuity
Every practical way to prompt scenarios, lock a character across clips, keep voice consistent, and build progressive multi-part stories.
Imagine Video 1.5 with References is the leap: text-to-video, multi image references, voice consistency, and native 1080p. SuperGrok Heavy / Plus get priority access; features continue to roll out by tier and region. This chapter is the production playbook.
Generation modes (know which dial you are turning)
| Mode | Inputs | Best for |
|---|---|---|
| Text-to-video (T2V) | Prompt only | New shots, establishing scenes, concepting |
| Image-to-video (I2V) | Still + motion prompt | Exact first frame control; animate a designed shot |
| Reference-to-video (R2V) | 1–7 refs + prompt | Same face/product/location in a new scene |
| Image + voice reference | Character image + voice sample | Dialogue / singing with locked timbre |
Durations are short by design (commonly 6s; longer tiers where available). Professional workflows are many short shots, not one endless take. Assemble with an editor or ffmpeg.
Prompt craft for different scenarios
Write present tense, one clear action, one camera move. Overloading motion causes mush.
Cinematic establishing
Dawn over a rain-slick neo-Tokyo alley, neon reflecting in puddles.
Slow aerial push-in toward a lone figure under a red umbrella.
Cinematic anamorphic bokeh, grounded physics, subtle ambient city hum.Product hero
A matte-black wireless earbud rotates slowly on a glass pedestal.
Soft studio key light, gentle camera orbit, crisp reflections,
premium commercial look, no text overlays.Action beat
The courier vaults a rooftop gap at night, coat flaring.
Handheld medium shot tracking beside her, light motion blur,
rain streaks, tense electronic underscore. Single jump only.Dialogue / performance
Close-up: the captain speaks calmly to camera on a starship bridge.
Subtle head movement, soft blinks, stable framing, shallow depth of field.
Keep lips natural; no rapid cuts.Music / singing
The same woman sings a single sustained chorus line on a dim stage.
Microphone in hand, soft backlight haze, locked face and voice reference,
camera slow dolly-in, no choreography changes.Game-style / Minecraft-ish adventure
Blocky explorer and a wolf crest a snowy hill at golden hour.
Gentle side-tracking shot, coherent character proportions,
charming adventure mood, stable silhouette.Step 1 — Build a character bible (do this once)
Continuity fails when you re-describe a person from memory every time. Create a master reference pack:
- Master portrait — neutral light, clear face, signature outfit, plain background.
- Turnarounds — front / 3⁄4 / side (edit from the master, do not re-roll from scratch).
- Expression set — calm, smile, determined (same lighting).
- Outfit card — bullet list of colors, materials, props.
- Voice sample — clean line reading for voice lock (when available on your tier).
Studio portrait of Mira Chen, 28, sharp emerald eyes, defined jaw,
messy auburn hair in a loose bun, small scar above left eyebrow,
tailored navy trench over white silk blouse, thin silver pendant.
Neutral gray background, Rembrandt lighting, 85mm, photoreal,
identity-defining details sharp.Step 2 — Lock identity across videos
Method A — Multi image references (primary)
Upload up to seven reference images. Each ref should lock one thing: face, full body outfit, location, key prop. Prompt the change (action/scene), not a full redescription of the face.
Using the character references for face and outfit:
Mira walks through a sunlit archive hall, dust in light shafts.
Medium shot, steady steadicam follow, calm confidence.
Keep exact face geometry, hair, trench coat, and pendant.
Change only location and walk cycle.Method B — Image-to-video from a composed still
For maximum first-frame control: generate/edit a perfect still (with character refs), then animate that still. Best for complex composition.
Animate: slow push-in, she turns her head slightly toward camera,
coat fabric settles, dust motes drift. No new characters, no outfit change.Method C — Voice + face
Pass character image + voice reference so dialogue scenes keep both likeness and timbre. Reuse the same voice pack for the whole series.
Method D — Edit chain for stills before video
Never re-generate the character from text alone for episode 2. Always image_edit / reference from the master: “Keep this exact character — change only the background to a night market.”
In every continuity prompt, list what must not change first (face, hair, scars, outfit colors, body proportions), then the single change. Models drift toward average faces without a freeze list.
Progressive multi-video series (episode pipeline)
Think like an episodic director. Each clip is one beat; the series is assembled later.
- Series bible — logline, 3-act beats, recurring locations, cast list with ref packs.
- Shot list — rows: shot id, duration, refs used, camera, action, dialogue.
- Still pass — keyframes for hard shots (fight, reveal, product insert).
- Animate pass — T2V/I2V/R2V per shot with the same refs.
- Select — keep only takes that hold identity; regenerate failures with tighter freeze lists.
- Assemble — editor or ffmpeg concat; color match lightly.
- Progressive continuity — episode N+1 reuses end-state wardrobe/injury/location from episode N’s final still as a new ref.
Ep1 refs: mira-master, mira-trench, archive-hall
Ep1 end still: mira with torn sleeve, holding brass key
Ep2 refs: mira-master, mira-torn-sleeve-still, brass-key, night-market
Prompt: Mira sprints through a night market clutching the brass key.
Neon bokeh, handheld chase energy, same face and hair, torn sleeve visible.
Do not heal the sleeve; do not change pendant.Carrying state forward (wardrobe, wounds, props)
| State | How to carry it |
|---|---|
| Outfit damage | Export last frame → add as ref; freeze torn regions in prompt |
| New prop | Still of prop alone + character ref; show grip contact |
| Time of day | Location ref + lighting words; keep character refs |
| Aging / arc | Subtle stepwise edits to master between arcs, not jumps |
| Second character | Separate ref pack; multi-ref with both faces |
Multi-reference strategy (up to 7)
- Ref 1: face close-up (identity)
- Ref 2: full-body outfit
- Ref 3: location plate
- Ref 4: critical prop
- Ref 5: secondary character face (if any)
- Ref 6: style frame (film stock / grade) if needed
- Ref 7: previous episode end state
Do not upload seven near-duplicate faces — you waste slots. Diversity of what is locked beats quantity of the same portrait.
Resolution, aspect, duration
- 1080p when available for finals; 720p for drafts.
- Match aspect to platform: 16:9 YouTube, 9:16 stories/reels, 1:1 feeds.
- Prefer 6s shots; chain many. Longer clips only when motion stays simple.
API sketch (automation)
# Pseudocode — check current xAI API docs for exact fields
client.video.generate(
model="grok-imagine-video-1.5",
prompt="...",
reference_image_urls=[face, outfit, location],
duration=6,
aspect_ratio="16:9",
resolution="1080p",
)CLI vs web for video
- Web / grok.com/imagine — fastest interactive refs + voice tools.
- Grok Build CLI — great when video is part of a larger repo (marketing site embeds, game trailers, ffmpeg assembly scripts). Use
/imagine-videowhen exposed; otherwise agent tools. - Build Mode apps — generate trailers that ship inside the product you are building.
Failure modes & fixes
| Symptom | Fix |
|---|---|
| Face morphs mid-series | Stronger face ref; shorter freeze list; less outfit change in same shot |
| Outfit drifts | Full-body outfit ref + name colors/materials every time |
| Busy warping | Simplify source still; move camera only; shorten action |
| Voice mismatch | Reuse one voice ref; avoid noisy samples |
| Lip sync soft | Closer shot, slower speech, shorter line |
| Style jumps | Add a grade/style ref; keep same lighting words |
Pair this with game design for mascot trailers, or websites for hero background loops (keep files light on Hostinger — poster frame + short muted loop).