Here's the mistake almost everyone makes on their first AI video.
They open a model, paste in their story, and type: "Here is my story. Make me a video."
What comes back is technically a video. The character's face changes between shots. The clothing changes. The camera does something arbitrary. The lighting resets. Ten seconds in, the audience has stopped believing any of it is the same person in the same world.
The problem isn't the model. The problem is that a story is not a production plan.
Creating a compelling AI video is not simply about writing a good prompt. The real workflow looks like this:
→ Raw story
→ Story refinement
→ Episode structure
→ Scene breakdown
→ Character and world bible
→ Image generation
→ Video generation
→ Frame-to-frame continuity
→ Voiceover and sound
→ Final edit
Ten stages. The prompt is one of them. This guide walks the whole pipeline — the same 37 steps as the downloadable PDF — using a single running example, so you can watch a rough idea become a shot list become a finished episode.
Phase 1 — Start With the Base Story
Your starting point can be almost anything: a short story, a novel chapter, a personal experience, a movie idea, a historical event, a children's story, a fantasy concept, a Reddit-style thread, a voice note, a rough plot outline, even a few paragraphs of ideas.
The first step is not image generation. First, understand the story.
Suppose your raw story says:
"I found an old trunk in my grandfather's farmhouse. Inside was a strange staff and an ancient book. When I touched the staff, I could suddenly read the book, and something changed inside me."
That's a story idea. It is not yet a video. It has to be transformed into something cinematic before a single frame gets generated.
Rewrite the Story for Video
AI video needs visual moments. A written story can spend three paragraphs on what a character thinks. A video cannot.
Instead of: "He became increasingly suspicious about what was inside the room."
Show:
→ He stops walking
→ His flashlight moves across the wall
→ He notices the doorway
→ He hesitates
→ He enters
→ His flashlight reveals the trunk
→ His expression changes
The rule: if the audience can see it, don't explain it. Turn exposition into action, expression, and visual storytelling.
Make the Narration Sound Human
A common failure is narration that reads like a screenplay rather than a person remembering something.
Weak: "I was eighteen. I lived on an old sailboat in a boatyard."
Technically correct. Completely unnatural.
Better: "Back then, my life was pretty simple. I was eighteen, finishing up the last few weeks of school, and living on my grandfather's old sailboat."
The second version feels like someone telling you what happened. Narration should sound like "this is what happened to me," not "here is the information you need to know."
Create the Episode Structure
If your story is long, don't try to turn the entire thing into one video. Divide it into episodes.
A strong episode generally has five beats:
→ Hook — something immediately interesting
→ Setup — introduce the character and situation
→ Escalation — something strange, dangerous, or unexpected happens
→ Payoff — the central event of the episode
→ Cliffhanger — a reason to watch the next one
For the trunk story, that expands into a nine-beat structure: the protagonist discovers a hidden room (hook), we learn about his unusual but peaceful life (setup), he finds an ancient trunk (discovery), inside are a staff and an impossible book (mystery), the staff activates (supernatural event), his body begins changing (transformation), he understands something has happened to him (realization), he discovers he can make objects disappear (new power), and he hears breaking news connected to what happened (cliffhanger).
Break the Episode Into Scenes
Don't think "I need one video." Think "I need a sequence of cinematic shots."
Episode 1 breaks into eleven scenes:
→ Scene 1 — Discovering the hidden room
→ Scene 2 — Life before the discovery
→ Scene 3 — Getting the trunk out
→ Scene 4 — Opening the trunk
→ Scene 5 — Discovering the staff and book
→ Scene 6 — Discovering the language
→ Scene 7 — The book takes control
→ Scene 8 — Transformation
→ Scene 9 — The following day
→ Scene 10 — Returning home
→ Scene 11 — Breaking news
Once the episode exists as a list of scenes, AI generation becomes dramatically easier. You are no longer asking a model to invent a story. You are asking it to render a shot you have already specified.
Every Scene Needs a Purpose
Before generating an image, ask four questions:
→ What happens? The physical action.
→ What does the character feel? The emotion.
→ What does the audience discover? The information.
→ What changes by the end? The story progression.
If nothing changes, the scene probably doesn't need to exist.
Phase 2 — Build a Character Bible
This is one of the most important steps in AI filmmaking, and the one most often skipped.
AI models don't inherently remember your character. If you generate "18-year-old boy" over and over, you will get a different person every time.
So define the character once, in full, and reuse the definition:
MAIN CHARACTER
Age: 18
Gender: Male
Skin: medium-brown
Build: lean athletic
Height: medium-tall
Face: oval
Eyes: deep-set brown
Eyebrows: thick
Hair: short, slightly wavy black hair
Facial hair: clean-shaven
Personality: intelligent, curious, introspective
Clothing: dark charcoal T-shirt, faded blue jeans, worn brown boots
That block becomes your character identity specification. It goes into every prompt, unchanged.
Build a World Bible Too
Do the same for important locations and objects.
Farmhouse — large and old, weathered wooden exterior, rural isolation, aging but maintained, screened porch, old basement, wooden interior.
Sailboat — 50 feet, old but beautiful, wooden interior, cozy cabin, small kitchen, quiet boatyard.
Trunk — seven feet long, three feet deep, ancient dark wood, heavy construction, wooden latch, no visible lock, unusual wooden hinges.
Staff — six feet long, dark natural wood, intricate wooden lattice, milky-white crystal, subtle supernatural glow.
Book — large and heavy, dark ancient cover, thin metallic pages, metallic page edges, strange multilingual writing.
These details should never randomly change. The bible is the contract between you and the model.
Phase 3 — Generate Your Reference Images First
Don't immediately ask for video. Create high-quality stills first.
Images give you control over character appearance, clothing, environment, camera angle, lighting, objects, and composition. Once you have the correct image, you can animate it. Once you have animated a wrong image, you have burned a generation.
The workflow: story → scene description → image generation → approve the image → video generation.
How to Write Image Prompts
A strong image prompt describes eight things:
→ Subject — who or what is visible
→ Environment — where they are
→ Action — what they are doing
→ Emotion — what they are feeling
→ Camera — what type of shot
→ Lighting — time of day and light source
→ Style — the visual aesthetic
→ Continuity — what must remain consistent
The formula:
[SUBJECT] + [ACTION] + [ENVIRONMENT] + [EMOTION] +
[CAMERA] + [LIGHTING] + [STYLE] + [CONTINUITY]
Don't Just Say "Cinematic"
"Cinematic" on its own is too vague to steer anything. Specify instead: photorealistic, 35mm film aesthetic, shallow depth of field, anamorphic composition, natural skin texture, atmospheric depth, realistic lighting, dramatic shadows, film grain, 16:9 widescreen.
The goal is to give the model a visual language, not a mood word.
Use a Global Style Prompt
Rather than reinventing the look every time, write one reusable block and prepend it to every scene:
Photorealistic cinematic fantasy thriller, grounded visual realism,
natural skin texture, realistic human proportions, atmospheric lighting,
subtle film grain, anamorphic cinematic composition, shallow depth of
field, dramatic shadows, muted earthy palette, high production value,
realistic practical lighting, 35mm film aesthetic, 16:9 widescreen.
Then add the scene-specific instructions underneath. This is what gives an episode a consistent visual identity.
A Real Prompt, Before and After
Instead of: "Boy finds an old trunk in a basement."
Use:
Photorealistic cinematic thriller scene inside an enormous old farmhouse basement. An 18-year-old young man with medium-brown skin, lean athletic build, oval face, deep-set brown eyes, thick dark eyebrows, short slightly wavy black hair and clean-shaven face stands cautiously inside a forgotten room, holding a flashlight toward an enormous seven-foot ancient wooden trunk. Dust particles float through the flashlight beam. The wooden walls are weathered and covered in cobwebs. The trunk is dark, massive and remarkably preserved despite its age. His expression shows cautious curiosity and uncertainty. Dramatic realistic flashlight illumination, deep shadows, atmospheric dust, cinematic 35mm photography, shallow depth of field, photorealistic fantasy thriller, 16:9.
Same scene. The second version gives the model enough information that it no longer has to guess — and guessing is where continuity dies.
Generate Multiple Shots Per Scene
A scene doesn't equal one image. "Discovering the trunk" is really five shots:
→ Shot A — Basement establishing shot
→ Shot B — Character entering the basement
→ Shot C — Discovering the hidden doorway
→ Shot D — First reveal of the trunk
→ Shot E — Close-up of the latch
Generating all five gives you editing flexibility later, which is worth far more than a single perfect frame.
Phase 4 — Move to Video Generation
Once the image is approved, take it to your video model. This is where most beginners make their biggest mistake. They write "animate this image."
That's not enough. You need to tell the model, explicitly, that this image is the starting frame.
The Single Most Important Instruction
Use language like this:
Use the supplied reference image as the exact opening frame of this video. Do not reinterpret, redesign, replace, or regenerate the starting composition. The first frame must visually match the reference image exactly. Preserve the character's face, clothing, body proportions, environment, objects, lighting and camera position. Begin motion from the exact state shown in the reference image.
Without it, the model may change the face, change the clothing, move objects, change the environment, change the camera, or start from an entirely different composition. This one paragraph is the difference between a film and a slideshow.
Tell the Model What Should Move
An image-to-video model needs to know what stays still, what moves, how fast, and how the camera moves.
Good: "The character slowly turns his head toward the doorway. His flashlight moves with his hand. Dust particles drift through the beam. The camera makes a slow push toward him."
Not: "Make the scene cinematic."
Don't Overload a Short Video
For a five- to ten-second generation, don't ask for twenty actions.
Bad: "He walks upstairs, opens the trunk, takes out the staff, reads the book, sees visions, becomes cold, turns to stone and disappears." That's several scenes, and the model will compress them into nonsense.
Better: "He slowly reaches toward the trunk's latch, hesitates, then lifts it. The camera pushes closer."
One action. One emotional beat.
Camera Movement Matters
You can direct the camera with plain cinematic language:
→ Slow push-in — camera gradually moves closer. Use for discoveries, realizations, emotional moments.
→ Pull-back — camera moves away. Use for revealing scale, isolation, endings.
→ Tracking shot — camera follows the character. Use for walking, running, exploration.
→ Orbit — camera slowly circles the character. Use for transformations, character reveals, supernatural moments.
→ Static — camera stays still. Use for suspense, dialogue, shock, horror.
Use Motion to Control Pacing
Slow movement creates mystery, suspense, emotion. Fast movement creates panic, action, urgency. Almost no movement creates unease, anticipation, horror.
For supernatural scenes, less is usually more.
Phase 5 — Continuity Is the Biggest Challenge
This is where AI video most often falls apart. Shot 1: the character has short black hair. Shot 2: suddenly longer hair. Shot 3: a different face. Shot 4: different clothing. Immersion is gone.
The fix is reference chaining.
The Reference-Chain Method
Instead of generating each shot independently from its own still:
Image 1 → Video 1
Image 2 → Video 2
Image 3 → Video 3
Chain them, so each shot inherits the visual state of the one before it:
Image 1 → Video 1
Final frame of 1 → Video 2
Final frame of 2 → Video 3
The final frame of each clip becomes the opening frame of the next. That creates a visual chain the model cannot drift away from.
Use Two References When Necessary
For important transitions, supply both the previous video's final frame and the next scene's reference image, then tell the model how to relate them:
The previous video's final frame is the immediate continuity frame. The supplied reference image is the intended visual destination. Preserve the character, environment and object continuity from both. Begin from the previous video's final frame and transition naturally toward the supplied reference composition.
This is particularly useful when changing camera angle, location, character position, or time of day.
The Ending Frame Rule
Every shot should have an intentional ending. Don't let a clip stop at random — decide what image the next shot needs to begin with.
→ Shot 1 ends: character looking at the hidden doorway → Shot 2 begins: character already standing at the doorway
→ Shot 2 ends: flashlight illuminates the trunk → Shot 3 begins: trunk fills the frame
That's continuity by design rather than by luck.
Write Scenes as Visual Handoffs
Every scene should specify four things:
→ Start — where does the shot begin?
→ Action — what happens?
→ End — what should the final frame look like?
→ Next — what does the next shot need to inherit?
This is one of the most powerful techniques in AI filmmaking, and it costs nothing but a few lines of planning.
Voiceover, Dialogue and Suspense
Voiceover Should Not Control the Visuals
Write the visual action first. Then write narration that supports it.
→ Visual: character discovers the trunk → Voiceover: "I wasn't looking for anything that night."
→ Visual: he touches the trunk → Voiceover: "I found it in a little side room off the basement."
The visuals carry the story. The narration adds context and emotion. Reverse that order and you get a documentary voiceover playing over stock footage.
Use Dialogue Sparingly
Dialogue is most effective when the character naturally reacts to something. Lines that feel real are short, instinctive, and grounded in the moment — not crafted to deliver information to the audience.
"Oh, come on…" — "Please don't fall through." — "What the hell is this?"
Don't turn every scene into exposition. Let reactions speak louder than explanations.
Build the Suspense
AI videos become far more engaging when information is revealed gradually. Don't show everything at once — reveal it layer by layer.
→ Conceal — hold back the full picture
→ Hint — drop clues that raise questions
→ Reveal — layer by layer, earn the payoff
The Question Chain
A good episode constantly makes the audience ask the next question: What's that? Why was it hidden? What's inside? What is the staff? Why is it warm? What is this book? What language is this? Why can he suddenly read it? What's happening to his body? What has he become? What's on the news?
That chain of unanswered questions is what keeps viewers watching.
Don't Explain the Mystery Too Early
Especially in fantasy and thriller.
Instead of: "The staff contains ancient magical energy that transforms people."
Show: the staff heats up. The crystal glows. His body changes. Then let the audience wonder what is happening. You can explain it later — or in episode three.
Cliffhangers
A good episode shouldn't simply stop. It should open a new question.
Weak ending: "He went home and went to sleep."
Strong ending: he is eating dinner → the television reports breaking news → his fork stops → he recognizes something → cut to black.
Now the audience needs episode two.
The Complete Production Pipeline
Assembled, the whole workflow is six phases:
→ Story — raw story, central conflict, cinematic rewrite, episode division
→ Pre-production — scene breakdown, character bible, location bible, prop bible, visual style
→ Image generation — create keyframes, concepts, and visual references
→ Video generation — turn approved visuals into motion scenes and shots
→ Continuity — keep characters, objects, and style consistent across episodes
→ Post-production — edit, refine, add sound, finalize the episode
Inside those phases the step sequence is: create scene prompts → generate images → select the best → fix continuity problems → write video prompts → set the image as the exact opening frame → generate a short video → check movement, character consistency and object consistency → save the final frame → use it as the next reference → generate the next shot → repeat. Then assemble clips, add voiceover, sound effects, background music, transitions, colour correction, and export.
The Two Prompt Templates
Image Prompt Template
CHARACTER [Character description from your bible]
LOCATION [Location description from your bible]
ACTION [What the character is doing]
EMOTION [What the character feels]
COMPOSITION [Wide shot / medium shot / close-up / overhead]
CAMERA [35mm / cinematic / low angle / eye level]
LIGHTING [Morning / sunset / moonlight / practical lighting]
VISUAL STYLE [Your global visual style block]
CONTINUITY [Elements that must remain identical]
FORMAT 16:9 widescreen, photorealistic cinematic film still.
Video Prompt Template
REFERENCE IMAGE — CRITICAL
Use the supplied reference image as the exact opening frame of this
video. Do not reinterpret, redesign, replace or regenerate the starting
composition.
Preserve the character's: face, hairstyle, skin tone, clothing,
body proportions.
Preserve: location, architecture, props, lighting, camera composition.
ACTION: [Describe only the movement that should happen.]
CHARACTER MOVEMENT: [Natural movement.]
CAMERA MOVEMENT: [Push-in / tracking / orbit / static.]
ENVIRONMENTAL MOVEMENT: [Wind / dust / water / clothing / background.]
EMOTION: [Subtle facial expression.]
ENDING FRAME: End with [specific visual state] so the next
shot can continue seamlessly.
Do not add subtitles, captions, logos or watermarks.
The Golden Rules of AI Filmmaking
→ Story first. Prompt second.
→ Never generate the entire film in one prompt.
→ Create your characters before creating your scenes.
→ Create reference images before video.
→ Treat every image as a production asset.
→ Tell the video model exactly what moves.
→ Control the camera.
→ End every shot intentionally.
→ Use the previous final frame for continuity.
→ Don't over-explain your mystery.
→ Make narration sound like a human telling a story.
→ One short video generation should contain one major action.
→ Consistency is more important than complexity.
→ Generate, review, correct, regenerate.
AI filmmaking is iterative. Your first generation doesn't have to be perfect — it has to be close enough to correct.
The Whole Workflow on One Line
Raw story → story refinement → episode breakdown → scene breakdown → character and world bible → visual style bible → image generation → approve image → video generation → check character and motion → save final frame → use final frame as reference → generate next shot → repeat → assemble → voiceover and sound design → music and editing → final cinematic video.
That's the pipeline. The prompt was never the hard part — the plan was.
