Hands-on · 2026-10-07

AI Film 101: How to Make a Realistic AI Video

Which model does which job, the four steps in order, and the prompts to paste - from my video How to Survive the AI Apocalypse.

Every prompt, every model and the 4 steps behind my video How to Survive the AI Apocalypse. Copy the prompts. Swap in your own face and places.

What I usedSix jobs, six tools

Me, on screen
3 stills cut from a 15 second phone clip
Reference sheets
GPT Image 2.5, inside ChatGPT
Places and the look
Nano Banana Pro
Video clips
Seedance 2.5, in draft mode
Final sharpness
One upscale at the end
Voice to text
Parakeet and Whisper, on a laptop

Seedance or Omni? Seedance makes a new clip. Gemini Omni changes a clip you already have. More on Omni here.

The recipe
4 steps, in this order
1. Put yourself in it. 2. Lock the look with stills. 3. Write the video prompt like a bystander filmed it. 4. Draft first, upscale last.

Step 1 of 4Put yourself in it

Film yourself for 15 seconds on your phone. Move around.

Give ChatGPT the clip and one example character sheet. Ask for a sheet of you like it. It pulls frames from the clip and builds the sheet.

Four frames from my phone clip next to the character sheet ChatGPT made from them
Left: the four frames ChatGPT pulled from my clip. Right: the sheet it made.

Cut 3 stills from the same clip. Face, waist up, full body. They go into every video run, so the person on screen is really you.

For your voice, give the video model 10 seconds of you talking. Add this line: "Take nothing visual from the reference video."

Three stills of me, one frame of the voice clip and one frame of the opening clip
The three stills, a frame of the voice clip, a frame of the opening clip.

Watch for: a big grid of poses loses the face.

Step 2 of 4Lock the look with stills

Make the places as pictures before any video. I put all four places in one picture, 2 by 2. Every video run gets it, so the places stop changing.

Test one prompt on two models first. Same prompt, same 4 reference pictures. Nano Banana Pro won for the places.

The same prompt on GPT Image 2.5 and on Nano Banana Pro
Same prompt, same references. Left: GPT Image 2.5. Right: Nano Banana Pro.

Paste my image prompt into your own. It names the colours I did not want and asks for worn rooms.

Open the image prompt
BOTH NEW FRAMES match the top row: same overcast morning, same damp weather, same cool green-grey grade (moss, sage, olive, grey-teal), things keep their own real colors only slightly muted. NOT golden, NOT yellow, NOT a blue tint. Exactly zero people and nothing shaped like a standing person. No text, numbers, logos or readable signs. No weapons, no smoke, no fire, no nuclear power station, no cooling towers, nothing from the 1940s.

Photographic realism above everything: unedited documentary photographs taken this week in real rooms with a real camera. Shallow focus falloff, motivated window light mixed with one warm lamp, fine sensor grain, real worn textures: scuffed paint, creased sheets, carpet wear, fingerprints on the laptop, a ring from a mug. No CGI, no 3D render, no concept art, no clean showroom surfaces.

Watch for: worn rooms read as photographs. Clean empty rooms read as 3D renders.

Step 3 of 4Write the video prompt like a bystander filmed it

Rule 1: the first line names who is filming. Film language gave me a movie trailer. A bystander with a phone gave me real footage.

Raw amateur vertical phone video, 9:16, 15 seconds, three shots, filmed by a bystander who quietly follows along a few steps behind him and to one side, phone held upright. The picture is the right way up and level for the whole clip: sky at the top of the frame, the ground flat along the bottom, the horizon level, no tilt, no rotated frame. Only the scene in front of the lens is ever in the picture. Unedited real footage: the picture bobs with each step, slightly shaky hands, motion blur, autofocus hunting once, natural light, phone-camera noise, no colour grading, no music, no on-screen text, no subtitles, no logos. Ordinary overcast morning, flat grey daylight, green grass on the dunes.

Rule 2: keep the camera out of the picture. This line goes in every prompt: "Only the scene in front of the lens is ever in the picture."

Rule 3: time the slow motion in seconds. The word "freezes" gave me no stop at all.

Open the slow motion prompt
SHOT 1, 0 to 5 seconds, the beach in @strand. Rust spots on the armoured trucks, a torn windsock on the pier.
0 to 1.2s, real speed: @roy sprints along the wet sand with a home-built exo rig strapped over his shirt: scratched grey steel struts on his arms and legs, ratchet straps, silver tape. The camera keeps pace beside him at chest height, his whole body in the picture. One walking machine closes in from the far side in long heavy strides. Six small quadcopter drones fly low overhead in one line.
1.2 to 3s, extreme slow motion: the machine drops its shoulder and drives it into his chest. The steel struts of his rig bend, one strap snaps and whips loose, a ring of wet sand and sea water bursts up from under his feet. Both his feet leave the ground. His head snaps back, his arms fly wide, and he is thrown up and backward through the air, face up to the sky.
3 to 4.4s: time stops. He hangs in mid-air one metre above the sand, completely still, sand grains and drops of water hanging around him, the machine stopped mid-stride, the drones stopped in the sky with their rotors not turning. Nothing in the scene moves. Only the camera still moves: the person filming walks slowly around him in a half circle and ends close on his face. No sound, only a low held tone.
4.4 to 5s: the whole shot snaps into fast reverse.
His voice over the picture, calm and dry, he does not speak on screen: "How to survive the AI apocalypse."

Rule 4: describe each thing once and paste it word for word. Give it its own reference picture too.

The machines are the same every time: a two-legged walking machine as tall as a man, matte dark grey body panels, exposed black joints, a glossy black faceplate, no markings. Heavy on its feet, every footfall sinks into the sand.

Rule 5: worn objects, plain light, real sound. Rust spots, a torn windsock, "ordinary overcast morning". No music.

Rule 6: short actions, in order. One per line, three or four per shot.

The whole prompt. Six shots in 15 seconds. Words with an @ are my reference pictures. Swap in your own.

Open the whole prompt
Raw amateur vertical phone video, 9:16, 15 seconds, six quick shots with hard cuts, about two seconds each, filmed by a bystander who quietly follows along a few steps from him, phone held upright. The picture is the right way up and level for the whole clip: sky at the top of the frame, the ground flat along the bottom, the horizon level, no tilt, no rotated frame. Only the scene in front of the lens is ever in the picture. Unedited real footage: slightly shaky hands, motion blur, autofocus hunting once, natural light, phone-camera noise, no colour grading, no music, no on-screen text, no subtitles, no logos. Ordinary overcast morning, flat grey daylight, green grass on the dunes. Fast pace: he keeps moving and talks quickly, no pauses.

The man is a real man, @roy (the same man shown in @royupper and @roybody): same face, messy dark brown hair, short beard and moustache, slim build, no hat, wearing his green patchwork work shirt with beige shoulder and sleeve panels and grey shorts, barefoot. Real skin, sand on his shins. He is dry and a little tired, sure of what he is saying, he does not smile.

The machines are the same every time, exactly the one in @machine: a four-legged walking machine the size of a large dog, matte dark grey body panels, exposed black joints, a boxy sensor head, no markings, no lights. In this clip four of them stand in a row on the beach in @strand, two steps apart, standing still, calm. They are identical except for one piece of plain grey gear each:
Machine one: one big round glass eye, a camera lens the size of a fist, on the front of its head.
Machine two: one long grey furry microphone, as long as a forearm, standing up from the top of its head.
Machine three: one flat screen the size of a paperback book on its side.
Machine four: all three: the glass eye, the furry microphone and the screen.

SHOT 1, 0 to 2.5 seconds, the wooden boardwalk steps down the dune in @strand. Filmed from the bottom of the steps, looking up at him.
@roy comes down the steps fast, toward the camera.
He swings a worn, patched olive canvas backpack onto his right shoulder.
He looks straight into the camera and says: "Rule two: use that to your advantage."

SHOT 2, 2.5 to 4.5 seconds. Hard cut. Close on machine one, he is crouched right beside its head.
He waves his hand in front of the glass eye and it turns to follow his hand.
He says to the camera: "Some AI can see."

SHOT 3, 4.5 to 6.5 seconds. Hard cut. Close on machine two, he is right beside it.
He snaps his fingers and the furry microphone swings toward the snap.
He says to the camera: "Some can hear."

SHOT 4, 6.5 to 9 seconds. Hard cut. Close on the screen on machine three, he is right beside it.
The screen shows a bright sunny moving picture of this same man surfing a blue wave. No words on the screen.
He taps the screen with one finger and says to the camera: "Some make pictures and video."

SHOT 5, 9 to 11 seconds. Hard cut. A little wider: machine four from head to feet, he stands right behind it.
All at once its glass eye turns to him, its microphone swings to him and its screen shows the surfing picture.
He points down at it with both hands and says: "Some do it all."

SHOT 6, 11 to 15 seconds. Hard cut. Filmed from two steps away, at his chest height, he and machine four both in the picture.
He twists the glass eye a quarter turn and it comes off in his hand with a click. The machine stays calm.
He holds it up beside his face, looks straight into the camera and says: "Know which is which, and you can build your own tools."
He drops it into the backpack.
Plain grey cloud above: no beams, no glow, no sparks, no smoke.

Real sounds only: wind, surf, bare feet on wood and sand, a faint whir when a machine head turns, one finger snap, one metal click. The reference video is only a sample of his real voice: in every line he sounds exactly like the man speaking in the reference video - the same voice, pitch, accent and rhythm, the same voice in all six shots. He says every line on camera and his lips move with the words. He says only the six lines in quotes, 34 words, and nothing else. Take nothing visual from the reference video: no cap, no grey t-shirt, no room, no framing.

Step 4 of 4Draft first, upscale last

Run every try in Seedance draft mode at 480p. It comes back faster, so you get more tries. A draft checks the action, the lines and the order.

Judge the look on a full quality run. Then upscale once, at the end.

The idea behind the videoLocal AI, cloud AI, and building your own

Local AI runs on your own device. Fast, private, works with the wifi off. These run on a laptop and are worth a look. How to set up Whisper.

Voice to text
Parakeet in 714 MB. Whisper large v3 turbo in 1.6 GB
Answers questions
Qwen 2.5 0.5B in 397 MB
Phone-sized
Gemma 4 E2B fits in 1 GB of memory

Cloud AI runs in the big data centers. Way more power.

The hard problems
Claude Opus 5.5 and GPT-6 Astra
Pictures
GPT Image 2.5 Flare for speed, Sunburst for precise edits
Video with sound
Seedance 2.5

Route it. A small model takes the easy stuff. The big one takes the hard stuff. That is the tool I build in the video. New kinds keep arriving too, like Meta's V-JEPA.

So build your own. With Opus 5.5 and GPT-6 Astra you can build like never before. Build with it and it stops being scary.

It's not about perfection, it's about iteration. Version one lasted nine seconds. That's not failing, that's learning.Rule three
Coming next

Get the next one in your inbox.

I write about the AI tools I'm testing, what just shipped, and what's actually worth your time. Weekly, no fluff.

Subscribe to the newsletter