ResourcesBy Baptiste O'Donovan3 min
The One Prompt That Got Claude to Animate My Whole Video
Resources
- Claude Code, runs the prompt, with Opus 5.5 on max reasoning
- Remotion, free for individuals and companies of up to 3 employees
- whisper.cpp, free, MIT licensed, transcribes on your own machine
- FFmpeg, free, handles the audio and video conversion
Everything moving in the top half of this reel was animated by Claude. There was no After Effects and no motion designer. Just one prompt.
It ran in Claude Code with Opus 5.5 on max reasoning. It took about 30 minutes, and it cost nothing on top of a Claude subscription. Claude wrote every frame as React code in Remotion, then timed each scene to the words in the voice track.
The full prompt is at the bottom of this page. It's the one people asked for when they commented DESIGN. Fill in your own video, script and beats, and run it. First, here's what each step of it does and why it's there.
What does each step of the prompt do?
Probe the footage: catch the format that breaks renders
iPhones and OBS on a Mac often record in HEVC. Headless Chrome can't decode it, so every Remotion still hangs until you convert a copy to H.264. The prompt makes Claude check this first, and check which Whisper you already have before it installs another.
Transcribe with word timestamps: where the timing comes from
Every cut in the animation lands on a spoken word, so the word timings have to be right. Claude transcribes each speech region on its own, because one pass over the whole file drifts by a second or more. It also checks that the last word sits near the end of the file. A cut-off transcript looks fine and quietly breaks every scene after it.
Make the spoken cut: keep the best take, clean the audio
Claude picks the best complete take of each line, trims dead air and only cuts inside silence. One constant speed-up covers the whole cut, never single lines, so your face never moves in slow motion. It also tracks your face, so you stay centred in the bottom half when you drift.
Design the animation: show it, don't label it
Scenes point at word positions in the transcript, not frame numbers, so a retime can't break the sync. Numbers count up and land on the frame you say them. If you talk about your voice, the waveform comes from your real audio. A box with a sentence in it is the exact failure the prompt warns against.
Check the Instagram safe zone: measure it, don't eyeball it
Instagram covers the top and bottom of a reel with its own buttons and captions. The prompt sets hard pixel limits for text and makes Claude measure rendered stills against them. Remotion Studio shows the full frame with none of the app on top, so eyeballing it misses the problems.
Verify before it says done: stills, overlaps and timing
Claude renders 40 to 60 stills across the video, looks at them and fixes what's wrong. It checks that no two graphics overlap, that captions run to the last word, and that key beats land on their word. It also logs how long the run took.
What's the full prompt?
Swap the four placeholders in angle brackets for your video path, your script, your beat table and your brand. Claude builds the project and starts Remotion Studio. You render the final video from Studio yourself.
You are going to animate a short vertical video entirely as code, in Remotion. I recorded myself talking. You make
everything the viewer sees above my head, timed word by word to my voice, in one continuous run. The video's claim is
that an AI animated it, so it must be true: no stock footage, no clip art, no templates, and nothing I touch up by hand.
Every frame is React code you write.
## Inputs
- VIDEO: `<VIDEO_PATH>` (a talking-head recording; it may contain retakes, false starts and dead air). READ ONLY.
Never modify, move or delete it.
- SCRIPT (what I meant to say; the recording is the truth where they differ):
```
<SCRIPT>
```
- BEAT TABLE (what should animate on each line; treat it as direction, and improve on it where you can):
| Line | What should animate |
|---|---|
<BEATS, e.g.
| "Everything above my head was animated by Claude" | Kinetic type: CLAUDE builds itself letter by letter |
| "When I say three, you get three" | A big 3, then three shapes drop in on the word "three" |
| "When I stop talking, it stops too" | Everything freezes mid-motion through the silence |>
- BRAND: `<BRAND: background, surface, text and accent colours, typeface; or "choose a look that suits it">`
## Output
One Remotion composition, 1080x1920 (9:16) at 30fps:
- **Top half (1080x960):** the animation, which is your work.
- **Bottom half (1080x960):** me, cropped from the recording, face centred.
- **Small word-level captions:** 1-2 words at a time, current word highlighted, sitting just above my head so they
never fight the animation.
Build it inside a fresh Remotion project in this folder (install from npm, pin the version, don't clone the repo).
Start Remotion Studio as soon as the project exists and give me the URL. **Do not render the full video**; I render
it from Studio. You may render single stills and very short ranges to check your work.
## Step 1: Probe the footage
Print duration, resolution, fps, video codec and audio sample rate with ffprobe.
- **If the video is HEVC/H.265** (iPhone and OBS on a Mac often are), transcode a working copy to H.264 inside the
project. Headless Chrome cannot decode HEVC, and every Remotion still will hang on `delayRender` until you do.
- **Whisper:** check what this machine already has (whisper.cpp `whisper-cli`, openai-whisper, faster-whisper) before
installing anything. Transcription must run locally.
## Step 2: Transcribe with word-level timestamps, and prove the timings
- **Word timings:** get per-WORD timestamps, not per-segment ones.
- **Transcribe per region:** split the file into speech regions first (a loudness/RMS scan), then transcribe each
region separately and offset its times. A whole-file pass drifts by a second or more.
- **Coverage:** the last word's time must be near the end of the file. A truncated transcript looks valid and
silently ruins every cut after it.
- **Time base:** if the tool uses voice-activity detection, its times may have the silence squeezed out. Spot-check
3-4 words against the audio's loudness envelope before trusting any of it.
- **Beat words:** word boundaries are estimates. For any word a beat lands on, find its real onset in the waveform.
- **Misheard names:** fix product and model names, e.g. "cloud" -> "Claude".
## Step 3: Make the spoken cut
- **Takes:** pick the best COMPLETE take of each line and drop retakes, false starts and stray words. Prefer one
continuous take over splicing fragments.
- **Dead air:** tighten it between phrases, but keep any pause the script depends on (e.g. the silence after "it
stops too").
- **Cut points:** cut only inside silence. Before you finalize each cut point, check that it doesn't clip the first
or last word of a line. Transcribe the first and last second of every kept segment and compare with the script. A
clipped "Stop" or "So" at the head of a line is the most common defect.
- **Speed:** apply ONE constant speed to the whole cut (1.0-1.15x; pitch preserved, e.g. ffmpeg `atempo`). Never
speed or slow individual lines: a slowed line puts the face in visible slow motion.
- **Voice:** master it lightly: high-pass ~75 Hz, gentle denoise, de-ess, small presence boost, 3:1 compression,
limiter at -1.2 dB.
- **Face and voice files:** output the face video (cropped for the bottom half) and the voice as WAV, plus a
`words.json` with every word's time on the NEW timeline.
- **Framing:** track my face across the take (e.g. OpenCV face detection every 0.25-0.5s, smoothed over ~1s). I will
drift left and right, so keep the face centred in the bottom box with a slight zoom (~1.1x) and a slow tracking pan.
A single fixed crop leaves me off-centre.
## Step 4: Design the animation (this is the real job)
- **Every scene change lands on a spoken word:** timeline.ts refers to word INDEXES from words.json, never fixed
frame numbers, so a retime can't break the sync.
- **Show, don't label:** a box with a sentence in it is the failure mode. Numbers count up and land on the frame the
number is said. Lists build item by item on each item's word. Callbacks reuse the same visual object.
- **Restraint:** some lines need nothing, or just the ambient motion.
- **Real data where the script refers to it:**
- If I talk about my voice, draw the waveform from the actual audio (RMS envelope of voice.wav), with words popping
off it as they're spoken.
- If I talk about code, show real code from this project.
- Numbers I say on camera (time taken, cost) are shown exactly as I say them.
- **Freezes:** if the script says the animation stops, implement it as one function: from that word until the next
word, every layer renders from the same frozen frame. Check two stills in the pause are pixel-identical.
- **Hook:** no title card. The animation IS the hook, so the first second must already be moving.
- **Ending:** a clean end card for the call to action, e.g. the comment keyword typed into a comment box.
- **Craft:**
- Kinetic type (big condensed words that build, break and get struck through) as the base.
- Geometric motion (grids, bouncing shapes, tiles) between type beats.
- Overshoot-and-settle entrances.
- Nothing sits still for long, and nothing is left on screen after its moment.
- **Components:** build one real component per idea, not one generic card fed different strings.
## Step 5: Instagram safe zone (hard rule, measured on rendered frames)
- **Readable area:** at 1080x1920, nothing readable above y=220 (app chrome) or below y=1470 (caption, username and
buttons), or outside x 35-980.
- **Centred rows:** a row centred on the frame must be <=880px wide, because the right-hand button rail is wider than
the left margin.
- **Bleed:** video and background may bleed past these lines; text may not.
- **Check the pixels:** render stills, find the non-background pixel bands, and print their extents against these
numbers. Studio draws the full frame with no app chrome, so eyeballing it will miss violations.
## Step 6: Verify before you say it's done
- **Stills:** render ~40-60 stills across the whole timeline, look at them (as contact sheets), and fix what you see:
overlaps, clipping, blank frames at cuts, text off the safe area.
- **Overlaps:** sort every graphic's start/end and diff neighbours; two graphics over the same seconds is the most
common bug.
- **Captions:** they must run to the last spoken word.
- **Beat sync:** confirm key beats land on their words by rendering the still AT the word's frame.
- **Media decoding:** use `<OffthreadVideo>` for video layers so renders are frame-exact.
- **Timing:** record your start and end time (`date`) and report the elapsed wall-clock time; it goes in the video.
- **Report:**
- The Studio URL and composition id.
- The final duration.
- The beat-by-beat list of what you built.
- What you verified, and anything uncertain.
- **Save this brief:** save the exact brief you worked from as PROMPT.md in the project.Motion design is one job Claude can take off your plate. To find the ones costing your team the most time, take the free AI bottleneck test, or read our insights on AI implementation.
