FreshStack

ResourcesBy Baptiste O'Donovan5 min

Opus 5.5 vs GPT-6 Astra: One Prompt, Two AI Video Edits


Resources

  • The AI video editing prompt, free PDF, the full prompt ready to copy
  • Claude Code, where Opus 5.5 ran the edit
  • ChatGPT, where GPT-6 Astra ran the edit
  • Remotion, free for individuals and companies of up to 3 employees
  • Whisper, free, MIT licensed, transcribes on your own machine
  • FFmpeg, free, handles the probing and the audio pass-through

Two AI models edited the same talking-head reel: Opus 5.5 in Claude, and GPT-6 Astra in ChatGPT. Same recording, same prompt, and nobody touched a frame of either result.

Every graphic in both edits was written as code in Remotion. Neither model got a script, a shot list or a storyboard. Each one had to work out from the audio alone what the viewer should see, and when to show nothing at all.

The exact prompt is a free PDF in the links above. It's the one people asked for when they commented CUT. Swap in your own recording and your own brand, and you can run the same test on your videos.

What was each line of the video testing?

The recording is the test. Every line forces an editing decision, and nothing in the prompt says what to do with any of them.

  • "Every graphic... was written as code": does the model show code or a build, or ignore the line?
  • "Almost 48,000": a number. A counter, a chart, or just the text?
  • "A picture is worth a thousand words": an idiom. Literal words on screen, or an actual picture?
  • "Three steps... the third step, you'll see": a list with a gap left on purpose. What does the model fill in?
  • "On this side... on this side": a gesture, pointing left then right. The prompt never mentions it.
  • "Most editors spend four hours": a time comparison. A clock, a timeline, or nothing?
  • "Automation. Remember that word": an anchor. Does the model mark it for later?
  • "Some lines don't need anything": restraint. Does it leave the frame clean?
  • "That word from earlier? Automation": a callback. Does it connect back to the anchor?

What does each step of the prompt do?

Look at what you have: the numbers everything hangs off

The model probes the source first and prints its duration, resolution, frame rate and audio sample rate. Every later check is measured against those numbers, so a wrong guess here breaks everything after it.

Transcribe it locally: word timings you can trust

Whisper runs on your own machine, never a hosted API, and it has to give a time for every word. The prompt makes the model prove the transcript reaches the end of the file and that the timings line up with the real audio. A transcript that stops early looks perfectly valid and quietly ruins every graphic after it.

Don't touch the timing: edit the picture, not the cut

The recording is already the final spoken cut. The model can't trim a pause, reorder a line or change the speed, and the output has to match the source's duration to the frame. That keeps the test fair: both models edit the exact same timeline.

Decide what the viewer sees: the real test

This is the step the prompt deliberately leaves open. There's no list of graphics to pick from. The model gets four rules of craft instead: land every graphic on its word, leave most lines bare, show the thing rather than label it, and make it work with the sound off. It then has to defend every choice in a table.

Build it in Remotion: one real component per idea

Every graphic is React code. The speaker stays on screen the whole time, captions are word by word and in sync, and everything keeps clear of the edges where Instagram puts its buttons. One generic card fed 12 different sentences is exactly what the prompt rules out.

Leave the audio alone: no music, no effects

The original audio goes straight through to the finished file. No music, no sound effects, no loudness processing. What you're judging is the picture.

Prove it before calling it done: stills, overlaps and duration

A passing build doesn't count. The model renders stills across the whole video and looks at them. It checks that no two graphics overlap and that the captions run to the last word. It also confirms the duration and audio match the source, then says plainly which checks it actually ran.

Where do I get the full prompt?

Download the prompt as a PDF. It's the full text, ready to copy. Paste the absolute path to your recording at the end. Then swap the brand block for your own colours and typeface, or delete it and let the model choose a look.

So which one edited it better?

There's no clear winner, and that's the honest answer. Each model made its own calls on the same lines, and which set you prefer comes down to taste. Watch both edits in the reel and tell me in the comments which one you'd post.

Video editing is one job AI can take off your plate. To find the ones costing your team the most time, take the free AI bottleneck test, or read our insights on AI implementation.

You made it to the end

Want to use AI in your business but don’t know where to start?

FreshStack builds AI systems that drive more revenue and cut the manual work slowing your team down. Book a 30-minute call and we’ll show you where AI fits first.