← Back to Blog
AI & Filmmaking · September 17, 2025

Stop Prompting AI Video Like It's Midjourney

By Pranav Arya · PAFP · #ai-video · #prompts · #tutorial · #workflow

Almost every AI video prompt I get asked to critique makes the same mistake. It describes a beautiful frozen picture, lighting, mood, styling, composition, and then just tacks a verb onto the end and hopes for the best. That's a Midjourney prompt wearing a video model's clothes, and it produces exactly what you'd expect: a gorgeous first frame that barely moves for four seconds.

Motion is the subject, not an afterthought

The models I use day to day respond dramatically better when camera movement and subject motion are described with the same specificity as the visual style. Instead of "a woman walking through a market, cinematic," I'll write something closer to "handheld camera tracks alongside a woman walking left to right at a steady pace, slight vertical bounce, market stalls blur past in the foreground at the edges of frame." That's a longer prompt, and it produces footage that actually looks shot rather than generated.

Three things worth naming every time

Camera behavior first: is it locked off, handheld, a slow push, a whip pan? Say it plainly. Subject motion second: walking speed, head turn, hand gesture, and roughly when in the clip it happens. Depth of field third, because static-image prompting habits tend to describe an entire scene in uniform focus, and real footage almost never looks like that.

Where this breaks down

Even with a well-structured prompt, most current models still struggle with complex multi-subject motion where two people need to move independently and believably relative to each other. I've mostly stopped fighting this and instead generate single-subject clips and composite multiple layers together in post when a shot needs more than one moving element. It's more work, but it's reliable work, and reliable beats clever on a deadline.

The underlying shift in mindset matters more than any single prompt template. You're not describing a picture anymore. You're directing four seconds of a scene, and the model needs blocking notes, not just a mood board.

Pranav Arya is a Berlin-based filmmaker producing AI video content for brands and social media, alongside real-world event, brand, and fashion shoots worldwide. He also teaches photography and videography to aspiring creators. Get in touch to work together.