I spent a chunk of last week on a client job rebuilding a character from scratch because I got lazy with reference images partway through a shot list. The face held. The jacket didn't. Nobody in the client review called it out specifically, but everyone felt something was off, and "something feels off" is the most expensive note you can get on a deadline. Character drift is still the thing that separates AI video that looks like a toy from AI video you can actually sell to a brand. There's finally a workflow that handles it properly.
Kling 3.0 Omni's Elements system is the first tool I've used where I didn't have to fight the model to keep a face, a wardrobe, and a voice locked across multiple generations. Kling 4.0 Flash has been in early access for a few weeks now, with full release landing soon. Now's the time to actually learn this workflow instead of limping along with your old 15-second habits.
Building the Element: your reusable character asset
The old way of doing character consistency in AI video involved a lot of guesswork. You'd write increasingly specific prompts, cross your fingers, generate ten takes, and cherry-pick the one where the character didn't grow a third eyebrow. Elements replace that guesswork with an actual asset.
An Element is a reusable reference you build once from source media, either images or a short clean video clip, and then call into every subsequent generation. For a one-off background character, two or three solid reference images will do. For anyone recurring in a campaign. Your spokesperson, your mascot, your client's "brand ambassador" avatar. I build the Element from a short clean video clip instead. It gives the model more to lock onto: how the face moves, not just how it sits still.
Stick to two to four reference images per Element. I tested six once, thinking more data equals more consistency. Wrong. The model started averaging features instead of locking them, and I got a character with slightly melted bone structure. I don't go past four anymore.
The prompt template that actually works
Once your Element exists, binding it inside a prompt is where the real control happens. Kling 3.0 Omni treats everything you feed it as one combined prompt. Images, clips, elements, text, audio. And it can lock each character or object independently within a single scene. That matters if you're running two named characters through a dialogue scene and don't want one of them to inherit the other's jacket by shot six.
Here's the template I've been running on client work, lightly adapted:
- Reference @Element1 as [character name]'s identity, face, wardrobe, and voice.
- Use [start frame / reference image] for the scene composition.
- [Shot type / camera movement] of [character] in [setting].
- [Character action over time].
- [Environmental motion].
- Audio: [dialogue, ambience, sound effects, music direction].
The part people skip is the negative prompt discipline. Lock your fixed descriptors (exact jacket color, exact haircut, exact accent or no-accent) and explicitly negative-prompt the things the model likes to improvise, like changing eye color under certain lighting or adding jewelry nobody asked for. Then use AI Multi-Shot to generate your angles off the same locked reference, and extend your scenes by carrying that same reference forward rather than starting a new generation cold. Build the master character first, lock the prompt, generate your angles, then extend forward from there. In roughly that order, because skipping the first part undoes the rest.
Why the 30-second jump changes your shot list
Kling 4.0 expanding clip length to 3-30 seconds, up from the old 15-second ceiling, sounds like a spec bump. It actually changes how you should be planning coverage. At 15 seconds you were basically forced into a cut-heavy edit style, lots of short punchy clips stitched together, because nothing could breathe. At 30 seconds you can hold a shot long enough for a character to deliver an actual line of dialogue with a reaction beat after it. That means fewer generations, fewer consistency failure points, and an edit that doesn't feel like a TikTok supercut by default. I'm already restructuring shot lists for an upcoming Web3 summit recap around this, planning for 20-25 second masters instead of chopping everything into 8-second fragments and praying the cuts hide the seams.
A quick note on what died and what's worth your budget
If you're still finding Sora tutorials online, skip them. OpenAI shut the Sora web and app experiences down, then pulled the Sora 2 API entirely not long after. That ecosystem is closed. Google, meanwhile, quietly stopped leading with Veo and now points to Gemini Omni Flash, which is sitting at number one on both Artificial Analysis and Arena blind-vote leaderboards. It's less urgent to learn today than Kling's Element workflow, but it's worth knowing the ground shifted.
On pricing: for a 30-second spot, Veo 3.1 Lite with audio comes in cheapest, Runway Gen-4.5 silent costs a bit more, Kling 3.0 with audio runs higher still, and Veo 3.1 Standard with audio is the most expensive of the bunch by a wide margin. Kling costs more per second than the Lite tier. But if character consistency is your actual bottleneck, that's not where you cut corners. I'd rather eat the extra cost per take than spend another afternoon rebuilding a jacket.
Higgsfield's Hell Grind premiered at Cannes this spring with zero A-list casting and zero Oscar pedigree, and it got written up because it held together as a watchable film made with these same kinds of tools. Not impressive for AI. Just a film people wanted to finish watching. Elements and the longer clip window are what's closing that gap, and the people who learn this workflow now will have a finished reel while everyone else is still rebuilding jackets on shot four.