I switched my Kling tab and my Veo tab side by side this week and ran the same prompt through both. Kling nailed the camera move on the first try. Veo needed a reference image and a rewrite before it stopped ignoring my lighting instructions. Neither of those sentences would have been true four months ago. The pecking order in AI video just got scrambled, and anyone planning a client pipeline off last quarter's assumptions is already behind.
Kling's quiet takeover
Kling v3 is sitting at the top of the live text-to-video arena leaderboard with a score of 1991, well clear of Happy Horse 1.0 at 1783 and Seedance 2.0 Fast at 1703. That ranking comes from blind votes, over a thousand of them, where people compare raw outputs with no idea which model made what. Nobody's picking Kling because it's trendy or because a big name is attached to it. People are just picking the clip that looks right. Kling's motion coherence has felt underrated for a while now, especially among Western creators who default to whatever OpenAI or Google ships without checking. This leaderboard is the data catching up to what a lot of people on set already suspected. If you're still defaulting to Sora or first-gen Veo for anything involving complex camera movement, it might be worth running a side-by-side before your next job.
Sora's app is gone, and the previz gap is where things get interesting
OpenAI killed the standalone Sora web and app back in April. It only survives inside ChatGPT now, and the API sunsets in September. Google and ByteDance are both moving into the space that leaves behind, from different angles. Veo 3.1 is still the only widely available model doing proper 48kHz synced dialogue instead of vague background noise, and its "Ingredients to Video" feature lets you feed in up to three reference images to hold a character or product steady across a whole sequence. It's the feature I reach for most on brand work: a client sends product shots, I lock the appearance, and I stop fighting model drift scene to scene. The more interesting launch this month is Seedance 2.5, out of ByteDance's Volcengine unit. Native 4K, 30-second outputs, local editing, all fairly incremental on their own. What made me sit up was the 3D pre-visualization tool. You block out camera moves and lighting on a virtual set before you render final textures, which isn't really a video generator feature so much as a director's tool. ByteDance's CEO has cited how film directors previz shoots as the direct inspiration. If it works as advertised, this is one of the first genAI tools built around how a working DP or director actually thinks rather than what looks good in a demo reel. I'd like to get my hands on it for an upcoming F1 activation shoot, where lighting continuity across a dozen setups is the whole headache.
Meta arrives late and brings a consent problem
Meta dropped Muse Image on July 7, built by Meta Superintelligence Labs and apparently code-named Mango internally. It's free in the Meta AI app, Instagram Stories, and WhatsApp, and it's landed well: number two on the Arena Elo human-preference rankings for image gen and editing as of July 5. Meta says Muse Video is "already in development," which reads less like news and more like a formality. The part that should bother more people is quieter than the launch itself. Muse lets you manipulate another public Instagram user's images with AI. Their profile just has to be public. I run a public account for my studio. So does most of the industry. Meta has effectively built a feature that turns "public" into "fair game for anyone's AI editing session," and it's being framed as a fun creative tool rather than what it actually is, which is a consent problem with a nice UI.
The culture war catches up with the tech
Three things this month made clear the fights are no longer theoretical. Tilly Norwood's studio announced its first feature, "Misaligned," a comedy-drama built as a hybrid production: real directors, writers, and editors working alongside AI specialists, with AI training baked into the workflow itself. SAG-AFTRA's response was blunt: not an actor, "a character generated by a computer program" with no life experience or emotion to draw from. It's a hard position to argue with, and it's also not going to slow any of this down. Then there's Charlie Curran, whose spoof videos have racked up tens of millions of views since January, using tools including Seedance 2.0 to make a Batman-riffing attack ad for an LA mayoral candidate. Political ads made with Chinese-developed video models, running in American local elections, is a sentence that deserves more scrutiny than it's currently getting. And Tribeca screened "Dreams of Violets," a feature-length "live action" film entirely AI-generated. No cameras, no actors, no sets, made for around $2,000, and a first for the festival. Whether it's actually good is up for debate, reports are mixed, but the more telling fact is that a festival built around traditional filmmaking just legitimized a $2,000 no-crew feature as something worth screening at all.
The workflow tip that's actually useful
Stop treating every model like a slightly different flavor of the same tool. Creators who are getting consistently better output right now are prompting each model according to what it's actually built to optimize for:
- Runway Gen 4.5 wants physics and camera language. Think in motion vectors and forces, not adjectives.
- Kling generates audio and video together, so write it like a timeline script with beat markers, not a text description.
- Veo 3.1 performs best with structured input. JSON-style schemas and reference "ingredient" lists beat prose every time.
- Sora 2 is closer to a physics simulator. Describe cause and effect, not just the end frame.
I've started keeping four separate prompt templates for the same shot concept instead of one prompt pasted everywhere. It's more setup work upfront, but it's cut my regeneration count on client jobs by roughly half.
None of this settles the bigger argument about what AI video is for, or who it's for. But it does raise a narrower, more immediate one: who gets to decide what "public" means on your own social profile, and whether that gets answered before it matters in an actual election.