← Back to Blog
AI · Entertainment · October 04, 2026

Kling 4.0: 30-Second Clips, 4K HDR, and Real Competition

By Pranav Arya · PAFP · #kling · #ai video · #news · #seedance
Kling 4.0: 30-Second Clips, 4K HDR, and Real Competition

Thirty seconds. That's the headline number out of Kuaishou this week, and if you've been cutting AI video into 5-second fragments and stitching them in Premiere like the rest of us, you'll understand why that number is a big deal.

Kling 4.0 was announced on September 28. It roughly doubles what Kling 3.0 could do in a single pass: 30-second native clips, up to 4K output in 10-bit HDR, 10 keyframes for story control, and an Omni Reference mode that takes 10 images, 5 videos, and 7 saved characters or objects in one task. For anyone building ad campaigns or brand films where consistency across shots has been the entire bottleneck, that reference ceiling matters more than the resolution bump. I've spent more hours than I'd like to admit re-uploading the same character reference into Kling 3.0, hoping it holds the jawline. If 4.0 actually nails 7 saved characters across a sequence, production teams get a real shift in how they work, not just another spec to brag about.

There's a lighter Kling 4.0 Flash already in limited early access, generating up to 20 seconds per clip, but it's locked to Ultra Yearly subscribers for now. The full 4.0 release doesn't have a confirmed date beyond "October 2026," which, given we're already four days into October, means this could land any week now. I'm watching this one closely. If you run a production pipeline that depends on Kling, you should be too.

Worth remembering: Kling isn't actually winning right now

Independent benchmarks on Artificial Analysis currently rank Kling 3.0 in 12th place, which is worth saying out loud since the narrative hasn't caught up to the leaderboard yet. If you've seen a creator confidently declare "Kling is the best model out there" recently, that claim is stale. The real competitive pressure right now comes from ByteDance's Seedance 2.5, which also does 30-second passes but pushes reference handling even further, up to 50 references in one task. Kling 4.0 is a big step forward for Kuaishou, but it's as much about catching up to Seedance's ceiling as setting a new one of its own. Benchmark it yourself before taking the press release at its word.

Sora is dead, and nobody built a proper replacement

The bigger story of the season isn't a new model. It's the absence of one. OpenAI discontinued the standalone Sora web and app back in April, then pulled the Sora 2 API entirely on September 24. No migration path, no "here's Sora 3." Just a shutdown notice. For a tool that a huge chunk of the creator economy had built workflows around less than two years ago, that's a significant exit, even if it's a quiet one.

Runway has absorbed most of that displaced traffic, and honestly, it's the smartest move they've made. Gen-4.5 is now the flagship, but the real pitch is the workspace itself. Runway now hosts 20+ third-party models in one place, including Kling 3.0, Seedance 2.0, Veo 3.1, and WAN, plus an agentic "Runway Agent" that handles prompt-to-finished-video without you babysitting every intermediate step. If you're still running Sora-shaped pipelines in your head, it's worth rebuilding them around Runway as the hub rather than any single model. Betting your whole pipeline on one vendor's video model is starting to look like a mistake, and infrastructure that doesn't care which model you plug in is clearly where this is heading.

What's actually usable on set this week

Setting the roadmap talk aside, here's what's actually worth putting in front of clients right now.

ElevenLabs v4 launched days ago and immediately took the #1 spot on Artificial Analysis' speech arena. The new architecture reads tone, pacing, and character straight from the text rather than needing heavy tagging, and it now covers 90+ languages, up from 70 in v3. It's already wired into Higgsfield and HeyGen pipelines, which means whispers, laughter, and sound effects can live directly in your script instead of being bolted on in post. Getting the emotional beats right in voice work used to take two extra passes. Now it's closer to one.

And here's the actual workflow trick worth stealing. In Kling 3.0, upload your starting image as the first frame, flip on multi-shot in the settings panel, and stack three shot prompts, each one just a camera direction plus a line of dialogue. The lip sync in Kling 3.0 is genuinely the best I've seen in any video model right now, but there's a real limitation: sentences over 10 words start drifting slightly at the end of the phrase. Keep your dialogue short and punchy. "Pull back, wide shot. 'We're not done yet.'" works better than a monologue. This single tip will save you more re-rolls than any prompt engineering trick I've seen this year.

One more data point worth noting: the AI-generated micro-drama "Feng Shui Tian Shi" hit over 100 million views on Douyin in 12 hours, and viewers genuinely didn't clock it as AI on first watch. Meanwhile HeyGen open-sourced a real-time avatar stack built on GPT-Live-1, running full-duplex speech-to-speech with tool calls appearing as animated overlays on screen. This isn't a demo reel anymore, it's a product running at distribution scale, and it changes what clients will expect from a "quick turnaround" by next quarter.

The gap between "AI video tool" and "production tool" closed faster this year than anyone in this industry predicted, and it closed from the China side first. If you're still benchmarking your workflow against what Sora could do a year ago, the issue isn't that you've fallen behind. It's that you're measuring against the wrong thing entirely.

Pranav Arya is a Berlin-based filmmaker producing AI video content for brands and social media, alongside real-world event, brand, and fashion shoots worldwide. He also teaches photography and videography to aspiring creators. Get in touch to work together.