Sora's API dies in September. I still have three client contracts referencing a Sora-based pipeline that was written in January like it was permanent infrastructure. It wasn't going to be, and nothing in this industry is permanent infrastructure right now. If you've been building your workflow around one model instead of one process, this is the month that bites you.
I've spent the last two weeks rebuilding a chunk of my own pipeline around Seedance 2.0. The model earned the switch, though I'll admit the Sora deadline forced my hand too. Here's what's actually working on real jobs right now.
The depth map trick is real, and it's dumber than you'd expect
Everyone's talking about this right now: before you feed reference footage into Seedance 2.0 for motion transfer, convert it to a grayscale depth map first. Don't send the raw clip. Strip it down to structure, then let Seedance rebuild appearance from your style reference separately.
I tried this on a Web3 summit recap where the client wanted a specific walk-and-turn camera move recreated on a CG-ish brand mascot. Feeding Seedance the raw reference clip gave me motion that was close but kept bleeding the original actor's proportions into the output. Swap in a depth map and the model stops fighting itself. It has nothing to reference but motion and geometry, so that's all it copies, while appearance comes from your first-frame style image. The separation gets a lot cleaner, and so does the result.
If you're in ComfyUI, there's now a node for this that fuses the depth map (structure), your first-frame reference (appearance), and your prompt (semantic intent) into one pass. The batch fan-out and collector nodes let you queue four style variations off a single depth map in one run, and that's the actual time-saver here. You're not re-running motion capture four times. You're re-running appearance four times off a motion skeleton you already solved.
Actionable version if you're not touching ComfyUI: run your reference clip through a depth estimation pass, export it as grayscale video, and use that as your Seedance motion input instead of the source footage. One extra step, noticeably fewer weird identity-bleed artifacts.
Omni Mode doesn't feel like a gimmick, which surprised me
Seedance 2.0's Omni Mode takes 9 images, 3 fifteen-second video clips, and 3 audio files simultaneously, alongside your prompt. A character photo anchors appearance. A film clip anchors camera movement. A track sets pacing. The model reads all of it at once instead of you stitching separate generations together in post.
I was skeptical until I used it on an F1-adjacent brand piece where the client wanted a very specific slow dolly-in matched to an old reference spot, set to a temp track that dictated the emotional beat of the cut. Normally that's three separate tools and a lot of manual timing in Resolve afterward. Omni Mode got me 80% of the way there in one generation. I still had to stabilize and de-flicker in DaVinci after, since that step hasn't disappeared, but the base plate was usable instead of a rough draft. It justified the cost on its own.
Start/end-frame interpolation is the other half worth mentioning. You define your opening and closing composition and Seedance fills the physics between them. This is genuinely good for match cuts and reveal shots where you know exactly where you're starting and landing but don't want to hand-animate the transition.
Know when to leave Sora behind, and where everything else sits
The sunset timeline is blunt: web and app access ended in April, API access dies September 24th. If you're a ChatGPT Plus or Pro subscriber, Sora 2 still lives inside ChatGPT itself. But any production pipeline built around the standalone app or API needs a plan this month, not in August when you're panicking before a deadline.
My current cheat sheet, drawn from jobs I've actually billed:
- Runway when the client needs granular control over a specific shot
- Kling when budget and iteration speed matter more than polish
- Luma for moody image-to-video work, especially anything atmospheric
- Veo when you need the best all-around output including audio
- Sora only if you're maintaining something already built on it
Betting on a single model is a losing strategy this year. Betting on a repeatable process has held up a lot better.
Speed as a technique, and why boards still save your budget
The Popeyes "Wrap Battle" case is nearly a year old now but it still holds up as the reference point for reactive production. Full ad, music included, done in under 3 days using Veo 3 and Suno, after the team abandoned image-to-video because it was too slow for the news cycle they were chasing. It's a technique you build a retainer around if you're working with any brand that wants cultural-moment reactivity.
But speed without structure is how budgets evaporate. The discipline I keep coming back to, and the one I'd push hardest on anyone starting out right now, is boarding before you generate motion. Boards cost cents. They decide composition, continuity, and shot intent before you've spent a single credit on video generation. Skip straight to motion on an un-boarded scene and you end up re-rolling the same shot fifteen times at $2 a pop, chasing a camera move you never actually defined.
Write the script. Board it scene by scene. Generate stills from beats the story has actually earned. Then generate motion. The models will keep trading the lead every quarter. It's Seedance now, it'll be something else by winter. Whatever pipeline you've built around evaluating and combining them is what actually stays yours.