A huge share of the videos I deliver get watched on a phone speaker in a noisy environment, or with the sound completely off and captions carrying the meaning, and mixing audio for that reality is genuinely different from mixing for a controlled viewing environment with good headphones. It took me a while to fully adjust my mixing habits to match how content actually gets consumed rather than how I'd prefer it to be consumed.
Compress harder than feels natural
Dynamic range that sounds rich and natural on studio monitors collapses into inaudible quiet passages on a small phone speaker in a loud environment, so I compress audio noticeably harder for social deliverables than I would for anything meant to be watched with real headphones or speakers. It sounds slightly less nuanced in a controlled listening environment and noticeably more usable in the actual real-world conditions most viewers are in.
Captions aren't a backup, they're a primary channel
I now build captions as a first-class part of the edit rather than an afterthought added at the end, timing them to actually support comprehension for someone watching with sound off entirely, which describes a meaningful share of viewers on some platforms. That's changed how I pace dialogue and even how I write on-camera talking points, keeping sentences shorter and more caption-friendly rather than optimizing purely for how something sounds.
- Compress harder than feels natural for controlled listening environments
- Build captions as a primary channel, not a backup for accessibility alone
- Test a mix on an actual phone speaker before calling it finished, every time
The test that actually matters
I now do a final listen on a phone's built-in speaker, no headphones, at a moderate volume, as a required step before calling any social deliverable finished. It's caught more real problems, dialogue getting lost under music, an effect that sounded fine on monitors disappearing entirely, than any amount of critical listening on good studio speakers ever did, purely because it matches how the thing is actually going to be experienced.