Content
Artificial intelligence has arrived in audio, and not as a future promise. What it can already do surprises me as a sound engineer. What it supposedly can do is often overrated. Here is my assessment, without the marketing fog.
Once a mix is finished with all the instruments in it, pulling the vocal back out used to be practically impossible. Now AI comes along and makes it possible.
It is like buying a finished cake at the supermarket, putting it in the microwave at home, and suddenly the individual ingredients come back out. Impossible really, but it works.
The joke is that this happens entirely without object-based audio. In an object-based mix all the individual objects would be there anyway, we are just denied access to them. The AI retrieves them from a finished stereo mix.
And it is already usable: Apple Music has a karaoke function that turns the voice down when you want to sing along yourself.
This is my boldest claim on the subject. Upmixers are getting better fast, and in future you may not need Dolby Atmos content at all.
The evidence sits in the car. Upmixing algorithms there already create impressive spatial experiences out of a plain stereo file. AI can do that better than sound people who are limited by time and budget.
What appeals to me about it: a version of a song that gets better over time is more interesting to me than being locked to the best mix that was possible on release day.
I have followed the progress of object-based audio for years and I am convinced its potential only unfolds through artificial intelligence. Despite the immense skill of countless sound engineers worldwide, the complexity of fully optimising object-based audio exceeds what humans can do alone.
That is why I consider the AI-based scene analysis in IAMF a remarkable step. It promises to optimise audio tracks automatically for different content.
The price is real. There is a concern in the audio community that the original artistic intent for certain scenes gets overridden, and the listening experience turns out differently than the makers intended. That balance between AI optimisation and individual preference is what will matter.
AI has changed little about my own workflow so far. There may even be more films, because high quality becomes more accessible. But immersive audio is so specific that you cannot really train an AI on it at the moment. Where that goes remains to be seen.
A second example of unfinished technology is personalisation. There is no universal HRTF that works equally well for everyone, which is why companies are researching dynamic or AI-based selection. That is still in development.
In ordinary listening situations many users cannot readily tell a generic HRTF from a personalised one. On top of that, Apple, Samsung and Ceva use different HRTF databases, which leads to inconsistent experiences across devices and content sources.
Not everything sold as AI is AI. The Klipsch T5 II True Wireless ANC offer gesture control: answer calls by nodding, skip songs with movement patterns.
Forbes magazine took that for artificial intelligence. It is head tracking.
Here I disagree with most of my own trade. Most sound people say AI cannot replace humans. I say of course it can.
In immersive audio I still have good hope for very specific use cases, namely my own, simply because there is not enough training data there to replace me. That is good for me right now.
For podcasts and even music production there will be an AI that does the whole job. What it cannot do is set up microphones and other physical things. And even where it fails today: imagine what it looks like in a year.
From that follows my test question for anyone in this profession. Focus on very specific use cases and ask yourself seriously whether an AI could replace that within the next twelve months. If the answer is yes, you should not be doing it. That is an unpopular opinion, I know.
Then there is the market. Film production is already struggling, and AI will increase the pace at which classic film productions decline. That hits sound people directly.
My advice: take AI seriously and use it to your advantage. Do not be afraid of AI, be afraid of a sound engineer who uses AI.
My forecast for the next ten years is not about formats, it is about control. Until now another person or a system decides what you get to hear and how.
I see a strong trend towards users getting more control over how they listen: the guitar louder because you like guitar, the film score quieter because it is sometimes too loud. Object-based audio is the technical basis for that, and AI plays the major role.
This goes beyond entertainment. Even phone calls and navigation become more personal, and you get addressed differently from every other user. Technically it is still very hard to give everyone different audio content. But synthetic audio created on demand is something I expect to see more and more of in the coming years.
Want to know what can sensibly be automated in your project and what cannot?
Get in touch →
back to blog