Content
Spatialization is the craft of giving a sound a position and a room. Panning puts it somewhere. Reverb makes the brain believe the somewhere is real. Most 3D productions get the first part right and the second part wrong, which is why so much spatial audio still sounds like it is stuck to the listener’s head.
New to the topic? How spatial audio works covers the basics. This guide is about the practice.
Disclaimer: the first draft of this article came out of an AI. I rewrote it against how I actually work.
For me, spatialization is two things at once: 3D panning plus reverb. Only both together produce a place. dearVR Pro understood that and gives you both in one tool. Dolby Atmos, on the other hand, is panning first of all — the object carries a position, but no information about the room it is in. That is one of the reasons Atmos does not convince over headphones: position alone is not enough for the brain, it needs the room information with it, or the sound stays glued to your head.
Spatialization means placing a sound source at a defined position in three dimensions and giving it the acoustic properties of the space it is supposed to be in. In a headphone context that is done with HRTF filtering plus room simulation. Over speakers it is done by distributing the signal across a loudspeaker array.
Start from the other end and it is easier to see. Mono is a single audio channel with no spatial information at all. You know neither where the sound is nor what kind of room it is in. It stays an abstract source in a dead room. Spatialization is everything you add to that.
This is the mistake I see in most 3D productions, and it is always the same one. People pan objects around the 3D space without giving them any room information.
Put a sound up and to the left and you will hear it up and to the left. But in reality you would simultaneously hear room reflections from behind you and a little from the right, because sound moves through the room. Without that information the brain concludes the source must be very close to your head.
That is the secret of externalisation. The brain needs the reverb information before it will accept audio as three-dimensional. My own proof is the dummy head demo: when I speak into the left ear, something happens on the right too, even if only a little. The brain needs that, because it has spent a lifetime learning it. Even when a car is on your right, you still hear something on the left.
Run it as a classic send. Turn the send all the way down and the voice sticks to your face again. Send too much and the speaker drifts further and further away and intelligibility goes with them. The reverb send is your distance control, and finding the right amount is a listening decision, not a preset.
Binaural sound helps us to localize sounds the same way in the real world, which is why the technique works at all.
Not everything belongs in the room, and this is where I disagree with Dolby.
Bass is hard to localise. In stereo you already want it mono and centred, not pulled wide. The same law applies in 3D audio, so bass does not belong externalised there either. Dolby tries to spatialize every sound, and my position is that certain sounds should not be spatialized.
There is a second reason, and it gets mentioned even less often. The moment you add spatialization, the timbre changes, because you are giving the object room information and separate information for the left and right ear. So I deliberately keep sounds out of the process that I do not want altered: the kick drum, a normal bass, anything with few frequency components.
My working method is a balance of some in-head localisation and some externalisation. You need a solid foundation, and from stereo we already know that certain sounds should be mono and are completely fine that way.
| Family | How it stores space | Consequence for you |
|---|---|---|
| Channel-based | Fixed speaker assignment: mono, stereo, 5.1 | Tied to a speaker layout |
| Binaural | Exactly two channels, left and right | Headphones only, works everywhere |
| Ambisonics | A sound field in 4, 9 or 16 channels | Rotatable, but objects are no longer separable |
| Object-based | Every object moves freely in space | Renderer decides the output at playback |
Ambisonics sits between the two extremes and is called sound-field based. The individual objects are in there, inside the sound sphere, but you can no longer reach them afterwards. That single property decides whether a project should be recorded in Ambisonics with 360 microphones or built object by object.
More background in Dolby Atmos and in the three-dimensional soundscape glossary.
If you want a sound between two loudspeakers, you cannot place it there directly. You put a bit of signal on both. When the signals are identical, the brain cannot separate them and decides the sound is in front. That is why two speakers sometimes make it feel as though the sound comes out of the screen, even though the screen produces no sound at all.
Over headphones this does not happen. The same identical signal lands as in-head localisation instead. Everything in headphone spatialization exists to work around that.
Tools matter less than the order does, but you do need the right tools at hand. For playback, 3D headphones for Dolby Atmos and surround sound and soundbars and smart speakers behave very differently, and a mix that only holds up on one of them is not finished.
Dynamic head tracking is the single largest improvement available in headphone spatialization, larger than personalized spatial audio. Without it, sounds stay fixed to your head instead of to the room, and front-back confusions multiply because the brain never receives the motion cues it expects. The technical background is in HRTF.
For music this matters less, because the mix was made for one listening position. For anything anchored to a picture it matters a great deal, which is also why automotive audio and VR music experiences pull in opposite directions.
Placing a sound source at a defined position in space and giving it the acoustic character of a room, so a listener perceives direction and distance rather than just left and right.
Not automatically. It sounds more spacious, and spatial processing usually also sounds louder, which people mistake for better. Matched for level, a bad spatial mix is worse than a good stereo mix.
Almost always missing reverb. You panned without giving the objects room information, so the brain places them at the closest possible distance. Add a shared room and set distance with the send.
The deliberate placement and movement of sounds within a piece. It has been a compositional device in electroacoustic music for decades, long before the streaming formats arrived. On why the streaming version often disappoints, see why Apple Music with Dolby Atmos sounds bad.
Yes, that is upmixing, and modern upmixers are getting good. It is not the same as a mix built spatially from the stems, because the upmixer has to guess what belonged where.
On the delivery format, at matched loudness, and on more than one pair of headphones. Then turn your head. If the image collapses when you move, the reverb is doing too little.
Back to the BlogGot stereo material that should sound spatial, and want it done properly?
Get in touch →