Content
A Dolby Atmos mix is not a wider stereo mix. It is a scene: a bed of channels plus objects with positions, rendered by the listener’s device into whatever layout it happens to have. This tutorial walks through the tool chain, the deliverable, and the mistakes that make an Atmos master arrive technically correct and still sound worse than the stereo it came from.
Full disclosure before we start: I do not produce Dolby Atmos music myself. I work with music in immersive contexts, but music production in Atmos is not my field, and for that part of my course I bring in a mixing engineer from Berlin who mixes chart records. What follows is the technical chain and the design questions, not a claim to be your mix engineer.
Two pieces of software, at minimum.
The Dolby Atmos Renderer, included in the Dolby Atmos Production Suite and the Dolby Atmos Mastering Suite. It renders audio and Atmos metadata coming out of your DAW and lets you monitor the Atmos mix. It also does the binaural rendering, which is intended for encoding content as Dolby AC-4 Immersive Stereo.
A panner. Either the Dolby Atmos Music Panner plug-in, or the Atmos panners built natively into Pro Tools and Nuendo. Instantiated on an object track, the panner feeds position and other metadata to the renderer.
That is the minimum. Everything else is monitoring.
Fewer people need a full room than the marketing implies, because the renderer’s binaural mode exists precisely so you can work on headphones. But if you do build a room, the component that matters most is not the speakers.
You need a device that records the master, and you need a monitor engine, because you no longer have just stereo but many speakers. You have to be able to set downmixes, hear how it sounds binaurally, in 5.1 and in 7.1, and solo individual paths, for example only the front height speakers, so you know what is actually running up there. The monitor engine plays an even more decisive role in an immersive studio than in a stereo one.
The signal path: from the DAW into the production tools, then into the immersive engine or renderer, which is format agnostic, then into a multichannel interface that goes either to the binaural output or to the speakers. On larger setups the rendering runs on a dedicated machine, which is what Dolby’s Mastering Suite approach is for. That suite has to receive 128 channels, because that is the maximum size of Dolby Atmos: a bed of ten channels plus up to 118 objects.
The bed is a fixed channel layout. Those are the ten channels of the 128 the renderer handles, which works out as 7.1.2. Objects are sources with coordinates that the renderer places at playback.
The working rule I would give anyone starting out is the same one that governs stereo. Bass is hard to localise, so you want it mono and centred, not pulled wide, and that law does not change in three dimensions. Add to that a fact people discover too late: the moment you spatialize a source, its timbre changes, because you have given it room information and separate information for the left and right ear. So keep out of the object domain anything you do not want recoloured. Kick drum, a normal bass, anything with few frequency components.
Dolby’s approach is to spatialize everything. Mine is that you need a solid foundation first, and from stereo we already know some sounds should be mono and are completely fine that way.
Mix big on speakers, then press the magic button labelled “binauralize”, and suddenly it is stereo. That is not how it works.
The binaural render is a separate design decision, not an export setting. In the renderer the binaural modes are Off, Near, Mid and Far. They describe the distance between the sound source and the listener’s head position, they are stored per object as metadata, and they only affect the headphone signal. On speaker playback they do nothing. If you never set them, every object is at the renderer’s default distance and the mix collapses towards your head.
Given that the overwhelming majority of listeners will hear your Atmos release on headphones, this is not a finishing touch. It is where most of the perceived quality lives.
Dolby comes from film. Many films are mixed in Atmos, and out of that came the idea to use the format for music too. What nobody asked first was what music would actually have needed in order to be immersive.
What Atmos Music does instead is recreate the feeling of sitting in a cinema. Listen to the releases and you notice the reverb and the room, but the punch is gone. You no longer have the pressure you know from a stereo production. I understand the intention. I just keep asking why you would want that for music.
There is a second structural problem, and it is worth knowing before you commit. In the bed, dialogue or a lead vocal can sit in the centre. Move your head and you hear the centre move with you. The bed was designed for a forward-facing listener, but with head tracking people do not just look forward, and then all objects and the whole bed start rotating, which was never the intent. Apple now enables head tracking for everything, so all Dolby mixes behave this way, and Dolby’s answer is that users should switch head tracking off.
According to Dolby, the delivery to streaming services has to be a BWF/ADM file exported from the renderer.
The sequence: after the production, record an Atmos master in the renderer. The metadata lands in the Dolby Atmos Master File Set, the DAMF set. Then export the .atmos master as a BWF/ADM .wav master. That single file carries the beds, the objects and all the automation, which is why the recipient needs nothing from your session. More on the underlying format in the Audio Definition Model.
Verify before you deliver, not after rejection. A file that opens is not the same as a file that is correct, and the failure mode is silent.
You can enable head tracking with a 360° video on YouTube, and on Facebook too. So you can upload your Dolby Atmos mix inside a 360° video and show people how the mix sounds with head tracking. The caveat: for music people this is not especially exciting, and you need some picture material for head tracking to engage at all. These days it also works without a headset and without video, using a head tracker or headphones with tracking built in.
Route object tracks through an Atmos panner to the Dolby Atmos Renderer, place the objects, keep the low end in the bed, set the binaural distance per object, then record an Atmos master in the renderer and export it as BWF/ADM.
That you are mixing for a renderer you do not control. The same master produces a different result on a soundbar, a 7.1.4 room and a pair of AirPods, so you have to check all three rather than perfecting one.
There is no single two-channel target to master to. What you are actually doing is checking that the render holds up across the downmixes, and that loudness is consistent between the Atmos version and the stereo version of the same track.
No. The Production Suite with binaural monitoring is enough to produce a valid master. A room gets you certainty about the speaker render, which is a different question from whether the mix is good.
You can upmix, and the algorithms are improving quickly. It is not the same as a mix built spatially from the stems, because the upmixer has to guess what belonged where.
For most Atmos music, no. The mixes were made for a forward-facing listener, and the bed rotating with your head is not something anyone designed for.
produce your Dolby Atmos Mix with this courseWorking on something immersive where music is one layer among several?
Get in touch →Related Articles
3D Audio - the immersive spatial soundtrack from all directions