Blau beleuchtete Frau trägt transparente Brille
MusicTools

Upmixing: Turning Stereo Into Surround and Immersive Audio

Content

    Upmixing is the least respected craft in immersive audio, and it is about to become the most important one.

    Every discussion about Dolby Atmos assumes the same thing: that immersive audio has to be mixed by hand, track by track, in a dedicated session. Upmixing — deriving a surround mix from existing stereo content — is treated as the cheap shortcut you take when the budget runs out. My own position is less comfortable than that. Upmixers are improving fast, and there is a realistic future in which we no longer need dedicated Dolby Atmos content at all. An algorithm has no deadline and no budget ceiling. A human mixing engineer has both.

    This article is about what upmixing actually does, how the tools differ, where it fails audibly, and when a real immersive mix is still the only honest answer.

    What upmixing actually is

    The clearest way to think about it is as a ladder of dimensions. Mono has zero dimensions — a single audio channel that carries no spatial information at all and does not know where it is. Stereo is one-dimensional: left and right, nothing in front, nothing behind. Take the height layer away from 3D audio and you get a surface, which is surround — two-dimensional. Full 3D adds front, back, left, right, above and below.

    Upmixing is the attempt to move material up that ladder without going back to the multitrack. You feed the process a stereo source and ask it to produce a surround mix that contains information the original audio never explicitly carried.

    That framing matters, because it makes the central problem obvious: the spatial data is not hidden in the file waiting to be found. It has to be inferred. Everything that separates a good upmix from a bad one comes down to how intelligently that inference is made — and how gracefully it fails when the assumption is wrong.

    It is also worth putting the size of the step into perspective. The biggest leap in the history of recorded audio was not surround to 3D — it was mono to stereo. You only notice how large it is when you switch a stereo song to mono and hear how much collapses. Stereo to surround is a clear step as well, because suddenly something happens behind you. Comparing stereo directly against 3D is a limping comparison anyway: we are simply used to stereo, especially on headphones where the sound arrives directly at the ear, and it is rarely clear what “better” is even supposed to mean — more natural, more instruments audible, more involving?

    An upmixer is asked to make that comparison for you, automatically, on material it has never heard before.

    Abstract blue fluid painting

    Does an upmix count as real immersive audio?

    Yes, with one condition. An upmix is a genuine surround mix in every technical sense — discrete channels, real spatial distribution, valid deliverables. What it cannot contain is spatial intent that the stereo source never implied. If the creative concept depends on a specific element being placed behind or above the listener, that decision has to be made by a person in a session, not inferred by an algorithm.

    How an upmixer derives channels that were never recorded

    Every upmix algorithm works from the relationship between the two channels of the stereo content rather than from the channels themselves.

    Material that appears identically in left and right — correlated, centre-panned content such as lead vocals, kick, snare, bass, dialogue — can be identified and extracted into a discrete centre channel. Material that differs between the two sides — reverb tails, room signal, wide synths, audience noise — is treated as ambience and distributed to the surround channels. The better the algorithm separates these two populations, the more convincing the result.

    The height layer is a further extrapolation on top of that. There is no height information whatsoever in a stereo source, so an enhanced upmix algorithm has to decide which parts of the ambience read as “above” rather than “around”. This is where approaches diverge most, and where the differences between tools become audible fastest.

    Three failure modes are worth listening for specifically:

    Centre extraction that pulls too hard. When the algorithm is aggressive, the lead vocal becomes a hard mono point source in the centre speaker and the mix loses width where it used to feel solid. When it is too gentle, the vocal stays phantom-imaged and the centre channel does nothing.

    Ambience that turns into gated noise. Extracted reverb is never clean. Push the surround level and you hear the extraction breathing behind the music, especially on sparse material.

    Phase and downmix behaviour. This is the one people discover too late. An upmix that sounds impressive in the room can collapse when a playback device folds it back down. If correlated content ends up spread across channels with shifted phase, the downmixed versions lose exactly the elements that carried the track. A high quality downmix is not a formality you check at the end — it is a constraint you mix against from the start.

    The Dolby Surround Upmixer and the alternatives

    Four approaches dominate professional work, and they are genuinely different in character rather than merely in branding.

    Approach Character Typical use
    Dolby Surround Upmixer Conservative, predictable, tightly integrated Inside the Dolby Atmos Renderer, bed generation
    Nugen Audio Halo Upmix Detailed control over separation and width Post, broadcast, music
    DTS Neural Surround UpMix Broadcast heritage, robust on live material Live, sports, broadcast chains
    Auro-Matic Height-forward, strong vertical impression Auro-3D deliveries, cinema

    The Dolby Surround Upmixer sits inside the Dolby Atmos Renderer and is the one most engineers meet first, because it is simply there when you need to turn a stereo or 5.1 bed into something wider. It is deliberately conservative — it rarely produces a spectacular result and it rarely embarrasses you either.

    Halo Upmix from Nugen Audio gives the most direct access to the underlying upmix parameters: how hard the centre is extracted, how wide the front image stays, how much energy reaches the surrounds. If you want to shape rather than accept the result, this is where you get individual channel output control. DTS Neural comes from a broadcast lineage and behaves accordingly — it holds together on unpredictable live material where a more aggressive coherent spatial upmix would wander. Auro-Matic leans hardest into the height layer, which is exactly what you want for an Auro-3D delivery and exactly what will sound overdone elsewhere.

    For scene-based work the logic changes again: an ambisonic upmix in Ambix format is not channel extraction at all but a spatial re-encoding, and it behaves differently under rotation and head-tracking.

    Two practical notes. First, none of these tools has a correct default — the right upmix parameters depend on genre and on material, not on the plugin. Second, the differences between them shrink dramatically once you match levels properly. Which brings me to the part most comparisons get wrong.

    Abstract blue wave forms glowing on a dark background

    What a blind listening test actually revealed

    Rather than argue about this from theory, I tested it. Together with sound colleagues I ran an extensive listening session over a 5.1.4 system using Tidal, comparing identical mixes for about an hour.

    The conclusion was sobering. Most of the tracks did sound more spacious and broader — and that did not do all of them any good. Above all, they lost the punch that the stereo version had. Some titles genuinely worked better in the immersive version. It turned out to be strongly dependent on the individual track and on genre.

    That result is worth sitting with, because those were not upmixes. Those were hand-made immersive mixes, produced deliberately by professionals. If human immersive mixing already trades punch for width on a significant share of material, then the honest comparison is not “upmix versus real mix”. It is “which of the two serves this particular track”.

    One methodological warning from that session: any comparison run without level matching is worthless. Louder reliably reads as better. If you are evaluating an upmix against the original audio, match levels first or you are measuring gain, not spatial quality.

    When upmixing is the right call

    Upmixing earns its place in more situations than its reputation suggests.

    Catalogue and archive. The multitrack is gone, the session is unopenable, the studio no longer exists. There is no alternative, and a careful surround mix derived from the stereo source beats no immersive version at all.

    Beds and ambience. Deriving a surround bed from a stereo atmosphere and placing discrete objects on top is standard practice, not a compromise. The bed carries space; the objects carry intent.

    Live and broadcast. Real-time immersive mixing of a live event with full manual control is often not realistic. A robust upmix in the chain is what makes an immersive broadcast feasible at all.

    Volume work. When hundreds of episodes need an immersive deliverable, hand-mixing each one is not a budget problem — it is an impossibility.

    When it is not

    Upmixing cannot invent intent. If the creative idea requires a specific element to move behind the listener, to sit above them, or to be somewhere the stereo mix never implied, no upmixer will produce it. It can only redistribute what correlation tells it is already there.

    It also cannot repair a mix. Upmixing a dense, heavily limited master gives you a wide, dense, heavily limited master. The spatial impression increases; the underlying problem does not go away, and often becomes easier to hear.

    And it cannot survive a bad monitoring decision. If the upmix is judged only in the room and never checked in its downmixed versions or on headphones, the delivery will surprise someone later.

    Most immersive mixes were originally made for loudspeakers, and headphone listeners get what is left over. That is a mixing decision, not a format limitation — and it applies to upmixed material just as much as to hand-made mixes.

    Painting of a person with headphones singing into a microphone while playing guitar Two chrome spheres in a symmetrically mirrored blue room

    A working order of operations

    If you take one practical thing from this article, take the sequence. Most upmixes fail because the checks happen in the wrong order, not because the wrong plugin was chosen.

    Start by listening to the stereo source on its own and deciding what actually needs to be wider. Material that is already dense rarely benefits. Then set the centre extraction before touching anything else, because every other parameter depends on how much of the correlated content has been pulled out. Judge it on the lead element — vocal, dialogue, snare — and stop as soon as it starts to sound detached from the rest of the mix.

    Only then open up the surrounds, and only then the height layer. Add the vertical component last and less than feels right in the moment; height is the parameter engineers most reliably overdo, because the first impression of it is so convincing. Finally, before you call it finished, fold the mix down and listen to the downmixed versions at matched level. If the fold surprises you, the problem is in the centre extraction, not in the fold.

    Downmix compatibility is the real quality gate

    Object-based delivery changed where the downmix happens. With formats such as MPEG-H, you produce for the largest playback setup and the adaptation to smaller configurations is performed by a downmix algorithm inside the renderer of the end device, with downmix factors that can be defined during production. Dolby Digital Plus with Joint Object Coding takes a related route, transmitting a multichannel downmix plus parametric information from which the decoder reconstructs the objects — carrying Atmos at bitrates as low as 384 kbit/s.

    For upmixing this has a direct consequence: your output audio is not the last word. Something downstream will fold it, and that fold is where a careless upmix falls apart. Check the two-channel result. Check the 5.1 result. If the centre extraction was too aggressive, the vocal will come back louder than intended. If the ambience extraction was overdone, the fold will sound hollow rather than wide.

    Where this is going

    We are, with 3D audio, roughly where stereo was in its early years. It took decades before stereo sounded as good as it does now, and the first attempts were more than bumpy. There are still barely any rules of thumb that transfer cleanly from stereo to 3D.

    That is the context in which I find upmixing genuinely interesting rather than merely pragmatic. Upmixing algorithms in cars already produce impressive spatial experiences from nothing more than a stereo file. Every listener with a modern playback chain is already hearing upmixed audio, whether or not anyone produced an immersive version for them.

    The version of a song that keeps improving over time is, to me, a more exciting prospect than being locked to the best mix that was possible on release day. That does not make hand-crafted immersive mixing obsolete — the blind test showed that hand-crafted is no guarantee either. It means the interesting question is no longer whether upmixing is legitimate. It is which material deserves a dedicated mix, and which is better served by an algorithm that will be better next year than it is today.

    If you want the broader picture of how sound is positioned in three-dimensional space in the first place, audio spatialization covers the underlying practice.


    This website uses cookies. If you continue to visit this website, you consent to the use of cookies. You can find more about this in my Privacy policy.
    Necessary cookies
    Tracking
    Accept all
    or Save settings