Content
Pretty much all technologies for playback of virtual reality audio on headphones can be put into three categories: channel-based, object-based and soundfield-based. Some examples, bevor we go into detail:
Those three are the textbook categories. What actually ships today is almost always a mixture: a channel bed carries the static part, objects carry whatever moves, and underneath there is often a soundfield layer as well. That is why hybrid formats get their own section further down.
If you are just getting started, read how to start with virtual reality sound first.
Which format you go with is rarely decided by how it sounds and almost always by the chain behind it: what the target platform accepts on upload, what the player renders from it, and how much of that actually reaches the listener. So the categories below are not a matter of taste, they are a decision about compatibility.
Every channel of an audio-file is routed to a fixed place for playback. For stereo, it is “left” and “right” and therefore has two channels. 5.1 Surround has six channels, which means there are two additional loudspeakers at the back of the listener, as well as a center between left and right, but also an LFE-channel for the subwoofer.
Binaural basically means that you are able to get a surround experience, just via headphones and therefore only has two channels. It simulates how a human is hearing by recording with special microphones (a dummy head with two artificial ears) or it can be calculated as a downmix from a surround-formats with something called HRTF (Head-Related Transfer Function).
| Pro | Contra |
|---|---|
| Can be played back on every platform, just like a normal stereo-file | No playback on loudspeaker – technically possible, but sounds weird |
| Quick possibility to listen to the tone of a spatial audio mix | No head-tracking possible, it can only deliver the sound of a fixed viewing direction |
| Pro | Contra |
|---|---|
| Can be played back not only on headphones for VR but on a conventional loudspeaker setup | positions in between the loudspeakers are only realized as phantom sound source |
| Easy setup, used and supported for ages | only two dimensional, doesn’t support height information from above or below |
Here, sounds are placed as so-called audio objects in 3D space without being bound to loudspeaker arrangements or channels. During the later playback, the position of the object in the room is calculated on the available loudspeakers, thus an acoustic irradiation with an almost unlimited number of loudspeakers is possible and represents in the arrangement in about a hemisphere.
Hardly any format that actually ships is purely object-based, though. Dolby Atmos is a channel bed plus objects, and G’Audio Lab uses channel, object and soundfield at the same time — strictly speaking, both entries below are hybrids.
The Dolby Atmos tools got transformed to be used for virtual reality. The Atmos-master file is transcoded to an ec3-file, which later supports playback with head-tracking.
| Pro | Contra |
|---|---|
| It is possible to playback the file without further decoding since it recognizes, the current situation and converts itself from surround to stereo | baked-in format, few possibilities for distribution, even Previewing your own video can be complicated |
| object-based approach on VR, playback during mixing is easily possible on surround loudspeaker setups | Its VR Transcoder can output ambisonics, but it is only first order and will sound worse than a Dolby’s ec3 |
The team behind it already started working on MPEG-H in 2005 involved in its binaural rendering. But they knew that MPEG-H is not perfect for VR since it is not possible to use channel-, object-, and soundfield-based audio at the same time.
| Pro | Contra |
|---|---|
| Uses the benefits of channel-, object-, and soundfield-based audio each | Not support on platforms, but encoding workaround to ambisonics possible |
| Can also be used for interactive VR, so moving away from the camera (6 instead of 3 degrees of freedom) | Mac-only at the moment |
For this format, I already wrote a more detailed article, right here Mixing with ambisonics for virtual reality audio is relatable to working with object-based audio. But the technologies are way different.
These two Ambisonics formats are pretty similar and compatible with each other, so I will not distinguish between them.
| Pro | Contra |
|---|---|
| High compatibility with other channel-based formats with decoder | No playback without decoder possible but it’s implemented on most platforms |
| Scalable with more channels, for 360° videos, four channels are already making fun | Music is usually static and not supposed to rotate with the scene, but stereo is only supported with a workaround which is not lossless (head-locked) |
It’s its own ambisonics format, which company got bought by Facebook and got introduced as its standard. It’s a hybrid higher order Ambisonics, which uses eight channels, a well thought through concept, but also some flaws. A real production with a dubbed version shows this in the Humboldt 360° case for the Goethe-Institut. More on that perhaps on my blog.
| Pro | Contra |
|---|---|
| Good compromise of channel number and possible resolution; it supports an additional static stereo-track which solves the ambisonics problem | baked-in format, it is difficult to bring it to other formats, but it was recently improved |
| complete pipeline from DAW to SDK, you won’t here big surprises throughout the process | nontransparent: what do the channels stand for, what kind of HRTF is being used etc. Free to use, but not open source |
Is a format, that can be classified somewhere between channel-based and soundfield-based. It relies on four stereo-files which represent four lines of sight at 0°, 90°, 180°, and 270°. During playback, the audios are interpolated for angles in between these fixed numbers. Although in the future it will probably be used less, it still has its right to exist.
| Pro | Contra |
|---|---|
| When programming apps, there is no need to implement an HRTF with a decoder, which saves resources | playback is mostly only a mix of the audio and therefore not very accurate |
| the stereo track at 0° represents a downmix (see binaural stereo) which can be useful as a preview without head-tracking | It is possible to go e.g. from ambisonics to quad-binaural, but not vice versa, so it’s a dead end for post-production |
How a file is split up decides how it behaves on playback: channel bed and soundfield layer rotate with your head, a head-locked track deliberately stays put, and objects only land in the right place if the player reads their metadata. Two Big Ears was already built that way, and so are Dolby Atmos VR and G’Audio Lab. Two newer formats push the principle much further.
Combines objects, channel-based audio and higher order ambisonics in one container, head-tracked and head-locked side by side, encoded as APAC (Apple Positional Audio Codec). On working with it, see my article about the ASAF plugins.
| Pro | Contra |
|---|---|
| All three approaches in a single authoring format, head-locked and head-tracked at the same time | Apple ecosystem: playback essentially only through Apple’s own chain, the Vision Pro above all |
| APAC is practically transparent and carries objects as well as ambisonics | For 360° inside APMP the object part has to be rendered onto an ambisonics bus first — a bed left empty ships silence |
The open counterpart, driven by Google and Samsung inside the Alliance for Open Media: channel beds, ambisonics and objects in one container. More on it in my article about IAMF and Eclipsa and in the sound check with real files.
| Pro | Contra |
|---|---|
| Royalty-free and openly specified, and YouTube accepts it on upload | The Base profile is capped at 18 channels across at most two audio elements — the example in the spec is third order ambisonics (16 channels) plus head-locked stereo (2 channels) |
| The Open Audio Renderer v1 from 30 July 2026 renders channel-based, scene-based and object-based audio | Objects only appear in the newer draft; they are not part of the path YouTube takes today. And for 360° video there is no finished player, you build it yourself |
Which format actually survives on which device and in which player is something I keep track of in my overview of spatial audio support in 360° players.
So, that was my little overview of cinematic virtual reality audio. If you have any questions, comments or feedback, feel free to write a mail.
The choice is rarely between good and bad, it is between compatible and controllable: channel-based plays everywhere and can do the least, ambisonics is the standard for 360° video and scales with its order, and hybrid formats like ASAF and Eclipsa can do the most while having the fewest players. So if you have to decide, work backwards from the playback path to the mix, not the other way round.
For the soundfield side there is my article on ambisonics; which format actually survives on which device is in my overview of the 360° players.
Are you working on a VR experience and the audio still feels flat?
Get in touch →