Content
MPEG-H Audio is an ISO-standardised audio codec for next generation audio. It carries channel-based, object-based and scene-based content in the same bitstream, and it lets the viewer change the mix at home. Unlike most spatial audio formats it is not a marketing term but a codec, which makes it one of the few things in this field you can compare like for like.
Disclaimer: the first draft of this article came out of an AI. I rewrote it and brought the deployment status up to date, because most of what is written about MPEG-H online is several years stale.
MPEG-H Audio is an ISO standard from the same family as MP3, MPEG-4 and AVC/H.264, developed at Fraunhofer IIS. It supports channel-based, object-based and scene-based transmission and any combination of those, plus interactivity and personalisation.
The last two words are the part that matters. With MPEG-H, the broadcaster can send the commentary as a separate object and let you turn it down. That is not a feature of the loudspeaker layout, it is a property of the bitstream.
No, and this is the confusion the whole comparison runs aground on. MPEG-H Audio is a codec. Dolby Atmos is an umbrella term for Dolby’s immersive sound experience, delivered over several different codecs: Dolby Digital Plus, Dolby TrueHD, AC-4.
So when you read “Dolby Atmos” you do not know which codec is underneath, the old E-AC-3 or the new AC-4. Technically, the only clean comparison is MPEG-H against AC-4.
This is where MPEG-H differs from formats that exist mainly in press releases. It is integrated into four television standards: ATSC, DVB, TTA in Korea and SBTVD in Brazil.
South Korea, May 2017. The first next generation audio codec deployed for a 4K UHD television service. Big events tend to be the springboard for new technology, and the 2018 Winter Olympics in Pyeongchang counted as a significant step. Rock in Rio and the Eurovision Song Contest have also been transmitted in MPEG-H.
Brazil, now. This is the strongest evidence that MPEG-H is not a laboratory format. Brazil’s TV 3.0 standard, marketed as DTV+, mandates MPEG-H Audio as its only audio system. President Lula fixed DTV+ as the country’s future television format by decree on 27 August 2025. Globo started the first pilot station in Rio de Janeiro in April 2025, and in June 2026 DTV+ moved from pilot to commercial operation, timed deliberately for the football World Cup. Globo has been broadcasting 24/7 in Rio de Janeiro, São Paulo and Brasília since then, with more regions following. Rede Amazônica and TV Cultura run continuous MPEG-H services too, and lean particularly on the personalisation options.
Two details with leverage. DTV+ requires every receiver to support MPEG-H Audio over HDMI, and the first sets from Aquário, Intelbras and Vivensis are on sale in Brazil. And alongside MPEG-H, DTV+ also mandates VVC, MPEG-5 LCEVC and ATSC 3.0 transport.
China is introducing the format in a new broadcast standard as well.
The DVB consortium decided to support exactly two next generation audio formats: Dolby AC-4 and MPEG-H. Both can carry multichannel and object-based content, and both do it at comparatively low bandwidth.
The advantage of object-based audio here is channel independence, because rendering happens in the end user’s playback system rather than in the studio. A system with NGA support needs a matching decoder for that. The whole thing is steered by metadata carried as a separate track, holding loudness relationships, the positions of the audio objects and interaction parameters. The framework used for that is the Audio Definition Model.
| Feature | What it means at home |
|---|---|
| Dialogue enhancement | Raise the voice against the crowd or the music |
| Audio description | A separate accessibility object, not a separate channel |
| Multiple languages | Selected in the receiver, not in the transmission |
| Object positioning | Move an object, for example the commentary, to a different place |
| Immersive playback | Rendered to whatever speaker layout is present, or to headphones |
None of these require a specific speaker setup, which is exactly the point. The same bitstream serves a soundbar, a 7.1.4 room and a pair of earbuds.
Fraunhofer’s MPEG-H Authoring Suite covers the production, delivery and playback chain. The tools you will meet:
The authoring tool is where the personalisation is defined: dialogue separation, audio description, multiple languages, interactive object positioning. All of it independent of the speaker configuration at the far end of the chain.
Here the picture is thinner. AC-4 and MPEG-H are both in use at various streaming services, but so far only for the immersive aspect, not the interactive one. The still modest catalogue of 3D music is usually only available at a surcharge on the normal subscription, and outside headphones there are few products that can play next generation audio at all.
So the honest summary is that MPEG-H’s strength is broadcast, and its music story is still waiting for a reason to exist. For the music side, see Dolby Atmos and the wider MPEG-H versus Dolby Atmos comparison.
MPEG stands for Moving Picture Experts Group, a subgroup of the ISO. MPEG-H is that group’s media transport and coding suite; the audio part of it is called MPEG-H 3D Audio.
The comparison is not well formed, because Atmos is an experience label and MPEG-H is a codec. Against AC-4, MPEG-H’s distinguishing feature is how far its interactivity and personalisation go, and that it is an open ISO standard rather than a licensed ecosystem.
Only occasionally. DVB supports it, but European broadcasters have largely stayed with existing formats. The regular services are in South Korea, Brazil and increasingly China.
The MPEG-H VVPlayer, part of Fraunhofer’s authoring suite, which plays encoded MPEG-H MP4 files with or without video so you can check a master before delivery.
A decoder, yes. In Brazil the DTV+ standard forces this by requiring every receiver to support MPEG-H over HDMI. Elsewhere it depends on the device.
It can be. It supports channel-based, object-based and scene-based content, and combinations of all three in one stream. That flexibility is why broadcasters picked it.
Delivering next generation audio and want to be sure the metadata survives the chain?
Get in touch →
Back to blog