Content
An immersive audiobook places voices and sounds around the listener instead of between two channels, so a scene has a direction and a distance. Technically it is binaural or object-based audio applied to spoken-word storytelling. Whether it is worth doing depends on one question, and it is not a technical one: does the story need the space?
Disclaimer: the first draft of this article came out of an AI and was mostly enthusiasm. I rewrote it around what I have actually learned producing spatial audio for narrative work.
A conventional audiobook is a voice in stereo, usually close-miked and dry, sometimes with music. An immersive audiobook adds two things: sources that have positions, and rooms that have acoustics. The narrator can stay centred while a character speaks from the doorway and rain falls on the roof above you.
The delivery format is almost always binaural, because binaural is a normal two-channel file that plays on any platform without a decoder, a subscription or a specific device. Front and back only work approximately, but compared to normal stereo it works pretty well, and that trade is the reason the format is practical at all.
| Normal audiobook | Immersive audiobook | |
|---|---|---|
| Voices | Centred, dry | Positioned, with room |
| Environment | Music or nothing | Ambience with direction |
| Playback | Anything | Headphones required |
| Production | Read, edit, master | Scripted spatially, then mixed as a scene |
The last row is the one that decides the budget. You cannot spatialize an existing audiobook convincingly, because the recording was made without the space in mind.
The whole concept has to be tellable in 3D audio. Otherwise it just sounds like someone moved 3D objects around fairly arbitrarily in post-production.
And the harder half of that rule: if the storytelling works just as well in stereo, it is better to do the project in stereo.
Three questions I settle before anything is recorded. How realistic should the worlds sound? How believable should they be? And above all, how can 3D audio support this story? If the answers do not point anywhere, the format is decoration.
The Virtual Barbershop is still the clearest demonstration of why this works. The makers used exactly the same technique as the ASMR community: a dummy head, scissors and similar sounds right at the ear.
What they did better was build a story. The piece gives you the feeling of getting a haircut. That is the difference, and it is not a technical one.
This is what I keep trying to get across. You should not use the technique because you can use it, but ask which use cases it opens up and what story you can tell. That is where the sweet spot is, and that is where the value is. Podcasts and everything happening on the audiobook platforms could benefit from 3D, when it is used inside a story that makes sense of it.
Two routes into a spatial audiobook, and most productions use both.
Record with a dummy head. Two microphones sit where a human’s ears would be, inside a head with modelled pinnae. The head shadows the far side, the pinna filters direction-dependently, and the time difference between left and right survives. Those three differences are what your hearing derives direction from, and they end up in an ordinary stereo file.
This is the right choice for anything diffuse and wide. You cannot really rebuild the sound of a river with a spatializer. You have to put a dummy head out there and record it, and the result is far more natural than anything you can reconstruct from a dozen close microphones.
Build it with a spatializer. Record dry, place each source with HRTF filtering afterwards. This gives you control and the ability to change your mind later. Dialogue almost always belongs here: a voice recorded dry stays intelligible, and intelligibility is not negotiable in spoken-word work.
In practice the two get combined. The room comes from a recording, the voices are placed into it. That way you keep the credibility of a real space and the control over the words.
The spatial decisions happen before anyone records, not after. Where does the narrator stand? Does a character move through the scene, and if so, along which path? Is the listener a participant or an observer? Those answers change the script, the session and the mix, which is why retrofitting an existing audiobook does not work.
Two practical constraints are worth knowing in advance. Headphones are required, so anyone listening on a kitchen speaker gets a downmix without the layer you spent the budget on. And the mix cannot be judged on a phone. If it is going to be approved, say up front which playback is required and supply the binaural version straight away, so nobody reviews the wrong file.
It pays off when the space carries meaning: a room the listener should feel enclosed by, a voice that has to come from behind, a scene where distance is part of the plot. Documentary and reportage benefit for the same reason, and so does anything built around presence rather than plot.
It does not pay off for a straight non-fiction read, for anything primarily consumed on speakers, or when the format is meant to be the selling point rather than the storytelling. The technology is never the reason. The story is.
If your project might be one of the first kind, the earlier we talk, the less has to be rescued later. How spatial sound works underneath all this is covered in how spatial audio works.
Get in touchRelated Articles
3D Audio - the immersive spatial soundtrack from all directions
360 Microphone for 3D Audio Recording in VR
Audio Definition Model (ADM) supercharges 3D Dolby Atmos
360 Production Sound - Creative Thoughts planning VR Audio
VRTonung learning - THE spatial audio course for immersive media