Illustrierte Person mit Kopfhörern vor einer Wand aus blau leuchtenden Daten
0DoFPodcast & ASMR

Exploring Immersive Audio Books: 3D Soundscapes for Spatial Storytelling

Content

    An immersive audiobook places voices and sounds around the listener instead of between two channels, so a scene has a direction and a distance. Technically it is binaural or object-based audio applied to spoken-word storytelling. Whether it is worth doing depends on one question, and it is not a technical one: does the story need the space?

    Disclaimer: the first draft of this article came out of an AI and was mostly enthusiasm. I rewrote it around what I have actually learned producing spatial audio for narrative work.

    What are immersive audiobooks? The term itself is worth a second look: immersive sound is more than 3D music.

    A conventional audiobook is a voice in stereo, usually close-miked and dry, sometimes with music. An immersive audiobook adds two things: sources that have positions, and rooms that have acoustics. The narrator can stay centred while a character speaks from the doorway and rain falls on the roof above you.

    The delivery format is almost always binaural, because binaural is a normal two-channel file that plays on any platform without a decoder, a subscription or a specific device. Front and back only work approximately, but compared to normal stereo it works pretty well, and that trade is the reason the format is practical at all.

    Woman with headphones in profile, blue particle structures floating around her head

    How do immersive audiobooks differ from a normal audiobook?

    Normal audiobook Immersive audiobook
    Voices Centred, dry Positioned, with room
    Environment Music or nothing Ambience with direction
    Playback Anything Headphones required
    Production Read, edit, master Scripted spatially, then mixed as a scene

    The last row is the one that decides the budget. You cannot spatialize an existing audiobook convincingly, because the recording was made without the space in mind.

    The rule I apply before saying yes

    The whole concept has to be tellable in 3D audio. Otherwise it just sounds like someone moved 3D objects around fairly arbitrarily in post-production.

    And the harder half of that rule: if the storytelling works just as well in stereo, it is better to do the project in stereo.

    Three questions I settle before anything is recorded. How realistic should the worlds sound? How believable should they be? And above all, how can 3D audio support this story? If the answers do not point anywhere, the format is decoration.

    The example everyone should listen to

    The Virtual Barbershop is still the clearest demonstration of why this works. The makers used exactly the same technique as the ASMR community: a dummy head, scissors and similar sounds right at the ear.

    What they did better was build a story. The piece gives you the feeling of getting a haircut. That is the difference, and it is not a technical one.

    This is what I keep trying to get across. You should not use the technique because you can use it, but ask which use cases it opens up and what story you can tell. That is where the sweet spot is, and that is where the value is. Podcasts and everything happening on the audiobook platforms could benefit from 3D, when it is used inside a story that makes sense of it.

    Person with headphones from behind, surrounded by blue rays of light

    The technical side: binaural recording

    Two routes into a spatial audiobook, and most productions use both.

    Record with a dummy head. Two microphones sit where a human’s ears would be, inside a head with modelled pinnae. The head shadows the far side, the pinna filters direction-dependently, and the time difference between left and right survives. Those three differences are what your hearing derives direction from, and they end up in an ordinary stereo file.

    This is the right choice for anything diffuse and wide. You cannot really rebuild the sound of a river with a spatializer. You have to put a dummy head out there and record it, and the result is far more natural than anything you can reconstruct from a dozen close microphones.

    Build it with a spatializer. Record dry, place each source with HRTF filtering afterwards. This gives you control and the ability to change your mind later. Dialogue almost always belongs here: a voice recorded dry stays intelligible, and intelligibility is not negotiable in spoken-word work.

    In practice the two get combined. The room comes from a recording, the voices are placed into it. That way you keep the credibility of a real space and the control over the words.

    Wireframe model of a head with headphones between two microphones

    What this costs you in production

    The spatial decisions happen before anyone records, not after. Where does the narrator stand? Does a character move through the scene, and if so, along which path? Is the listener a participant or an observer? Those answers change the script, the session and the mix, which is why retrofitting an existing audiobook does not work.

    Two practical constraints are worth knowing in advance. Headphones are required, so anyone listening on a kitchen speaker gets a downmix without the layer you spent the budget on. And the mix cannot be judged on a phone. If it is going to be approved, say up front which playback is required and supply the binaural version straight away, so nobody reviews the wrong file.

    When it is worth it, and when it is not

    It pays off when the space carries meaning: a room the listener should feel enclosed by, a voice that has to come from behind, a scene where distance is part of the plot. Documentary and reportage benefit for the same reason, and so does anything built around presence rather than plot.

    It does not pay off for a straight non-fiction read, for anything primarily consumed on speakers, or when the format is meant to be the selling point rather than the storytelling. The technology is never the reason. The story is.

    If your project might be one of the first kind, the earlier we talk, the less has to be rescued later. How spatial sound works underneath all this is covered in how spatial audio works.

    Get in touch

    This website uses cookies. If you continue to visit this website, you consent to the use of cookies. You can find more about this in my Privacy policy.
    Necessary cookies
    Tracking
    Accept all
    or Save settings