On television a stadium always sounds equally far away. You hear a wall of crowd plus the commentary, and both of them sit at the front, inside the set. Put the same broadcast on headphones with binaurally recorded sound and something simple yet unfamiliar happens: you are suddenly sitting in a seat. The cheering comes from behind at an angle, the drum stands to the left, and when the ball goes in you hear where it happened.
For a live sports stream I built exactly that part. It was about the atmosphere of a big match, recorded from three positions and delivered as a binaural track inside the stream. I am not allowed to say anything about the client or the competition, but I can talk about the work.
Almost everything else I do gets a second attempt. A film is mixed, watched, changed. A VR application is built, tested, improved. A live broadcast has none of that. It runs for nearly three hours in one go, and whatever sound comes out of that time is the result.
On top of that, a stadium is not a friendly place acoustically. It is loud, the level jumps twenty decibels within seconds, and the interesting events do not announce themselves. A microphone set correctly for the normal state distorts on the goal, and one set for the goal delivers mostly noise for the remaining two hours. The answer lies in the resolution: at twenty-four bit you can record far more cautiously than would ever have paid off in stereo, and lift the quiet part afterwards without it hissing. So you deliberately record too quietly and keep the peaks the room they need.
The third point is the actual reason for the effort. A stadium does not have one sound, it has several. Behind the home goal it sounds different from behind the away goal, and both sound different from the main stand. Anyone recording only one position gets an opinion about how the match sounded, and anyone recording three can decide in the mix.
Different methods stood at the two goals, and that was deliberate. On the home side an ambisonic microphone, a full sound field in four channels that can be turned in any direction in post. On the away side a device recording front and rear separately, from which an all-round picture can be assembled. In the stand a recorder capturing mid-side and XY at the same time, two stereo methods in parallel out of the same capsule arrangement.
That doubling is not a luxury but an insurance policy. Mid-side can still be changed in width afterwards, XY cannot, but XY is less prone to phase problems when you fold down to stereo later. On a recording that cannot be repeated you do not settle such questions in advance, you take both.
Three recorders running independently for three hours bring a second problem you have to solve in advance. Their clocks drift slightly apart, and what is in sync at the start sits visibly offset at the end. With music that is obvious immediately, with a crowd it is not, but it sounds diffuse and you cannot find why. It helps to have a common event in the material that all tracks can be aligned to, and to check the offset across the whole length rather than only at the start.
Then there is the plain logistics. Recording positions in a stadium are not freely chosen, they have to be accessible, safe and able to run unattended during the match. A recorder you cannot reach in the second half has to be set up beforehand so that it runs through without intervention. That sounds banal and is the part where recordings like this most often fail.
In the stream there were two layers in the end. The commentary stays fixed at the front, because you have to be able to follow it and because a voice that wanders when you turn your head is simply annoying in a broadcast. The atmosphere lies binaurally around it. The contrast between the two is what works: because the voice sits clearly in one place, the rest becomes readable as room.
What does not happen is just as important. I did not inflate the atmosphere. The temptation is strong to mix stadium sound so that it impresses, with plenty of level and plenty of width. After twenty minutes that is exhausting, and the commentary has to fight its way through, which makes it harsh. A livestream is not judged by its peak but by the hour before it.
What remains is a binaural sum across the full length of the broadcast, close to three hours, plus the individual positions as material for whatever comes out of it later.
Recordings like these are not worthless after the final whistle, quite the opposite: edits, trailers and retrospectives draw on them, and then it makes a difference whether there is a stereo track or a sound field you can serve any viewing direction from.
One thing has to be thought through honestly. Binaural only unfolds over headphones, and part of the audience watches on laptop speakers or a television. A binaural track does not sound wrong there, but the effect you went to the trouble for disappears. So a production like this always comes with the question of what share of viewers actually wears headphones, and whether it is worth keeping a separate version for the rest.
Sport in spatial sound is not new territory for me. For ranFIGHTING I recorded a boxing match in 360°, for a Bundesliga spot I made the VR version, and on wheelchair basketball the task was capturing a fast game with ambisonics. What was different here was the length, and the fact that it went out live. For the same client I had previously run ambisonic microphones from three positions at another live sports broadcast, where the sound went straight into the stream.
On a live production everything is decided on the recording day, because there is no second one. Let’s talk beforehand about how many positions it really takes.
Get in touch