Sound arrives before speech
The interpreter hears the floor a fraction before the audience does, and that fraction is engineered.
The interpreter hears the floor a fraction before the audience does, and that fraction is engineered.
The deliberate gap
Walk into any conference hall running simultaneous interpretation and you will notice something odd: the ambient sound in the room feels slightly out of sync. The speaker's voice reaches the audience through loudspeakers a beat after it leaves their mouth. That delay is not an accident or a flaw in the system — it is a design decision, and it belongs to the interpreter.
Inside a booth, the interpreter receives the floor audio directly from the microphone circuit, with almost no processing latency. The audience, by contrast, hears the same speaker through amplification and room acoustics that introduce a measurable delay — typically in the range of tens to a few hundred milliseconds depending on the system. The interpreter is therefore slightly ahead of everyone else in the room. That head start is what makes simultaneous interpreting possible.
Sightline, air changes, sound insulation and console layout are all specified, because the booth is the working environment and not a cupboard. A booth is a specification
The cognitive demand of simultaneous work is brutal: the interpreter must parse meaning, hold it in short-term memory, begin producing an equivalent in the target language, and monitor their own output — all while the next clause is already arriving. Even a fraction of a second of additional lead time reduces the pressure. Accuracy degrades fast under sustained load, and any architectural buffer helps.
Engineering the offset
The principle has a name in audio engineering: acoustic delay compensation. Loudspeakers in large venues are stacked, arrayed and timed so that the sound from every source reaches the listener's ear as a single coherent wave — a technique formalized in modern sound design through standards covering large-venue reinforcement. Where booths are present, the configuration is adjusted so that floor reinforcement runs slightly late relative to the raw microphone feed reaching the interpreter's headset console.
The ISO specifications covering booth design govern sightlines, acoustics and air changes; the audio routing that creates this offset is a separate layer, managed by the conference audio engineer in coordination with the venue. It is not guaranteed by the booth standard alone — it requires deliberate commissioning.
The first equipment was a philanthropist's idea, an engineer's build and a great deal of cable. Filene, Finlay, a box of headsets
The effect is modest and the interpreter does not consciously experience it as "extra time." What they experience is the absence of a specific kind of panic: the sensation of being simultaneously behind and expected to speak. Remove the offset — route the floor audio through the same processing chain as the loudspeakers — and interpreters report the work becoming harder. The milliseconds are real.
The audience hears a fluid voice in their language and assumes fluency. What they are really hearing is the downstream product of a carefully engineered latency gap that nobody in the room is supposed to notice.