Thirty Minutes
Why the booth is always staffed in pairs, and what happens when it isn't
Interpreters swap at roughly half-hour intervals because accuracy measurably degrades after that, which is why the job is staffed in pairs.
Every interpreter booked for a conference is booked as half of a unit. The programme says "two interpreters per booth"; the contract, whether issued under AIIC guidelines or an institutional framework, specifies the same minimum. The reason is not comfort or tradition. It is that human cognitive output under simultaneous interpreting load degrades measurably after roughly thirty minutes, and the work does not forgive degraded output.
What thirty minutes marks
Simultaneous interpreting is one of the most cognitively demanding tasks that can be performed routinely. The interpreter listens to continuous speech in one language, holds the incoming meaning in working memory, reformulates it in a second language, monitors their own output for accuracy, watches the room for pace and register cues, and does all of this in real time with no pause between cycles. The lag between hearing and speaking — the ear-voice span — is typically two to four seconds for fluent simultaneous work, but during dense passages, technical terminology or fast delivery it can stretch, and the cognitive cost of stretching it is immediate.
When no interpreter has both languages, the speech goes through a third — and every relay adds delay and one more chance to lose something. Relay, and what it costs
Research into the psycholinguistics of simultaneous interpreting has tracked this degradation in laboratory and field conditions for decades. The most consistent finding is that errors of omission, mistranslation and target-language fluency breakdown cluster in the period beyond approximately thirty minutes of continuous work. The number is not a hard cliff — individual capacity varies, subject matter affects it, pace of delivery affects it — but thirty minutes is the working consensus, adopted by AIIC as its standard guideline for continuous time in the active seat. The partner in the booth is not there to rest; they are there to take over before the first interpreter's output starts to slip.
The swap itself is silent and wordless. The active interpreter raises a hand, or taps a notepad, or simply catches the partner's eye and nods toward the microphone. The partner takes over mid-sentence if necessary. The transition should be invisible to the delegate floor: no gap in audio, no change in quality. Achieving that invisibility requires that the partner has not been absent from the booth or the channel — they have been following the proceedings the entire time, monitoring the incoming speech, keeping their own working memory primed, noting speaker names, terminology and any numbers that have come up. Passive time in the booth is not rest; it is low-load attention, held ready.
What happens when the pairing fails
The thirty-minute interval is easy to understand in principle and difficult to protect in practice. The pressures against it are real: sessions run long, a speaker delivers a monologue that should have been three separate agenda items, the meeting room does not break when it was scheduled to break. An interpreter who has been active for fifty or sixty minutes without rotation is not dramatically impaired in any visible sense — they can still speak, still maintain a surface fluency — but the error profile changes. Numbers are the first thing to slip, then proper nouns and lists, then the precise connective tissue of an argument: whether a causal claim is affirmed or negated, whether a condition is necessary or merely sufficient. These are exactly the categories of precision that conference delegates most need to receive correctly.
Under load the first things to go are numbers, names and lists, which is why they are written down rather than remembered. What the ear drops first
In the formal institutional context — the United Nations Office at Geneva, the Directorate-General for Translation and its interpreting counterpart, the Court of Justice of the European Union in Luxembourg — the pairing rule is enforced by meeting management, not left to individual discretion. At the court in particular, where the spoken record feeds directly into legally binding judgement, there is no procedural tolerance for an undermanned booth. The institutional cities where this work is concentrated — Geneva, Brussels, Luxembourg — have built their conference infrastructure around the assumption that interpreting is a paired discipline, and their meeting schedules are constructed accordingly.
In commercial and smaller conference settings the pairing rule is fragile. A client who has not been told why interpreters come in pairs sees a second person sitting apparently idle and wonders if the cost is justified. The answer is that the second person is not idle, and the cost of dispensing with them is not savings but risk — specifically, the risk that a critical passage of a long session arrives at the eighty-minute mark and is rendered with the precision of someone who has been running a cognitive sprint since the opening gavel. This is why reputable conference organisers build the pairing into the brief from the start, before the event design has locked in a schedule that cannot accommodate it.
The number behind the norm
Thirty minutes is the number that practitioners name, but the discipline behind it is older and more carefully built than a rule of thumb. The psycholinguistics of simultaneous interpreting — attention, working memory, the costs of divided processing — were being studied systematically from the 1970s onward, and the literature includes both laboratory studies under controlled conditions and field observations at operational booths. The consistent picture is of a performance curve that remains relatively stable through the first portion of a session and then bends: errors increase, self-monitoring decreases, output fluency holds up longer than accuracy does. That last point matters, because a delegate listening to an interpreter who is still speaking fluently has no signal that the accuracy has begun to erode.
Simultaneous interpreting is one of the most cognitively demanding tasks that can be performed routinely
The pairing model is the structural response to that hidden degradation. It does not ask the individual interpreter to know when they have peaked — cognitive load research suggests that accurate self-monitoring is itself one of the first casualties of sustained overload. It builds rotation into the system so that the threshold is never individually determined. Thirty minutes, and the other person takes the microphone.
What this means for anyone who commissions or manages conference interpreting is that the pair is not a staffing preference or a union requirement, though it appears in professional agreements because it is well-founded. It is a technical specification, in the same category as soundproofed booths and audible floor sound at the correct level: the conditions under which the output will be what you paid for. Remove any of them and the output changes, quietly, in ways that are invisible until they are not.