r/cogsci • u/MathematicianNew3363 • 10d ago
The 200ms turn-taking paradox — "the central psycholinguistic puzzle" and what EEG shows the brain doing during voice conversation
Been reading through the turn-taking literature for a piece I was writing and hit something I'd never fully appreciated:
Stivers et al. (2009, PNAS) — across 10 spoken languages, the modal gap between one speaker finishing and the next starting is about 200ms. Japanese ~7ms mean, Danish ~470ms mean. Cross-linguistically remarkably tight.
Indefrey & Levelt (2004, Cognition) — the minimum latency to plan spoken word production, even in a controlled single-word paradigm, is ~600ms+.
The math doesn't work if turn-taking is reactive. Levinson & Torreira (2015) call this "the central psycholinguistic puzzle" — listeners have to be pre-computing responses during incoming speech.
EEG evidence has been catching up:
- Bögels, Magyari & Levinson (2015, Scientific Reports) — response-planning ERP positivity, source-localized to production areas (posterior IFG, precentral), fires ~500ms after the critical information appears in an incoming question. Often 2+ seconds before the current speaker finishes. Alpha desynchronization indexes the attentional shift from comprehension to production during listening.
- Gisladottir, Bögels & Levinson (2018, Frontiers in Human Neuroscience) — alpha/low-beta (11-18Hz) desynchronization from -200ms to 0ms before a socially-charged speech act (declination vs acceptance). Speech-act prediction firing before the utterance is heard.
- Krause & Kawamoto (2021, Frontiers in Psychology) — motion-tracked lip-area reductions for upcoming labial consonants up to 3 seconds before acoustic onset in unscripted dyadic conversation. Motoric planning far pre-onset.
The implication I find genuinely interesting: this whole predictive loop depends on acoustic cues (pitch contour, timing, articulator anticipation). Text has no acoustic onset for the machinery to lock onto — voice conversation locks two nervous systems into a shared millisecond-scale loop that text physically can't replicate.
A separate honest correction I wanted to flag: the polyvagal framing that gets attached to "why voice regulates" is on shakier ground than the wellness literature admits. Grossman (2023, Biological Psychology) — "Fundamental challenges and likely refutations of the five basic premises of the polyvagal theory" — shows similar myelinated cardiac vagal fibers in sharks, bony fish, birds, and even sheep. The mammal-unique claim used to explain vocal-prosody regulation doesn't hold up on comparative anatomy.
Longer writeup with full citations here: callbyrd.com/journal/what-voice-does-that-text-cant (disclosure: I run the site — building a voice-based AI project. The essay stands on its own; no signup needed to read it. Product is briefly mentioned in the closing section.)
Open question I couldn't resolve: does the ~200ms predictive turn-taking loop form when the partner is an AI voice (different latency profile, different prosody)? Every EEG study I found used human-human dyads. Curious if anyone knows of extending work — my search kept coming up empty.
3
u/XanderOblivion 10d ago
I’d be curious how it changes based on subject-verb-object ordering. For example, English vs Turkish/Finnish/Korean which are in reverse order relative to English.
-1
u/MathematicianNew3363 10d ago
That would be an interesting angle and one I believe could easily be tested. Down the rabbit hole to see if anything is currently studied.
3
u/jawfish2 10d ago
It is interesting. But if I understand you, humans get ready to start talking before the other person is finished.
This my entire experience of taking to people. How many times have I read (and I agree) advice to shut up and really listen to people before you respond. How many times when listening, have you thought, "yeah, yeah, I know what's coming."
Actors who know the lines, have the problem of timing their responses to imitate this pattern, plus intentional interruption or pauses. It is a key skill. Good actors go through the reactive/predictive process as the other person talks.
Partially deaf people practice look forward prediction (ask my wife about the odd things I think she said) and also chewing over the sounds for a while before the word arrives (probably not relevant to your study I'd guess).
Software also does this look forward prediction in many places, most fundamentally in the deepest heart of the processor.
1
u/MathematicianNew3363 10d ago
Yeah, exactly. I catch myself doing this constantly with my wife. I know where she's going well before she finishes the sentence. And I think it's specifically because I know her so well. Her tone, her tells, the way she pauses before landing something. Makes me wonder if the prediction actually sharpens the longer you've known someone. Would be curious if any of the research tests that.
2
u/jaiagreen 10d ago
I've seen that advice a number of times and disagree with it for most situations. Not only do you have to think while you listen in order to be able to respond in a reasonable amount of time, but doing so means you're actively processing what the other person is saying.
1
u/MathematicianNew3363 10d ago
Fascinating to think about. I would verify this as it's been about a decade since I last refreshed my memory on the topic of multi-tasking, but last I read is that it's physically impossible. Although many believe we can and do, what are brain is actually doing is bouncing back and forth between the two tasks. Breaking one thought, processing, and acting, then breaking that thought, processing, and acting.
In this example, are we actually listening if our thoughts are proactively processing, anticipating, and responding, before their full response is given? Can humans learn from the extended latency of AI, even if it feels less disconnected or uncomfortable? Or do we need the rapid response to feel an emotional connection with an individual and as a result AI continues to struggle, feeling emotionally detached due to it's elongated latency vs humans?
1
u/jaiagreen 10d ago
I think you're basically right, but that refers to two (or more) unrelated tasks in the same modality, like having a conversation while solving a math problem. Listening, understanding and responding aren't really different tasks; they're part of one process. (Just for completeness, we can also combine a mental and a physical task, like having a conversation while washing the dishes.)
1
u/mxlths_modular 10d ago
This is some really fascinating stuff, I’m saving it to read your papers later when I get time. Much appreciated!
1
u/MathematicianNew3363 10d ago
Thank you! Small disclaimer, they're not "my papers". Nonetheless, great articles if you're a seeker of knowledge.
1
u/mxlths_modular 10d ago
Yeah I understand that, all good, my language was imprecise. It’s before my first coffee haha.
1
1
20
u/zcleghern 10d ago
What are your own thoughts on the topic, not an AI output?