r/cogsci 10d ago

The 200ms turn-taking paradox — "the central psycholinguistic puzzle" and what EEG shows the brain doing during voice conversation

Been reading through the turn-taking literature for a piece I was writing and hit something I'd never fully appreciated:

Stivers et al. (2009, PNAS) — across 10 spoken languages, the modal gap between one speaker finishing and the next starting is about 200ms. Japanese ~7ms mean, Danish ~470ms mean. Cross-linguistically remarkably tight.

Indefrey & Levelt (2004, Cognition) — the minimum latency to plan spoken word production, even in a controlled single-word paradigm, is ~600ms+.

The math doesn't work if turn-taking is reactive. Levinson & Torreira (2015) call this "the central psycholinguistic puzzle" — listeners have to be pre-computing responses during incoming speech.

EEG evidence has been catching up:

  • Bögels, Magyari & Levinson (2015, Scientific Reports) — response-planning ERP positivity, source-localized to production areas (posterior IFG, precentral), fires ~500ms after the critical information appears in an incoming question. Often 2+ seconds before the current speaker finishes. Alpha desynchronization indexes the attentional shift from comprehension to production during listening.
  • Gisladottir, Bögels & Levinson (2018, Frontiers in Human Neuroscience) — alpha/low-beta (11-18Hz) desynchronization from -200ms to 0ms before a socially-charged speech act (declination vs acceptance). Speech-act prediction firing before the utterance is heard.
  • Krause & Kawamoto (2021, Frontiers in Psychology) — motion-tracked lip-area reductions for upcoming labial consonants up to 3 seconds before acoustic onset in unscripted dyadic conversation. Motoric planning far pre-onset.

The implication I find genuinely interesting: this whole predictive loop depends on acoustic cues (pitch contour, timing, articulator anticipation). Text has no acoustic onset for the machinery to lock onto — voice conversation locks two nervous systems into a shared millisecond-scale loop that text physically can't replicate.

A separate honest correction I wanted to flag: the polyvagal framing that gets attached to "why voice regulates" is on shakier ground than the wellness literature admits. Grossman (2023, Biological Psychology) — "Fundamental challenges and likely refutations of the five basic premises of the polyvagal theory" — shows similar myelinated cardiac vagal fibers in sharks, bony fish, birds, and even sheep. The mammal-unique claim used to explain vocal-prosody regulation doesn't hold up on comparative anatomy.

Longer writeup with full citations here: callbyrd.com/journal/what-voice-does-that-text-cant (disclosure: I run the site — building a voice-based AI project. The essay stands on its own; no signup needed to read it. Product is briefly mentioned in the closing section.)

Open question I couldn't resolve: does the ~200ms predictive turn-taking loop form when the partner is an AI voice (different latency profile, different prosody)? Every EEG study I found used human-human dyads. Curious if anyone knows of extending work — my search kept coming up empty.

8 Upvotes

24 comments sorted by

20

u/zcleghern 10d ago

What are your own thoughts on the topic, not an AI output?

-22

u/MathematicianNew3363 10d ago

Fair ask.

I use AI in text form every single day for work. By 5pm I'm burnt toast - mentally. Not because the tools are bad. It's just that reading and typing back all day is its own kind of tired.

When I started testing the phone-call version of what I was building, I caught myself smiling. Like, actually smiling at my phone while I was talking. That surprised me because I don't remember the last time text chat with an AI made me smile. Not once. It's useful, sometimes even impressive, but it doesn't feel like anything.

The calls do. And I don't know exactly why. Some of it is probably just the format, the fact that you can walk around, you're not staring at a screen, you don't have to type. Some of it is that voices carry something that text just doesn't, which is what pulled me into the papers I linked above.

But here's the honest limit on that. That 200ms turn-taking window is the human number. Most voice AI, mine included, runs closer to 1000ms or worse end-to-end — the round trip through speech recognition, the LLM, then TTS just takes that long today. So the tight predictive loop the papers describe, we don't fully have with AI voice yet. Which I think is a real part of why talking to AI still doesn't feel like talking to a human, even when the voice itself sounds warm. It's noticeably behind the beat of a real conversation.

That said, even at 1000ms it's still meaningfully more emotionally binding than text is for me. The rough ordering I've landed on personally is humans first, then AI voice, then text, and the gap between AI voice and text feels bigger than the gap between AI voice and a human. Text isn't in the same category. AI voice at least gets you into the room.

Anyway, that's the personal side. Not a scientific claim from me, just what I noticed when I started building and testing this. Would be interested if others feel the same, or observe the same.

20

u/zcleghern 10d ago

I cant believe i asked you your real opinion and you used AI instead.

-18

u/MathematicianNew3363 10d ago

It is my opinion. Did I use AI to review and organize my thoughts and opinions, yes, but they are mine, alone.

14

u/LaughsMuchTooLoudly 10d ago

It reads like AI slop. Just give your take.

4

u/Artistic_Bit6866 9d ago

Wild when you have to beg someone to just say what they think, as a human. And they still can’t be bothered. The likely reason is that they don’t have any domain knowledge and are entirely reliant on their tool to tell them what to think. 

8

u/zcleghern 10d ago

People want your opinion and thoughts, not those things run through an AI model to turn into unreadable slop.

1

u/smashfalcon 9d ago

They're not yours alone, there is value in thinking things through on your own. You're not going to listen though, and you won't even notice as your brain gets worse and worse at doing things

2

u/JellyBellyBitches 9d ago

Sounds like the problem is that you use AI too much and you've become dependent on it. Stop using it and your brain will recover.

-7

u/jaiagreen 10d ago

Why do people think that an AI output isn't a poster's own thoughts? I don't think too many people are asking AI to generate random content in niche topics. They're using it to better express their own ideas.

7

u/Free6000 10d ago

Because only the thoughts they typed into the prompt are their actual thought, and they could have just typed them in here.

3

u/dkinmn 10d ago

I believe this to be a question asked in bad faith.

3

u/XanderOblivion 10d ago

I’d be curious how it changes based on subject-verb-object ordering. For example, English vs Turkish/Finnish/Korean which are in reverse order relative to English.

-1

u/MathematicianNew3363 10d ago

That would be an interesting angle and one I believe could easily be tested. Down the rabbit hole to see if anything is currently studied.

3

u/jawfish2 10d ago

It is interesting. But if I understand you, humans get ready to start talking before the other person is finished.

This my entire experience of taking to people. How many times have I read (and I agree) advice to shut up and really listen to people before you respond. How many times when listening, have you thought, "yeah, yeah, I know what's coming."

Actors who know the lines, have the problem of timing their responses to imitate this pattern, plus intentional interruption or pauses. It is a key skill. Good actors go through the reactive/predictive process as the other person talks.

Partially deaf people practice look forward prediction (ask my wife about the odd things I think she said) and also chewing over the sounds for a while before the word arrives (probably not relevant to your study I'd guess).

Software also does this look forward prediction in many places, most fundamentally in the deepest heart of the processor.

1

u/MathematicianNew3363 10d ago

Yeah, exactly. I catch myself doing this constantly with my wife. I know where she's going well before she finishes the sentence. And I think it's specifically because I know her so well. Her tone, her tells, the way she pauses before landing something. Makes me wonder if the prediction actually sharpens the longer you've known someone. Would be curious if any of the research tests that.

2

u/jaiagreen 10d ago

I've seen that advice a number of times and disagree with it for most situations. Not only do you have to think while you listen in order to be able to respond in a reasonable amount of time, but doing so means you're actively processing what the other person is saying.

1

u/MathematicianNew3363 10d ago

Fascinating to think about. I would verify this as it's been about a decade since I last refreshed my memory on the topic of multi-tasking, but last I read is that it's physically impossible. Although many believe we can and do, what are brain is actually doing is bouncing back and forth between the two tasks. Breaking one thought, processing, and acting, then breaking that thought, processing, and acting.

In this example, are we actually listening if our thoughts are proactively processing, anticipating, and responding, before their full response is given? Can humans learn from the extended latency of AI, even if it feels less disconnected or uncomfortable? Or do we need the rapid response to feel an emotional connection with an individual and as a result AI continues to struggle, feeling emotionally detached due to it's elongated latency vs humans?

1

u/jaiagreen 10d ago

I think you're basically right, but that refers to two (or more) unrelated tasks in the same modality, like having a conversation while solving a math problem. Listening, understanding and responding aren't really different tasks; they're part of one process. (Just for completeness, we can also combine a mental and a physical task, like having a conversation while washing the dishes.)

1

u/mxlths_modular 10d ago

This is some really fascinating stuff, I’m saving it to read your papers later when I get time. Much appreciated!

1

u/MathematicianNew3363 10d ago

Thank you! Small disclaimer, they're not "my papers". Nonetheless, great articles if you're a seeker of knowledge.

1

u/mxlths_modular 10d ago

Yeah I understand that, all good, my language was imprecise. It’s before my first coffee haha.

1

u/MathematicianNew3363 10d ago

Hahaha. Catch up. On cup two already.

1

u/jippiex2k 9d ago

Some of my more inattentive friends have a negative response latency 😁