Skip to content

Latest commit

 

History

History
121 lines (61 loc) · 17.3 KB

File metadata and controls

121 lines (61 loc) · 17.3 KB

The Voice Behind the Voice: Why AI Narration Exhausts Listeners

It started with a personal experience: I watch a lot of instructional videos on programming, AI engineering, and linguistics. I noticed a pattern: sometimes the narrator's voice seems very correct and articulated but it requires an unusually high amount of effort to process the meaning. I thought for a while that it was because of the complexity of topic, but that wasn't the case — other lectures with the same difficulty came much easier. I started to investigate, and it turned out that these 'tiresome' videos use AI-generated audio. I listened further and got an insight into what was wrong. Somehow, the speaker did not seem to understand what they were saying, and something about that felt fundamentally off — as if the voice tried to transmit understanding that was never there.

There is a widespread assumption behind the push for better AI voices: that listener fatigue is essentially an acoustic problem. If the voice sounds more natural, more human-like, processing it should require less effort. Better synthesis, easier listening.

The data don't support this. In controlled experiments using pupillometry — a physiological measure of cognitive effort — the highest-quality AI synthesizer produced more cognitive load than lower-quality systems, and more than human speech (Govender, 2023). The effect shows up quickly, and persists even when listeners can transcribe the audio without errors. Intelligibility is intact. Something earlier in the process is costing effort.

The usual suspects don't hold up under scrutiny. The uncanny valley predicts that discomfort should decrease as synthesis quality improves — the cognitive load data point the other way. A simple attitude or habit explanation is equally hard to sustain, since the effect registers in physiological measures rather than self-reported preference.

So what is it?


Prosody

To understand what's happening, it helps to start with prosody — the melody of speech: its rhythm, pitch, pauses, and emphasis.

Prosody is the backbone of spoken language. Listeners use prosodic cues to recognize words, resolve syntactic ambiguity, and organize what they hear into meaningful chunks. When prosodic signals are absent or unnatural, the brain loses the supporting structure it relies on automatically. Processing becomes shallower, and paradoxically, more effortful (Cutler, Dahan & van Donselaar, 1997).

But prosody does something beyond organizing structure: it carries information about the speaker's communicative act.

Experiments show that listeners can identify what a speaker is doing — criticizing, suggesting, warning, expressing doubt — from the prosody of a single word, without context and without lexical meaning. The identification is reliable, fast, and independent of emotion: even after the emotional coloring of an utterance is statistically removed, the illocutionary character of the speech act remains readable in the acoustic signal. Listeners hold an internalized repertoire of prosodic patterns for different communicative acts, and they match incoming speech against it — continuously, automatically, as a reflex rather than a choice (Hellbernd & Sammler, 2016).

This is a foundational capacity: infants use prosody to express communicative intentions before they have words. The capacity doesn't disappear with language acquisition — it persists beneath it, running in the background of every listening experience.


The trace of understanding

Human prosody is shaped by many factors, but a speaker's understanding of the material is central to how emphasis, pacing, and information structure get organized. A lecturer often slows down where the concept is difficult to articulate. Emphasis falls where the speaker considers something important. A pause before a key term reflects real weighing of words. These are functional traces of the speaker's cognitive and communicative process.

The speaker has a relationship to the material. The speaker is modeling what the listener knows and doesn't know, and shaping the delivery accordingly. This does not necessarily require expert knowledge of the subject. If complete subject-matter understanding were required for motivated prosody, professional voice actors would sound almost as problematic as TTS systems when reading unfamiliar material. They generally do not — because a skilled reader can still construct a communicative model of the text: recognizing where a new term is introduced, where a key claim is made, where an argument turns. What matters is not expertise in the subject, but the ability to build and express that internal representation. A TTS system does not build it. The prosody it produces is statistically plausible — trained on human speech, it reproduces the patterns associated with different sentence types and structures. But the patterns are generated without the layered communicative representation that determines where emphasis falls in human speech. There is a signal, but no communicative event behind the signal.

This distinction matters because the listener's brain is built to read exactly that signal — and to use it to construct a model of another mind. Neurolinguistic research shows that understanding speech is inseparable from processing the speaker's intention. When listeners hear an utterance, they don't just decode its semantic content — they automatically build a model of what the speaker intends to communicate, what the speaker believes, what the speaker expects the listener to do with the information. This processing involves the brain's theory of mind network, and it runs for every utterance, regardless of whether the listener consciously thinks about the speaker's intentions (Tomasello, Boux & Pulvermüller, 2025). Critically, it cannot be switched off by knowing the voice is synthetic.

Neurophysiological measurements show that this intention-reading begins within approximately 100 milliseconds of a key word — essentially instantaneous by the standards of language processing (Tomasello et al., 2022).

What happens when this mechanism encounters AI-generated speech? The prosodic signal is present — and the intention-reading system activates in response, as it does for any speech. The hypothesis here is that the signal, however acoustically plausible, points to nothing: no communicative event behind it, no intentional state that generated it. This is what might be called an indexical vacuum — prosodic cues are present, but the intentional state that normally produces them is not. The mismatch, repeated at every step of listening, may itself be what costs effort.


What the evidence shows

The phenomenon has been measured, even if it hasn't been fully explained.

A 2023 doctoral study used pupillometry — continuous measurement of pupil dilation, a reliable physiological index of cognitive effort — to compare human and synthetic speech across a range of TTS systems, including state-of-the-art neural models. Human speech consistently demanded the least cognitive load. More unexpectedly, in quiet listening conditions, the highest-quality synthesizer produced the greatest load of all. The researchers noted that in these conditions, the brain appeared to be indexing "some other resource" than working memory, and left the question open for future research (Govender, 2023).

Where Govender measures the effort required to process what the voice delivers, Gong (2023) captures another dimension: depth of engagement. When listeners hear human speech, the brain engages more deeply (higher activity in regions associated with auditory processing, language comprehension, and working memory) than when they hear AI-generated speech, even when they rate the AI voice as equally comprehensible. The difference is stronger for emotionally inflected content, where the speaker's intention and attitude are most salient.

Both lines of evidence share a critical finding: the gap between human and AI speech is not about intelligibility. Listeners can transcribe synthetic speech as accurately as human speech. The difference appears somewhere between decoding the signal and constructing meaning — precisely where prosodic intention would normally do its work.

The implications for learning are specific and worth stating plainly. Cognitive load research consistently shows that when processing demands are high, recall suffers — not because the content wasn't heard, but because the resources needed for retention are consumed by the act of listening itself. Studies confirm that AI-narrated content impairs recall even when comprehension appears intact (Govender, 2023). The listener follows along, but the information doesn't hold in memory.

There is also a subtler loss, most acute in educational contexts. Research on pedagogical speech shows that children as young as five can detect teaching intent from prosody alone — before words, before context. A speaker who intends to teach produces prosodic patterns that signal this intention, and listeners orient to them (Bascandziev, Shafto & Bonawitz, 2025). Current AI narration is unlikely to produce these signals consistently, because they normally arise from genuine pedagogical intent.


Where it works — and why

At this point, a reasonable objection arises: surely some AI voices work fine? Navigation prompts don't exhaust anyone. Plenty of AI-narrated podcasts hold attention without difficulty. If the problem is fundamental, why doesn't it show up everywhere?

The answer may lie in what the task asks of the listener — and the clearest illustration comes from an unexpected source. NotebookLM's AI-generated podcasts sound remarkably natural: the emphasis lands where it should, the speakers seem genuinely engaged with the material, the prosody feels motivated. What distinguishes it is not the synthesis technology itself, but the architecture of the input.

NotebookLM doesn't use strict text-to-speech — it generates a conversational script first, a dialogue where two speakers work through the content together. The communicative structure is built into the script itself: one speaker expresses surprise, the other asks for clarification, both orient toward a listener who needs the material explained. Whatever the underlying mechanism, the result sounds motivated in a way that simple TTS typically does not. Communicative structure was built before synthesis began — and that seems to be what makes the difference.

This points toward a broader principle. Navigation prompts, procedural instructions, pronunciation guides — these don't ask the listener to track an argument or calibrate their understanding against a speaker's grasp of the material. The communicative demand is minimal and transactional; the indexical vacuum is shallow. The problem deepens with content that has discursive complexity: a concept being explained for the first time, an argument built across several steps, a narrative that requires holding a developing picture in mind. Here, the listener normally relies on prosodic signals to know where to direct attention, when something is especially important, where the difficulty lies. A human expert's prosody is, among other things, a continuously updated map of the conceptual terrain. AI narration provides a map that looks right but wasn't drawn from the territory.


The limits of better synthesis — and what might lie beyond

The natural response to all of this is that AI voices will keep improving, and the problem will eventually disappear. This assumes the problem is acoustic — and it isn't.

The field has made enormous progress. Recent work has pushed TTS toward semantic awareness: systems like Llama-VITS integrate large language model embeddings into speech synthesis, prompting the model to classify the emotion, illocutionary type, and speaking style of each sentence before synthesis (Feng & Yoshimoto, 2024). The results are measurably better, particularly for expressive content.

And yet the step stops short of what the listening brain is looking for. Classifying a sentence as a warning or a suggestion is not the same as understanding what makes this particular point significant in this particular explanation. Motivated prosody doesn't emerge from sentence-level labels — it emerges from discourse-level understanding: what is new information here, what is the conceptual crux, where does this argument turn. That understanding is what generates emphasis, pacing, and pause in human speech. Semantic embeddings, however rich, don't model it. The prosody that emerges is more expressive, but it is still generated, not motivated. The indexical vacuum narrows; it doesn't close.

What would close it points in a different direction. The industry currently optimizes for naturalness, speaker similarity, and expressiveness. The standard measure — MOS, Mean Opinion Score — asks listeners to rate how natural a voice sounds. It doesn't ask whether they remembered what it said. Recent work suggests MOS-based metrics may not even reliably predict perceptual indistinguishability from human speech, let alone comprehension or retention (Varadhan et al., 2025). But the next question is different: not how natural does the voice sound, but how motivated is the prosody. Does the emphasis land where the content actually demands it?

The nearest reframes the task, as NotebookLM already demonstrates: generate communicative structure first, synthesize second. A script where speakers have roles and relationships to the material gives the synthesis system something genuine to follow.

Further out lies what current architectures are approaching but haven't reached: discourse-level semantic analysis feeding into prosodic markup. Prosodic annotation already exists as a practice — tools and frameworks for marking pitch accents, boundaries, and stress patterns are well established in speech research. What they annotate, however, is how speech sounds, not why it sounds that way. Large language models can already identify information structure — what is given, what is new, what carries conceptual weight in a passage. A pipeline connecting that analysis to prosody generation would produce emphasis that reflects the content, not the sentence type. The missing piece is not the analysis — it already exists. It is the systematic connection between that analysis and how the text gets spoken.

Further still lies an open question: whether training on recordings of speakers who deeply understood their material might shape prosodic patterns in ways that current architectures do not.


The gap that remains

The current benchmarks for TTS quality ask: how human-like does the voice sound? The field has answered this question impressively.

One interpretation of the evidence points toward a question that hasn't yet been systematically asked: how motivated is the prosody? Does the emphasis reflect what the content actually demands — or only what sentences of this type statistically tend to sound like?

That question may not have a clean engineering answer today. But it reframes what the next generation of voice technology needs to solve. The gap between human and AI narration is deeper than the quality of the voice. What needs to be modelled is what humans intend to communicate when they speak.


References

Bascandziev, I., Shafto, P., & Bonawitz, E. (2025). Prosodic cues support inferences about the question's pedagogical intent. Open Mind: Discoveries in Cognitive Science, 9, 340–363. https://doi.org/10.1162/opmi_a_00192

Cutler, A., Dahan, D., & van Donselaar, W. (1997). Prosody in the comprehension of spoken language: A literature review. Language and Speech, 40(2), 141–201. https://pure.mpg.de/rest/items/item_68854/component/file_503903/content

Feng, X., & Yoshimoto, A. (2024). Llama-VITS: Enhancing TTS synthesis with semantic awareness. arXiv:2404.06714. https://arxiv.org/abs/2404.06714

Gong, C. (2023). AI voices reduce cognitive activity? A psychophysiological study of the media effect of AI and human newscasts in Chinese journalism. Frontiers in Psychology, 14, 1243078. https://doi.org/10.3389/fpsyg.2023.1243078

Govender, N. (2023). Evaluating synthetic speech and its impact on the end-user: Cognitive load assessment of text-to-speech systems. PhD thesis.

Hellbernd, N., & Sammler, D. (2016). Prosody conveys speaker's intentions: Acoustic cues for speech act perception. Brain and Language, 158, 1–10. https://doi.org/10.1016/j.bandl.2016.04.006

Tomasello, R., Grisoni, L., Boux, I., Sammler, D., & Pulvermüller, F. (2022). Instantaneous neural processing of communicative functions conveyed by speech prosody. Cerebral Cortex, 32(21), 4885–4901. https://doi.org/10.1093/cercor/bhab522

Tomasello, R., Boux, I., & Pulvermüller, F. (2025). Theory of Mind and the brain substrates of direct and indirect communicative action understanding. Philosophical Transactions of the Royal Society B, 380, 20230497. https://doi.org/10.1098/rstb.2023.0497

Varadhan, P. S., Thomas, S., Teja M S, S., Bhooshan, S., & Khapra, M. M. (2025). The state of TTS: A case study with human fooling rates. arXiv:2508.04179. https://arxiv.org/abs/2508.04179


Kseniia Briling, June 2026