CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Latency and turn-taking in Spoken Dialogue Systems: Evaluating Strategies for Natural Real-time Conversations with AI
Jönköping University, School of Engineering, JTH, Department of Computer Science and Informatics.
2026 (English)Independent thesis Basic level (university diploma), 10 credits / 15 HE creditsStudent thesis
Abstract [en]

Spoken Dialogue System (SDS) utilize on a sequential processing pipelines that introduce compounding latency across Automatic Speech Recognition (ASR), Large Language Model (LLM), and Text-to-Speech (TTS) components, making natural real-time voice interactions difficult to achieve. This study evaluates how the choice of a LLM and turn-taking strategy affect the SDS latency when it is optimized for cold calling purposes. Together, four LLMs were tested under controlled conditions, which were GPT-4o-mini, GPT-5-Nano, Gemini-2.5-Flash-Lite, and GPT-5. Furthermore, three turn-taking strategies were also tested, specifically: Conservative, Rule-based, and Neural network. The results show that GPT-4o-mini produced the lowest and most consistent latency with a mean of 497.7 ms, while GPT-5 reached a mean of 1673.1 ms despite being a newer model and showing relatively consistent response times. For turn-taking, the Neural network strategy achieved the lowest mean Turn Latency at 2087 ms and notably reduced LLM Latency to 336 ms, partly attributed to AssemblyAI’s speaker classification output providing a cleaner conversational context to the language model. No component configuration successfully managed to bring down the end-to-end latency within the 200 ms threshold humans associate with natural conversation. This confirms that sequential pipeline architecture remains the primary structural bottleneck regardless of the performed component level optimization within the pipeline.

Place, publisher, year, edition, pages
2026. , p. 51
Keywords [en]
Spoken Dialogue Systems, Latency, Large Language Model, Turn-taking, Automatic Speech Recognition, Cold Calling, Real-time Voice AI, Sequential Pipeline, End-of-Turn Detection, System Responsiveness
National Category
Computer Sciences Artificial Intelligence Software Engineering
Identifiers
URN: urn:nbn:se:hj:diva-73132OAI: oai:DiVA.org:hj-73132DiVA, id: diva2:2080868
External cooperation
Uplinx AB
Subject / course
JTH, Computer Engineering
Supervisors
Examiners
Available from: 2026-08-13 Created: 2026-06-28 Last updated: 2026-08-13Bibliographically approved

Open Access in DiVA

fulltext(3210 kB)38 downloads
File information
File name FULLTEXT01.pdfFile size 3210 kBChecksum SHA-512
fdc73240494c77b687e7c1c2645d1bb901313afa3f31c8fdcb3e0b588cbb2dd5c3c65ba29edff576353a5472cfdd3e90dd0b823b7570cb4c50a1b465010e58f6
Type fulltextMimetype application/pdf

By organisation
JTH, Department of Computer Science and Informatics
Computer SciencesArtificial IntelligenceSoftware Engineering

Search outside of DiVA

GoogleGoogle Scholar
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 2175 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf