"Uhuh, ik begrijp je!"

Stef
Laurent

Sociale robot gaat met ouderen in gesprek om eenzaamheid tegen te gaan

Een persoon praat met een robot

Sociale isolatie en eenzaamheid bij ouderen vormen een groeiend maatschappelijk probleem. Een masterstudent Burgerlijk Ingenieur van de UGent ontwikkelde een sociale robot die met ouderen in gesprek kan gaan. Het doel? Eenzaamheid tegengaan door een luisterend oor te bieden.

Vele ouderen kampen met een gebrek aan sociale interactie. Uit onderzoek van de wereld gezondheid organisatie blijkt dat eenzaamheid niet enkel de mentale gezondheid schaadt, maar ook kan zorgen voor een verhoogd risico op dementie, hart- en vaatziekten, ...
Daarom werd een sociale robot ontworpen die in staat is om met ouderen in gesprek te gaan, niet om het sociale contact te vervangen maar aanvullend bij de fysieke gesprekken in een woonzorgcentrum. Dit lijkt misschien een eenvoudige opdracht, want voor ons, de mens, verloopt dit heel natuurlijk, maar dat is het zeker niet.

De paradox van Moravec

Tijdens een menselijke interactie gebeuren er onbewust heel wat dingen tegelijk. Deze voelen heel natuurlijk aan, maar zijn enorm moeilijk om na te bootsen voor een robot. Dit wordt vaak beschreven als de paradox van Moravec. Die formuleert dat complexe vraagstukken, zoals wiskundige berekeningen, vaak eenvoudig op te lossen zijn voor computers. Daarentegen zijn heel natuurlijke en gewone vaardigheden voor de mens vaak enorm moeilijk om na te bootsen voor een computer.

Denk maar aan het juist aanvoelen wanneer je aan de beurt bent in een gesprek, het verstaan van streekgebonden dialecten of het aangeven dat je effectief luistert. Dat laatste doen mensen door gebruik te maken van kleine woordjes of geluiden, zoals ‘uhm’, ‘ja’, ‘ahzo’ …Of zelfs door kleine knikgebaren te tonen.In de literatuur noemt men dit ook wel ‘backchannels’. Veel sociale robots kunnen vaak al vlot praten met een mens, maar er zijn er weinig die het gedrag beheersen om actief te luisteren.

Toonhoogtes, stiltes en AI

Om de robot deze vaardigheid te geven, werden er twee methodes uitgetest uit de literatuur om te bepalen op welk moment de robot best backchannels zou produceren. Enerzijds werd een methode getest waar, op basis van de toonhoogte en stiltes in het gesprek, het juiste moment werd bepaald. De onderzoekers die deze methode ontwikkelde, baseerden zich op data verzameld van gesprekken tussen mensen. Anderzijds werd ook een AI-gebaseerde methode getest waarbij een AI-model detecteert hoe waarschijnlijk het is dat er op dat moment een backchannel zou moeten worden geproduceerd.

Van zodra het juiste moment is gevonden, moet er ook bepaald worden welk specifiek woord zou worden gereproduceerd. Hiervoor werd een eigen en efficiënt AI-model getraind op basis van gespreksdata tussen mensen, dat gebruik maakt van de voorgaande zin om te bepalen welke backchannel er dan effectief moet worden uitgesproken. Zo werd er een sociale robot gecreëerd die niet alleen in staat was om het juiste moment te bepalen om deze bevestigingen te geven, maar ook om te bepalen welk exact type backchannel er moest worden gebruikt.

Door het actieve luistergedrag van de robot praatten de deelnemers gemiddeld langer, wat erop zou kunnen wijzen dat ze meer betrokken waren bij het gesprek.

Meer betrokkenheid

Twintig deelnemers gingen elk ongeveer twintig minuten in gesprek met de robot om de technologie te testen. Hoewel het aantal deelnemers nogal beperkt was om harde bewijzen te geven, leverde het experiment waardevolle inzichten op.

Door het actieve luistergedrag van de robot praatten de deelnemers gemiddeld langer, wat erop zou kunnen wijzen dat ze meer betrokken waren bij het gesprek. Maar niet alle deelnemers waren even enthousiast over de backchannels. Ze hadden het niet meteen verwacht of vonden dat dit te veel aanwezig was. Dit bewijst nog maar eens dat zelfs dit heel onmerkbare gedrag van de mens, heel complex is om juist te krijgen voor een robot.

Blik op de toekomst

Ondanks de grote ontwikkelingen in AI- en spraaktechnologie, voelt een gesprek met een robot nog lang niet even vlot en natuurlijk aan. Het wisselen van beurt en het verstaan van de persoon tijdens de conversatie loopt vaak stroef. Daarnaast hebben robots vaak nog moeite met het reproduceren van backchannels.

Ondanks dat er in dit onderzoek een poging werd gedaan om het actieve luistergedrag van een robot te verbeteren, met behulp van deze backchannels, is er nog een lange weg te gaan om robots natuurlijk te doen aanvoelen tijdens een sociale interactie.

Bibliografie

Liu C., M. Su, Y. Xiang, Y. Huang, Y. Yang, K. Zhang, and M. Fan. “Toward Enabling Natural Conversation with Older Adults via the Design of LLM-Powered Voice Agents that Support Interruptions and Backchannels.” In: Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (2025), pp. 1–22. url: https://dl.acm.org/doi/10.1145/3706598.3714228.

R N.. “Mastering Turn Detection and Interruption Handling in Voice AI Applications.” In: (2025). url: https://comparevoiceai.com.

Moravec H.. “Mind Children: The Future of Robot and Human Intelligence.” In: (1988). url: https://books.google.be/books?id=56mb7XuSx3QC.

Grimm M. and K. Kroschel. “Robust Speech: Recognition and Understanding.” In: (2007). url: https://books.google.be/books?id=OuSgDwAAQBAJ.

Shi M., Y. Shu, L. Zuo, Q. Chen, S. Zhang, J. Zhang, and L.-R. Dai. “Semantic VAD: Low-Latency Voice Activity Detection for Speech Interaction.” In: (2023). url: http://arxiv.org/abs/2305.12450.

Ekstedt E. and G. Skantze. “TurnGPT: a Transformer-based Language Model for Predicting Turn-taking in Spoken Dialog.” In: Findings of the Association for Computational Linguistics: EMNLP 2020 (2020), pp. 2981–2990. url: http://arxiv.org/abs/2010.10874.

Liao B., Y. Xu, J. Ou, K. Yang, W. Jian, P. Wan, and D. Zhang. “FlexDuo: A Pluggable System for Enabling Full-Duplex Capabilities in Speech Dialogue Systems.” In: (2025). url: http://arxiv.org/abs/2502.13472.

Pinto M. J. and T. Belpaeme. “Predictive Turn-Taking: Leveraging Language Models to Anticipate Turn Transitions in Human-Robot Dialogue.” In: 2024 33rd IEEE International Conference on Robot and Human Interactive Communication (ROMAN) (2024), pp. 1733–1738. url: https://ieeexplore.ieee.org/document/10731379/.

Lin Y., Y. Zheng, M. Zeng, and W. Shi. “Predicting Turn-Taking and Backchannel in Human-Machine Conversations Using Linguistic, Acoustic, and Visual Signals.” In: (2025). url: http://arxiv.org/abs/2505.12654.

Russell S. O. and N. Harte. “Visual Cues Enhance Predictive Turn-Taking for Two-Party Human Interaction.” In: (2025). url: https://arxiv.org/abs/2505.21043.

Openai. “Voice activity detection (VAD) - OpenAI API.” In:. url: https://platform.openai.com.

Cumbal R., R. Kantharaju, M. Paetzel-Prüsmann, and J. Kennedy. “Let Me Finish First - The Effect of Interruption-Handling Strategy on the Perceived Personality of a Social Agent.” In: Proceedings of the 24th ACM International Conference on Intelligent Virtual Agents (2024), pp. 1–10. url: https://dl.acm.org/doi/10.1145/3652988.3673916.

Cao S., J. Moon, A. Mahmood, V. N. Antony, Z. Xiao, A. Liu, and C.-M. Huang. “Interruption Handling for Conversational Robots.” In: (2025). url: http://arxiv.org/abs/2501.01568.

Garg K., L. M. Mathews, and P. Kambli. “InterConv: Conversational Agent That Handles Interruptions Like Humans.” In: 2025 IEEE International Conference on Interdisciplinary Approaches in Technology and Management for Social Innovation (IATMSI) 3 (2025), pp. 1–6. url: https://ieeexplore.ieee.org/abstract/document/10985683.

“Backchannel (linguistics).” In: Wikipedia (2026). url: https://en.wikipedia.org/w/index.php?title=Backchannel_(linguistics)&oldid=1331606746#cite_note-12.

Maitra A., D. French, and K. von der Wense. “Dialogue Acts as a Lens on Human–LLM Interaction: Analyzing Conversational Norms in Model-Generated Responses.” In: Proceedings of the Fourth Workshop on Bridging Human-Computer Interaction and Natural Language Processing (HCI+NLP) (2025), pp. 317–325. url: https://aclanthology.org/2025.hcinlp-1.25/.

Inoue K., D. Lala, G. Skantze, and T. Kawahara. “Yeah, Un, Oh: Continuous and Real-time Backchannel Prediction with Fine-tuning of Voice Activity Projection.” In: (2025). url: http://arxiv.org/abs/2410.15929.

Paierl M., M. Hagmüller, and B. Schuppler. “Continuous prediction of backchannel timing for human-robot interaction.” In: Interspeech 2025 (2025), pp. 3020–3024. url: https://www.isca-archive.org/interspeech_2025/paierl25_interspeech.html.

Arnold K.. “Humming Along.” In: Contemporary Psychoanalysis 48(1) (2012), pp. 100–117. url: https://doi.org/10.1080/00107530.2012.10746491.

Yngve V. H.. “ON GETTING A WORD IN EDGEWISE.” In: The Sixth Regional Meeting [of the] Chicago Linguistic Society (1970), pp. 567–578.

Drew P. and J. Heritage. “Conversation Analysis: Turn design and action formation.” In: (2006). url: https://books.google.be/books?id=tEIcAQAAIAAJ.

Park Y.-H., W. Liermann, Y.-S. Choi, S. H. Kim, J.-U. Bang, S. Yun, and K. J. Lee. “Backchannel prediction, based on who, when and what.” In: Interspeech 2024 (2024), pp. 3570–3574. url: https://www.isca-archive.org/interspeech_2024/park24b_interspeech.html.

Poppe R., K. P. Truong, D. Reidsma, and D. Heylen. “Backchannel Strategies for Artificial Listeners.” In: Intelligent Virtual Agents 6356 (2010), pp. 146–158. url: http://link.springer.com/10.1007/978-3-642-15892-6_16.

Ward N. and W. Tsukahara. “Prosodic features which cue back-channel responses in English and Japanese.” In: Journal of Pragmatics 32(8) (2000), pp. 1177-1207. url: https://www.sciencedirect.com/science/article/pii/S0378216699001095.

Blomsma P., J. Vaitonyté, G. Skantze, and M. Swerts. “Backchannel behavior is idiosyncratic.” In: Language and Cognition 16(4) (2024), pp. 1158–1181. url: https://www.cambridge.org/core/journals/language-and-cognition/article/backchannel-behavior-is-idiosyncratic/F75D0AEEAF258399166A58E3DDCC7D7E.

Choi Y.-S., J.-U. Bang, and S. H. Kim. “Joint streaming model for backchannel prediction and automatic speech recognition.” In: ETRI Journal 46(1) (2024), pp. 118–126. url: https://onlinelibrary.wiley.com/doi/abs/10.4218/etrij.2023-0358.

Ekstedt E. and G. Skantze. “Voice Activity Projection: Self-supervised Learning of Turn-taking Events.” In: (2022). url: https://arxiv.org/abs/2205.09812.

Adiba A. I., T. Homma, and T. Miyoshi. “Towards Immediate Backchannel Generation Using Attention-Based Early Prediction Model.” In: ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (2021), pp. 7408–7412. url: https://ieeexplore.ieee.org/document/9414193.

Wang J., L. Chen, A. Khare, A. Raju, P. Dheram, D. He, M. Wu, A. Stolcke, and V. Ravichandran. “Turn-taking and Backchannel Prediction with Acoustic and Large Language Model Fusion.” In: (2024). url: http://arxiv.org/abs/2401.14717.

Benus S., A. Gravano, and J. Hirschberg. “THE PROSODY OF BACKCHANNELS IN AMERICAN ENGLISH.” In: (2007).

Ward N.. “Backchannel Facts.” In: Backchannel Facts (2017). url: https://www.cs.utep.edu/nigel/bc/.

Zeng A., Z. Du, M. Liu, K. Wang, S. Jiang, L. Zhao, Y. Dong, and J. Tang. “GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.” In: arXiv.org (2024). url: https://arxiv.org/abs/2412.02612v1.

Mai L. and J. Carson-Berndsen. “Real-Time Textless Dialogue Generation.” In: (2025). url: http://arxiv.org/abs/2501.04877.

Zhang D., S. Li, X. Zhang, J. Zhan, P. Wang, Y. Zhou, and X. Qiu. “SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.” In: (2023). url: http://arxiv.org/abs/2305.11000.

Chen J., Y. Hu, J. Li, K. Li, K. Liu, W. Li, X. Li, Z. Li, F. Shen, X. Tang, M. Wei, Y. Wu, F. Xie, K. Xu, and K. Xie. “FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations.” In: (2025). url: http://arxiv.org/abs/2509.06502.

Zhang H., W. Li, R. Chen, V. Kothapally, M. Yu, and D. Yu. “LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems.” In: (2025). url: http://arxiv.org/abs/2502.14145.

Vu T. H. M. and T. B. N. Tran. “Backchannel Across Cultures: A Review and Implications.” In: VietTESOL International Convention Proceedings 1 (2021). url: https://proceedings.viettesol.org.vn/index.php/vic/article/view/189.

Park Y., D. Kim, and H. Song. “I Know You're Listening: Designing Visual Backchannels for Voice User Interfaces.” In: Companion Proceedings of the 30th International Conference on Intelligent User Interfaces (2025), pp. 56–59. url: https://dl.acm.org/doi/10.1145/3708557.3716343.

Engwall O., R. Cumbal, and A. R. Majlesi. “Socio-cultural perception of robot backchannels.” In: Frontiers in Robotics and AI 10 (2023), pp. 988042. url: https://pmc.ncbi.nlm.nih.gov/articles/PMC9909394/.

Lala D., P. Milhorat, K. Inoue, M. Ishida, K. Takanashi, and T. Kawahara. “Attentive listening system with backchanneling, response generation and flexible turn-taking.” In: Proceedings of the 18th Annual SIGdial Meeting on Discourse and Dialogue (2017), pp. 127–136. url: https://aclanthology.org/W17-5516/.

Blache P., M. Abderrahmane, S. Rauzy, and R. Bertrand. “An integrated model for predicting backchannel feedbacks.” In: Proceedings of the 20th ACM International Conference on Intelligent Virtual Agents (2020), pp. 1–3. url: https://dl.acm.org/doi/10.1145/3383652.3423948.

Roy R., J. Raiman, S.-G. Lee, T.-D. Ene, R. Kirby, S. Kim, J. Kim, and B. Catanzaro. “PERSONAPLEX: VOICE AND ROLE CONTROL FOR FULL DUPLEX CONVERSATIONAL SPEECH MODELS.” In: (2026).

A. Ł. and A. Gut. “rom robots to chatbots: unveiling the dynamics of human-AI interaction.” In: Frontiers in psychology (2025). url: https://doi.org/10.3389/fpsyg.2025.1569277.

Inoue K., M. Elmers, Y. Fu, Z. H. Pang, D. Lala, K. Ochi, and T. Kawahara. “Prompt-Guided Turn-Taking Prediction.” In: (2025). url: https://arxiv.org/abs/2506.21191.

Reece A., G. Cooney, P. Bull, C. Chung, B. Dawson, C. Fitzpatrick, T. Glazer, D. Knox, A. Liebscher, and S. Marin. “The CANDOR corpus: Insights from a large multimodal dataset of naturalistic conversation.” In: Science Advances 9(13) (2023), pp. eadf3197. url: https://www.science.org/doi/10.1126/sciadv.adf3197.

Fukunaga Y., R. Nishimura, K. Ohta, and N. Kitaoka. “Backchannel prediction for natural spoken dialog systems using general speaker and listener information.” In: Interspeech 2025 (2025), pp. 1078–1082. url: https://www.isca-archive.org/interspeech_2025/fukunaga25_interspeech.html.

Lin T.-H., H. Dinner, T. L. Leung, B. Mutlu, J. G. Trafton, and S. Sebo. “Connection-Coordination Rapport (CCR) Scale: A Dual-Factor Scale to Measure Human-Robot Rapport.” In: Proceedings of the 2025 ACM/IEEE International Conference on Human-Robot Interaction (2025), pp. 869–879. url: https://dl.acm.org/doi/10.5555/3721488.3721594.

“Methods for subjective determination of transmission quality.” In: (2000). url: https://cds.cern.ch/record/789314.

Doyle P. R., I. Gessinger, J. Edwards, L. Clark, O. Dumbleton, D. Garaialde, D. Rough, A. Bleakley, H. P. Branigan, and B. R. Cowan. “The Partner Modelling Questionnaire: A Validated Self-Report Measure of Perceptions toward Machines as Dialogue Partners.” In: ACM Trans. Comput.-Hum. Interact. 32(4) (2025), pp. 39:1–39:33. url: https://dl.acm.org/doi/10.1145/3729170.

Saad L., E. Roesler, E. Phillips, and J. G. Trafton. “Choosing the “Perfect” Scale: A Primer to Evaluate Existing Scales in HRI.” In: J. Hum.-Robot Interact. 15(2) (2026), pp. 42:1–42:30. url: https://dl.acm.org/doi/10.1145/3772066.

Inoue K., D. Lala, K. Ochi, T. Kawahara, and G. Skantze. “Towards Objective Evaluation of Socially-Situated Conversational Robots: Assessing Human-Likeness through Multimodal User Behaviors.” In: International Cconference on Multimodal Interaction (2023), pp. 86–90. url: http://dx.doi.org/10.1145/3610661.3617151.

Qian L. and G. Skantze. “Aligning Backchannel and Dialogue Context Representations via Contrastive LLM Fine-Tuning.” In: (2026). url: http://arxiv.org/abs/2604.16622.

June 2025. url: https://www.who.int/publications/i/item/978240112360

Pinto-Bernal M. J., K. Boels, and T. Belpaeme. “Fostering Social Connection and Emotional Engagement to Mitigate Social Isolation in Care Homes Using Empathetic Human-Robot Conversation.” In:.

Axelsson A., H. Buschmeier, and G. Skantze. “Modeling Feedback in Interaction With Conversational Agents—A Review.” In: Frontiers in Computer Science Volume 4 - 2022 (2022). url: https://www.frontiersin.org/journals/computer-science/articles/10.3389/fcomp.2022.744574.

Veluri B., B. N. Peloquin, B. Yu, H. Gong, and S. Gollakota. “Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents.” In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (2024), pp. 21390–21402. url: https://aclanthology.org/2024.emnlp-main.1192.

Download scriptie (9.13 MB)
Universiteit of Hogeschool
Universiteit Gent
Thesis jaar
2026
Promotor(en) en begeleiders
Tony Belpaeme, Maria Jose Pinto Bernal, Eva Verhelst
Kernwoorden