Back to Research papers
Research paper index

The World According to a Social Robot -- Augmenting Human-Robot Dialogue With Vision Language Models

Thomas Sievers

arXiv:2607.16318Published July 15, 20260 citations
  • cs.HC
  • cs.RO
  • action
  • robot

Abstract

Vision Language Models (VLMs) enable robots to visually perceive their environment as well as the actions and characteristics of their conversation partner or humans in collaboration. Especially for social robots deployed in everyday settings and for uncomplicated, natural use, it is essential that the robot has an understanding of situations that is appropriate to human customs. This paper presents initial experiences with the application of a Mistral AI language model with a Pepper robot for Human-Robot Interaction (HRI) in dialogue, as well as an investigation of the effects of additional visual information on response time in different models. The results show that incorporating visual information adds context to the dialogue with only a moderate increase in response time, enabling both the robot and the human to take into account unspoken elements of the situation. Furthermore, using an LLM hosted in Europe offers a solution that complies with European data protection regulations and can therefore facilitate real-life applications more easily.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.