Back to Research papers
Research paper index

RGB-D Video Generation for Improving Human-to-Robot Object Handover Prediction

Tianyu Sun, Zhoujie Fu, Zihui Gao, Bang Zhang, Guosheng Lin

arXiv:2608.13028Published August 13, 2026Updated September 1, 20260 citations
  • cs.CV
  • cs.RO
  • robotic
  • sim-to-real
  • robot

Abstract

Human-to-robot (H2R) object handover is a fundamental capability for human-robot collaboration, yet progress is hindered by the scarcity of large-scale, human-centric datasets and the significant sim-to-real gap. To address these challenges, we introduce Hand2Bot, an RGB-D video dataset that provides rich contextual information such as body posture and facial expressions, specifically collected for handover scenarios with real-world noise patterns. We further propose PassGen, a generative pipeline that leverages stable video diffusion and an Intention-Aware Temporal Face Encoder to synthesize realistic handover sequences while ensuring hand-object consistency. To bridge the sim-to-real gap, we implement a morphology-based depth editing strategy that replicates realistic sensor noise found in physical depth maps. Experimental evaluations demonstrate that our framework achieves high intention identification accuracy and low false trigger rates in both ablation studies and real-world deployment on a physical robot platform. Our results confirm that training on PassGen allows for robust zero-shot transfer and earlier intention anticipation compared to traditional hand-centric baselines, effectively enabling socially aware robotic behavior in shared workspaces.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.