Back to Research papers
Research paper index

Human-JEPA: A Human-Centric Vision Model that Perceives and Anticipates

Hui Wei, Licai Sun, Guoying Zhao

arXiv:2608.21160Published August 21, 20260 citations
  • cs.CV
  • cs.LG
  • action

Abstract

Machines that understand humans should perceive the present and anticipate the future. Existing human-centric vision model are pretrained on human images, set the state of the art in static dense perception, so motion and anticipation are out of reach. Here we present Human-JEPA, a human-centric vision model trained on video by anchored forecasting: dense targets are pinned to a frozen copy of the initialization, preventing a silent collapse of dense perception, and block masks are replaced by a pure past-to-future split, avoiding a five-point action tax and a seventeen-point re-identification collapse. Under frozen probes, Human-JEPA leads the pixel-anchored specialists on pose and person re-identification at 2.7 times fewer parameters, conceding high-resolution dense parsing, and its released predictor head is the first that does not degrade anticipation. A single safely adapted model thus serves both halves of understanding humans.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.