Back to Research papers
Research paper index

Per-Stroke Temporal Control for Text-to-Motion via Action Units and Action-Detection Guidance

Euijun Jung, Youngki Lee

arXiv:2607.15717Published July 17, 20260 citations
  • cs.CV
  • cs.GR
  • action

Abstract

Text-to-motion models are competent at the action a prompt names but unreliable at when each stroke lands: four punches alternating left and right rarely return four separable strokes. We introduce typed temporal events called Action Units (AUs) that make the individual stroke -- its body track, action class, time window, and impact timing -- an explicit conditioning signal. We ground a frozen text-to-motion backbone on the AU set through a lightweight gated adapter injecting two streams (per-stroke tokens and a per-frame phase channel), and at inference close residual timing errors with a training-free classifier gradient from a frozen frame-level detector. We measure per-stroke control on StrokeBench, whose prompts specify count, ordering, track, and core-frame placement, paired with an audited stroke corpus. AU grounding markedly raises the rate of correctly placed single strokes over the strongest prior interface, at the best motion quality among text-, interval-, and frame-level baselines. The prompted core frame emerges as a further steerable axis.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.