Back to Research papers
Research paper index

SynVAR: Synergizing Spatial and Semantic Alignment in Visual Autoregressive Model

Zhennan Chen, Tianxing Shi, Pengcheng Xu, Kepan Nan, Qian Wang, Zili Yi, Jian Yang, Ying Tai

arXiv:2608.07948Published August 8, 20260 citations
  • cs.CV

Abstract

VAR has gained widespread popularity due to its next-scale prediction paradigm. However, it faces substantial performance bottlenecks when handling complex scenes with multiple objects and attributes. Existing diffusion-based enhancement methods fail to adequately address the unique challenge of cross-scale error propagation and accumulation in VAR. To this end, we propose SynVAR, the first training-free enhancement framework specifically tailored for the VAR paradigm, which introduces a spatial-semantic collaborative control strategy to effectively suppress propagation error and improve generation quality. SynVAR comprises three key components: (1) Global guidance to ensure reasonable spatial structure in the early stages, (2) Receptive field constraints to mitigate early-stage semantic confusion, (3) High-frequency compensation to recover fine-grained details. Extensive quantitative and qualitative experiments demonstrate the significant improvements in the ability of SynVAR to enhance the VAR's capability for complex scene modeling.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.