Back to Research papers
Research paper index

Accelerating Unified Multimodal Models with Core-Expansion Routing and Unified Computation Scheduling

Wengyi Zhan, Chenqian Yan, Songwei Liu, Mingbao Lin, Rongrong Ji

arXiv:2608.29291Published August 29, 20260 citations
  • cs.AI

Abstract

Unified multimodal models jointly support understanding and generation, but incur substantial redundant computation across tokens, layers, and generation timesteps. Through token-importance probing, we identify an asymmetric core-expansion structure: understanding exhibits a stable importance component, while generation largely shares this component but requires progress-dependent corrections. We therefore propose CE-Router, which uses a task-shared core scorer and progress-conditioned generation expansions, optimized through generation decomposition and cross-task core alignment. At inference, CE-Router compacts token computation and supplies a learned routing signal to Unified Computation Scheduling, which coordinates layer skipping, FFN pruning, diffusion-head cache reuse, and denoising-step early exit. Experiments on two representative UMM architectures demonstrate consistent quality--efficiency improvements across both tasks, retaining 98.03\% of dense understanding performance with a 1.93$\times$ end-to-end inference speedup.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.