Back to Research papers
Research paper index

AnimeAdapter: Fine-grained and Consistent Zero-shot Anime Character Generation

Yixuan Han

arXiv:2605.20237Published May 17, 20260 citations
  • cs.CV
  • vision-language

Abstract

We present a lightweight appearance adapter for Stable Diffusion that enables controllable and consistent anime character generation under diverse editing conditions. Instead of relying on large-scale vision-language models or per-subject fine-tuning, our method injects fine-grained visual features from a single reference image into the diffusion process. Based on CLIP emergent local spatialization, we develop semantic-selective local attention. To further disentangle character appearance from spatial layout, we incorporate pose-aware conditioning during adapter training. The resulting pretrained adapter remains compact, modular, and fully compatible with Stable Diffusion community workflows, while requiring no additional fine-tuning at deployment time. Furthermore, we present a high-quality anime character dataset based on curated and restructured Danbooru prompts, and evaluate our method across several practical character editing scenarios. Our code, model weights, and dataset will be publicly released upon acceptance.

Read the original paper

This page indexes public paper metadata. The manuscript remains with its original publisher and authors.