2025/07/06 by Kyoungmin Lee, Ji‐Hun Park, Lee, Kyoungmin +11
Computer Science · #Authorship Attribution and Profiling #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Recommender Systems and Techniques #Speech Recognition and Synthesis
paper · pdf · doi:10.48550/arxiv.2507.04482
openalex publication_date 2025/07/06 · openalex created_date 2025/10/20 · openalex updated_date 2026/08/03
We present a training-free framework for style-personalized image generation that operates during inference using a scale-wise autoregressive model. Our method generates a stylized image guided by a single reference style while preserving semantic consistency and mitigating content leakage. Through a detailed step-wise analysis of the generation process, we identify a pivotal step where the dominant singular values of the internal feature encode style-related components. Building upon this insight, we introduce two lightweight control modules: Principal Feature Blending, which enables precise modulation of style through SVD-based feature reconstruction, and Structural Attention Correction, which stabilizes structural consistency by leveraging content-guided attention correction across fine stages. Without any additional training, extensive experiments demonstrate that our method achieves competitive style fidelity and prompt fidelity compared to fine-tuned baselines, while offering faster inference and greater deployment flexibility.