2025/12/19 by Hayeon Jeong, Jong-Seok Lee, Jeong, Hayeon +1
Computer Science · Neuroscience · #Aesthetic Perception and Analysis #Diffusion #Face Recognition and Perception #Generative Adversarial Networks and Image Synthesis #Generative grammar #Key (lock) #Phenomenon #Stability (learning theory) #Term (time) #cs.AI #cs.LG
paper · pdf · open access · doi:10.48550/arxiv.2512.20666
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/12/19 · openalex created_date 2025/12/26 · openalex updated_date 2026/07/28
Text-to-image diffusion models have attracted significant attention for their ability to generate diverse, high-fidelity images. However, in multi-concept generation, one concept token often dominates the output while others are suppressed-a phenomenon we term the Dominant-vs-Dominated (DvD) imbalance. To systematically study this failure mode, we introduce DominanceBench and examine its underlying causes from both data and internal-mechanistic perspectives. Our controlled fine-tuning study, which mimics concept learning during diffusion-model training, shows that concepts learned from visually homogeneous (low-variation) concept-specific training images exhibit stronger dominance when composed with others. Cross-attention analysis indicates that dominant tokens concentrate attention in early denoising steps, followed by reduced representation of competing concepts. Head-ablation analysis further shows that this dominance is distributed across attention heads rather than localized. Overall, these findings characterize DvD as a systematic concept-level failure mode and provide a basis for more reliable and controllable multi-concept generation. DominanceBench will be released upon publication.