vix.ing
·
top
·
new
·
best
·
stats
·
spec
Advancing Aesthetic Image Generation via Composition Transfer
2026/04/29 by
Kai Zou
,
Zhiwei Zhao
,
Bin Liu
+1
paper
· doi:10.1007/s11263-026-02862-8
Citations
CoT-lized Diffusion: Let's Reinforce T2I Generation Step-by-step
PosterCraft: Rethinking High-Quality Aesthetic Poster Generation in a Unified Framework
POSTA: A Go-to Framework for Customized Artistic Poster Generation
EliGen: Entity-Level Controlled Image Generation with Regional Attention
Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Draw Like an Artist: Complex Scene Generation with Diffusion Model via Composition, Painting, and Retouching
ControlNet++: Improving Conditional Controls with Efficient Consistency Feedback
MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis
Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs
Ranni: Taming Text-to-Image Diffusion for Accurate Instruction Following
Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack
IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models
SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
JourneyDB: A Benchmark for Generative Image Understanding
Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis
Grounded Text-to-Image Synthesis with Attention Refocusing
Visual Instruction Tuning
ImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generation
Adding Conditional Control to Text-to-Image Diffusion Models
GLIGEN: Open-Set Grounded Text-to-Image Generation
LAION-5B: An open large-scale dataset for training next generation image-text models
Scaling Autoregressive Models for Content-Rich Text-to-Image Generation
MANIQA: Multi-dimension Attention Network for No-Reference Image Quality Assessment
BNAI, NO-TOKEN, and MIND-UNITY: Pillars of a Systemic Revolution in Artificial Intelligence
High-Resolution Image Synthesis with Latent Diffusion Models
LoRA: Low-Rank Adaptation of Large Language Models
Learning Transferable Visual Models From Natural Language Supervision
Photo Aesthetics Ranking Network with Attributes and Content Adaptation
Microsoft COCO: Common Objects in Context
SLIC Superpixels Compared to State-of-the-Art Superpixel Methods
RealCompo: Balancing Realism and Compositionality Improves Text-to-Image Diffusion Models