2025/06/10 by Kazuki Kawamura, Kawamura, Kazuki, Jun Rekimoto +1
Computer Science · Engineering · Neuroscience · #68T05 #Aesthetic Perception and Analysis #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #H.5.2 #Human Motion and Animation #Human-Computer Interaction (cs.HC) #I.2.7 #K.3 #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.2506.08443
openalex publication_date 2025/06/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
While current AI illustration tools can generate high-quality images from text prompts, they rarely reveal the step-by-step procedure that human artists follow. We present SakugaFlow, a four-stage pipeline that pairs diffusion-based image generation with a large-language-model tutor. At each stage, novices receive real-time feedback on anatomy, perspective, and composition, revise any step non-linearly, and branch alternative versions. By exposing intermediate outputs and embedding pedagogical dialogue, SakugaFlow turns a black-box generator into a scaffolded learning environment that supports both creative exploration and skills acquisition.