2025/06/05 by Jan Ackermann, Kiyohiro Nakayama, Ackermann, Jan +7
Computer Science · Engineering · #Generative Adversarial Networks and Image Synthesis #3D Shape Modeling and Analysis #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.2506.05210
Multimodal foundation models have demonstrated strong generalization, yet their ability to transfer knowledge to specialized domains such as garment generation remains underexplored. We introduce VLG, a vision-language-garment model that synthesizes garments from textual descriptions and visual imagery. Our experiments assess VLG's zero-shot generalization, investigating its ability to transfer web-scale reasoning to unseen garment styles and prompts. Preliminary results indicate promising transfer capabilities, highlighting the potential for multimodal foundation models to adapt effectively to specialized domains like fashion design.