2023/12/19 by Lingjun Zhang, Zhang, Lingjun, Xinyuan Chen +7 · 1 voice · 16 citations
Computer Science · Engineering · Mathematics · #Algorithm #Artificial intelligence #Bayesian probability #Computer Vision and Pattern Recognition (cs.CV) #Computer science #Computer vision #Constraint (computer-aided design) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Human Motion and Animation #Image (mathematics) #Image editing #Mathematics #Multimodal Machine Learning Applications #Natural language processing #Naturalness #Object (grammar) #Pattern recognition (psychology) #Prior probability #Sketch
paper · pdf · doi:10.48550/arxiv.2312.12232
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2023/12/19 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Recently, diffusion-based image generation methods are credited for their remarkable text-to-image generation capabilities, while still facing challenges in accurately generating multilingual scene text images. To tackle this problem, we propose Diff-Text, which is a training-free scene text generation framework for any language. Our model outputs a photo-realistic image given a text of any language along with a textual description of a scene. The model leverages rendered sketch images as priors, thus arousing the potential multilingual-generation ability of the pre-trained Stable Diffusion. Based on the observation from the influence of the cross-attention map on object placement in generated images, we propose a localized attention constraint into the cross-attention layer to address the unreasonable positioning problem of scene text. Additionally, we introduce contrastive image-level prompts to further refine the position of the textual region and achieve more accurate scene text generation. Experiments demonstrate that our method outperforms the existing method in both the accuracy of text recognition and the naturalness of foreground-background blending.