Testing chatbots on the creation of encoders for audio conditioned image generation
2025/09/09 by Jorge Esquiche León, León, Jorge E., Miguel Carrasco +1
Computer Science · #Audio and Speech Processing (eess.AS) #Augmented Reality Applications #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2509.09717
openalex publication_date 2025/09/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
On one hand, recent advances in chatbots has led to a rising popularity in using these models for coding tasks. On the other hand, modern generative image models primarily rely on text encoders to translate semantic concepts into visual representations, even when there is clear evidence that audio can be employed as input as well. Given the previous, in this work, we explore whether state-of-the-art conversational agents can design effective audio encoders to replace the CLIP text encoder from Stable Diffusion 1.5, enabling image synthesis directly from sound. We prompted five publicly available chatbots to propose neural architectures to work as these audio encoders, with a set of well-explained shared conditions. Each valid suggested encoder was trained on over two million context related audio-image-text observations, and evaluated on held-out validation and test sets using various metrics, together with a qualitative analysis of their generated images. Although almost all chatbots generated valid model designs, none achieved satisfactory results, indicating that their audio embeddings failed to align reliably with those of the original text encoder. Among the proposals, the Gemini audio encoder showed the best quantitative metrics, while the Grok audio encoder produced more coherent images (particularly, when paired with the text encoder). Our findings reveal a shared architectural bias across chatbots and underscore the remaining coding gap that needs to be bridged in future versions of these models. We also created a public demo so everyone could study and try out these audio encoders. Finally, we propose research questions that should be tackled in the future, and encourage other researchers to perform more focused and highly specialized tasks like this one, so the respective chatbots cannot make use of well-known solutions and their creativity/reasoning is fully tested.
Citations
- Effectively obtaining acoustic, visual and textual data from videos
- Web-Bench: A LLM Code Benchmark Based on Web Standards and Frameworks
- Large Language Models for Code Generation: A Comprehensive Survey of Challenges, Techniques, Evaluation, and Applications
- Line Goes Up? Inherent Limitations of Benchmarks for Evaluating Large Language Models
- DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
- AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models
- The Llama 3 Herd of Models
- A Survey on Large Language Models for Code Generation
- Are Models Biased on Text without Gender-related Language?
- Look, Listen, and Answer: Overcoming Biases for Audio-Visual Question Answering
- Scaling Rectified Flow Transformers for High-Resolution Image Synthesis
- Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models
- Towards audio language modeling -- an overview
- Cacophony: An Improved Contrastive Audio-Text Model
- A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications
- Audiobox: Unified Audio Generation with Natural Language Prompts
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- CoDi-2: In-Context, Interleaved, and Interactive Any-to-Any Generation
- Mustango: Toward Controllable Text-to-Music Generation
- Pre-trained Speech Processing Models Contain Human-Like Biases that Propagate to Speech Emotion Recognition
- Intuitive Multilingual Audio-Visual Speech Recognition with a Single-Trained Model
- The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)
- Benchmarking Cognitive Biases in Large Language Models as Evaluators
- NExT-GPT: Any-to-Any Multimodal LLM
- RenAIssance: A Survey into AI Text-to-Image Generation in the Era of Large Model
- SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
- Word-Level Explanations for Analyzing Bias in Text-to-Image Models
- AudioToken: Adaptation of Text-Conditioned Diffusion Models for Audio-to-Image Generation
- Any-to-Any Generation via Composable Diffusion
- ImageBind: One Embedding Space To Bind Them All
- From Association to Generation: Text-only Captioning by Unsupervised Cross-modal Mapping
- BiCro: Noisy Correspondence Rectification for Multi-modality Data via Bi-directional Cross-modal Similarity Consistency
- Audio-Text Models Do Not Yet Leverage Natural Language
- Text-to-image Diffusion Models in Generative AI: A Survey
- BLAT: Bootstrapping Language-Audio Pre-training based on AudioSet Tag-guided Synthetic Data
- LLaMA: Open and Efficient Foundation Language Models
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models
- AudioLDM: Text-to-Audio Generation with Latent Diffusion Models
- Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
- Noise-aware Learning from Web-crawled Image-Text Data for Image Captioning
- NLIP: Noise-robust Language-Image Pre-training
- Robust Speech Recognition via Large-Scale Weak Supervision
- TimbreCLIP: Connecting Timbre to Text and Images
- I Hear Your True Colors: Image Guided Audio Generation
- AudioGen: Textually Guided Audio Generation
- A Survey on Generative Diffusion Model
- A Survey on Generative Diffusion Models
- Multimodal Learning with Transformers: A Survey
- Towards Understanding Grokking: An Effective Theory of Representation Learning
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
- Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets
- Multimodal Image Synthesis and Editing: The Generative AI Era
- Comparison and Analysis of Image-to-Image Generative Adversarial Networks: A Survey
- High-Resolution Image Synthesis with Latent Diffusion Models
- High-Resolution Image Synthesis with Latent Diffusion Models
- Wav2CLIP: Learning Robust Audio Representations From CLIP
- A Survey on Audio Synthesis and Audio-Visual Multimodal Processing
- AudioCLIP: Extending CLIP to Image, Text and Audio
- Large Pre-trained Language Models Contain Human-like Biases of What is Right and Wrong to Do
- Learning Transferable Visual Models From Natural Language Supervision
- Zero-Shot Text-to-Image Generation
- Image-to-Image Translation: Methods and Applications
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- Jukebox: A Generative Model for Music
- Deep Audio-Visual Learning: A Survey
- On-Device Machine Learning: An Algorithms and Learning Theory Perspective
- MirrorGAN: Learning Text-to-image Generation by Redescription
- Don't Just Assume; Look and Answer: Overcoming Priors for Visual Question Answering
- Attention Is All You Need
- Deep Residual Learning for Image Recognition
- U-Net: Convolutional Networks for Biomedical Image Segmentation
Related