- Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
2026/07/21 by Xinjie Zhang, Peng Zhang, Shicheng Zheng +21 · 1 voice · 1 citation
#cs.CV #cs.AI #cs.LG #cs.MM #eess.IV
- Semantic Context Matters: Analysis of Color Names Across Domains
2026/07/19 by Adilet Yerkin, Elnara Kadyrgali, Malika Ziyada +5 · 1 voice
#cs.CV #cs.HC #cs.MM
- Scalable Visual Pretraining for Language Intelligence
2026/07/10 by Yiming Zhang, Zhonghan Zhao, Wenwei Zhang +14 · 1 voice
#cs.CV #cs.AI #cs.MM
- Paper2Video: Automatic Video Generation from Scientific Papers
2025/10/06 by Zeyu Zhu, Zhu, Zeyu, Kevin Qinghong Lin +3 · 8 voices · 6 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · #Biomedical Text Mining and Ontologies #Mathematics, Computing, and Information Processing #Video Analysis and Summarization #cs.AI #cs.CL #cs.CV #cs.MA #cs.MM
- FakeParts: a New Family of AI-Generated DeepFakes
2025/08/28 by Ziyi Liu, Liu, Ziyi, Firas Gabetni +13 · 2 voices · 1 citation
Social Sciences · #Misinformation and Its Impacts #cs.AI #cs.CV #cs.MM
- The Hidden Cost of an Image: Quantifying the Energy Consumption of AI Image Generation
2025/06/20 by Giulia Bertazzini, Chiara Albisani, Bertazzini, Giulia +7 · 2 voices · 2 citations
#cs.LG #cs.MM
- The JPEG XL Image Coding System: History, Features, Coding Tools, Design Rationale, and Future
2025/06/06 by Jon Sneyers, Jyrki Alakuijala, Sneyers, Jon +21 · 7 voices · 5 citations
#cs.MM
- SciCom Wiki: Fact-Checking and FAIR Knowledge Distribution for Scientific Videos and Podcasts
2025/05/12 by Tim Wittenborg, Constantin Sebastian Tremel, Wittenborg, Tim +9 · 1 voice
Computer Science · #Computation and Language (cs.CL) #Digital Libraries (cs.DL) #FOS: Computer and information sciences #Multimedia (cs.MM) #cs.CL #cs.DL #cs.MM
- ZJUKLAB at SemEval-2025 Task 4: Unlearning via Model Merging
2025/03/27 by Haoming Xu, Shuxun Wang, Xu, Haoming +15 · 1 voice · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimedia (cs.MM) #cs.AI #cs.CL #cs.CV #cs.LG #cs.MM
- Visual and Auditory Aesthetic Preferences Across Cultures
2025/02/20 by Harin Lee, Lee, Harin, Eline Van Geert +12 · 1 voice · 1 citation
Computer Science · Psychology · #Color perception and design #FOS: Computer and information sciences #Multimedia (cs.MM) #Multisensory perception and integration #cs.MM
- Music for All: Representational Bias and Cross-Cultural Adaptability of Music Generation Models
2025/02/11 by Atharva Mehta, Mehta, Atharva, Shivam Chauhan +9 · 1 voice · 5 citations
#cs.SD #cs.AI #cs.CL #cs.LG #cs.MM
- When End-to-End is Overkill: Rethinking Cascaded Speech-to-Text Translation
2025/02/01 by Anna Min, Chenxu Hu, Min, Anna +5 · 2 voices · 1 citation
Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Sound (cs.SD) #cs.AI #cs.CL #cs.MM #cs.SD #eess.AS #electronic engineering #information engineering
- Frechet Music Distance: A Metric For Generative Symbolic Music Evaluation
2024/12/10 by Jan Retkowski, Jakub Stępniak, Retkowski, Jan +3 · 1 voice · 5 citations
#cs.SD #cs.AI #cs.MM #eess.AS
- Video-Guided Foley Sound Generation with Multimodal Controls
2024/11/26 by Ziyang Chen, Prem Seetharaman, Chen, Ziyang +12 · 4 voices · 17 citations
Computer Science · #Speech and Audio Processing #cs.CV #cs.MM #cs.SD #eess.AS
- The Sound of Water: Inferring Physical Properties from Pouring Liquids
2024/11/18 by Piyush Bagad, Bagad, Piyush, Makarand Tapaswi +5 · 1 voice · 2 citations
#cs.CV #cs.MM #cs.SD #eess.AS
- Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction
2024/10/28 by Qintong Zhang, Zhang, Qintong, Bin Wang +13 · 1 voice · 13 citations
Computer Science · #cs.MM #cs.AI #cs.CL #cs.CV
- Beyond Coarse-Grained Matching in Video-Text Retrieval
2024/10/16 by Aozhu Chen, Chen, Aozhu, Hazel Doughty +5 · 1 voice · 1 citation
Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimedia (cs.MM) #cs.CL #cs.CV #cs.MM
- PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music Processing
2024/09/17 by P. E. Long, Phillip Long, Zachary Novack +2 · 1 voice · 5 citations
Computer Science · Neuroscience · #Music Technology and Sound Studies #Music and Audio Processing #Neuroscience and Music Perception #cs.AI #cs.LG #cs.MM #cs.SD #eess.AS
- Sequential Contrastive Audio-Visual Learning
2024/07/08 by Ioannis Tsiamas, Tsiamas, Ioannis, Santiago Pascual +5 · 1 voice · 1 citation
#cs.SD #cs.CV #cs.LG #cs.MM #eess.AS
- Unsupervised Multimodal Clustering for Semantics Discovery in Multimodal Utterances
2024/05/21 by Hanlei Zhang, Zhang, Hanlei, Hua Xu +7 · 2 citations
#cs.MM #cs.AI #cs.CL
- Evaluating Text-to-Visual Generation with Image-to-Text Generation
2024/04/01 by Zhiqiu Lin, Deepak Pathak, Lin, Zhiqiu +13 · 1 voice · 114 citations
Arts and Humanities · Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Digital Humanities and Scholarship #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimedia (cs.MM) #cs.AI #cs.CL #cs.CV #cs.LG #cs.MM
- Vlogger: Make Your Dream A Vlog
2024/01/17 by Shaobin Zhuang, Zhuang, Shaobin, Kunchang Li +11 · 1 voice · 14 citations
Computer Science · #Human Pose and Action Recognition #Multimodal Machine Learning Applications #Video Analysis and Summarization #cs.AI #cs.CV #cs.LG #cs.MM
- Re:Draw -- Context Aware Translation as a Controllable Method for Artistic Production
2024/01/07 by Joao Liborio Cardoso, Francesco Banterle, Paolo Cignoni +1 · 1 voice
#cs.CV #cs.AI #cs.GR #cs.MM
- MotionCtrl: A Unified and Flexible Motion Controller for Video Generation
2023/12/06 by Zhouxia Wang, Ziyang Yuan, Wang, Zhouxia +11 · 1 voice · 102 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimedia (cs.MM) #cs.AI #cs.CV #cs.LG #cs.MM
- X-Adapter: Adding Universal Compatibility of Plugins for Upgraded Diffusion Model
2023/12/04 by Lingmin Ran, Xiaodong Cun, Ran, Lingmin +13 · 1 voice · 1 citation
#cs.CV #cs.AI #cs.MM
- LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents
2023/11/09 by Shilong Liu, Liu, Shilong, Hao Cheng +23 · 1 voice · 39 citations
#cs.CV #cs.AI #cs.CL #cs.LG #cs.MM
- From Capture to Display: A Survey on Volumetric Video
2023/09/11 by Yili Jin, Kaiyuan Hu, Jin, Yili +7 · 1 voice · 2 citations
Computer Science · Engineering · #FOS: Computer and information sciences #FOS: Electrical engineering #Image and Video Processing (eess.IV) #Multimedia (cs.MM) #Networking and Internet Architecture (cs.NI) #cs.MM #cs.NI #eess.IV #electronic engineering #information engineering
- Terrain Diffusion Network: Climatic-Aware Terrain Generation with Geological Sketch Guidance
2023/08/31 by Zexin Hu, Kun Hu, Hu, Zexin +7 · 1 voice · 1 citation
#cs.CV #cs.AI #cs.MM
- ImageBind: One Embedding Space To Bind Them All
2023/05/09 by Rohit Girdhar, Girdhar, Rohit, Alaaeldin El-Nouby +11 · 2 voices · 183 citations
Computer Science · #cs.CV #cs.AI #cs.LG #cs.MM
- Steps towards prompt-based creation of virtual worlds
2022/11/10 by Jasmine Roberts, Andrzej Banburski-Fahey, Roberts, Jasmine +3 · 3 voices
#cs.HC #cs.AI #cs.LG #cs.MM
more