vix.ing · top · new · best · stats · spec
  1. Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
    2026/07/21 by Xinjie Zhang, Peng Zhang, Shicheng Zheng +21 · 1 voice · 1 citation
    #cs.CV #cs.AI #cs.LG #cs.MM #eess.IV
  2. Semantic Context Matters: Analysis of Color Names Across Domains
    2026/07/19 by Adilet Yerkin, Elnara Kadyrgali, Malika Ziyada +5 · 1 voice
    #cs.CV #cs.HC #cs.MM
  3. Scalable Visual Pretraining for Language Intelligence
    2026/07/10 by Yiming Zhang, Zhonghan Zhao, Wenwei Zhang +14 · 1 voice
    #cs.CV #cs.AI #cs.MM
  4. Paper2Video: Automatic Video Generation from Scientific Papers
    2025/10/06 by Zeyu Zhu, Zhu, Zeyu, Kevin Qinghong Lin +3 · 8 voices · 6 citations
    Biochemistry, Genetics and Molecular Biology · Computer Science · #Biomedical Text Mining and Ontologies #Mathematics, Computing, and Information Processing #Video Analysis and Summarization #cs.AI #cs.CL #cs.CV #cs.MA #cs.MM
  5. FakeParts: a New Family of AI-Generated DeepFakes
    2025/08/28 by Ziyi Liu, Liu, Ziyi, Firas Gabetni +13 · 2 voices · 1 citation
    Social Sciences · #Misinformation and Its Impacts #cs.AI #cs.CV #cs.MM
  6. The Hidden Cost of an Image: Quantifying the Energy Consumption of AI Image Generation
    2025/06/20 by Giulia Bertazzini, Chiara Albisani, Bertazzini, Giulia +7 · 2 voices · 2 citations
    #cs.LG #cs.MM
  7. The JPEG XL Image Coding System: History, Features, Coding Tools, Design Rationale, and Future
    2025/06/06 by Jon Sneyers, Jyrki Alakuijala, Sneyers, Jon +21 · 7 voices · 5 citations
    #cs.MM
  8. SciCom Wiki: Fact-Checking and FAIR Knowledge Distribution for Scientific Videos and Podcasts
    2025/05/12 by Tim Wittenborg, Constantin Sebastian Tremel, Wittenborg, Tim +9 · 1 voice
    Computer Science · #Computation and Language (cs.CL) #Digital Libraries (cs.DL) #FOS: Computer and information sciences #Multimedia (cs.MM) #cs.CL #cs.DL #cs.MM
  9. ZJUKLAB at SemEval-2025 Task 4: Unlearning via Model Merging
    2025/03/27 by Haoming Xu, Shuxun Wang, Xu, Haoming +15 · 1 voice · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimedia (cs.MM) #cs.AI #cs.CL #cs.CV #cs.LG #cs.MM
  10. Visual and Auditory Aesthetic Preferences Across Cultures
    2025/02/20 by Harin Lee, Lee, Harin, Eline Van Geert +12 · 1 voice · 1 citation
    Computer Science · Psychology · #Color perception and design #FOS: Computer and information sciences #Multimedia (cs.MM) #Multisensory perception and integration #cs.MM
  11. Music for All: Representational Bias and Cross-Cultural Adaptability of Music Generation Models
    2025/02/11 by Atharva Mehta, Mehta, Atharva, Shivam Chauhan +9 · 1 voice · 5 citations
    #cs.SD #cs.AI #cs.CL #cs.LG #cs.MM
  12. When End-to-End is Overkill: Rethinking Cascaded Speech-to-Text Translation
    2025/02/01 by Anna Min, Chenxu Hu, Min, Anna +5 · 2 voices · 1 citation
    Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Sound (cs.SD) #cs.AI #cs.CL #cs.MM #cs.SD #eess.AS #electronic engineering #information engineering
  13. Frechet Music Distance: A Metric For Generative Symbolic Music Evaluation
    2024/12/10 by Jan Retkowski, Jakub Stępniak, Retkowski, Jan +3 · 1 voice · 5 citations
    #cs.SD #cs.AI #cs.MM #eess.AS
  14. Video-Guided Foley Sound Generation with Multimodal Controls
    2024/11/26 by Ziyang Chen, Prem Seetharaman, Chen, Ziyang +12 · 4 voices · 17 citations
    Computer Science · #Speech and Audio Processing #cs.CV #cs.MM #cs.SD #eess.AS
  15. The Sound of Water: Inferring Physical Properties from Pouring Liquids
    2024/11/18 by Piyush Bagad, Bagad, Piyush, Makarand Tapaswi +5 · 1 voice · 2 citations
    #cs.CV #cs.MM #cs.SD #eess.AS
  16. Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction
    2024/10/28 by Qintong Zhang, Zhang, Qintong, Bin Wang +13 · 1 voice · 13 citations
    Computer Science · #cs.MM #cs.AI #cs.CL #cs.CV
  17. Beyond Coarse-Grained Matching in Video-Text Retrieval
    2024/10/16 by Aozhu Chen, Chen, Aozhu, Hazel Doughty +5 · 1 voice · 1 citation
    Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimedia (cs.MM) #cs.CL #cs.CV #cs.MM
  18. PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music Processing
    2024/09/17 by P. E. Long, Phillip Long, Zachary Novack +2 · 1 voice · 5 citations
    Computer Science · Neuroscience · #Music Technology and Sound Studies #Music and Audio Processing #Neuroscience and Music Perception #cs.AI #cs.LG #cs.MM #cs.SD #eess.AS
  19. Sequential Contrastive Audio-Visual Learning
    2024/07/08 by Ioannis Tsiamas, Tsiamas, Ioannis, Santiago Pascual +5 · 1 voice · 1 citation
    #cs.SD #cs.CV #cs.LG #cs.MM #eess.AS
  20. Unsupervised Multimodal Clustering for Semantics Discovery in Multimodal Utterances
    2024/05/21 by Hanlei Zhang, Zhang, Hanlei, Hua Xu +7 · 2 citations
    #cs.MM #cs.AI #cs.CL
  21. Evaluating Text-to-Visual Generation with Image-to-Text Generation
    2024/04/01 by Zhiqiu Lin, Deepak Pathak, Lin, Zhiqiu +13 · 1 voice · 114 citations
    Arts and Humanities · Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Digital Humanities and Scholarship #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimedia (cs.MM) #cs.AI #cs.CL #cs.CV #cs.LG #cs.MM
  22. Vlogger: Make Your Dream A Vlog
    2024/01/17 by Shaobin Zhuang, Zhuang, Shaobin, Kunchang Li +11 · 1 voice · 14 citations
    Computer Science · #Human Pose and Action Recognition #Multimodal Machine Learning Applications #Video Analysis and Summarization #cs.AI #cs.CV #cs.LG #cs.MM
  23. Re:Draw -- Context Aware Translation as a Controllable Method for Artistic Production
    2024/01/07 by Joao Liborio Cardoso, Francesco Banterle, Paolo Cignoni +1 · 1 voice
    #cs.CV #cs.AI #cs.GR #cs.MM
  24. MotionCtrl: A Unified and Flexible Motion Controller for Video Generation
    2023/12/06 by Zhouxia Wang, Ziyang Yuan, Wang, Zhouxia +11 · 1 voice · 102 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimedia (cs.MM) #cs.AI #cs.CV #cs.LG #cs.MM
  25. X-Adapter: Adding Universal Compatibility of Plugins for Upgraded Diffusion Model
    2023/12/04 by Lingmin Ran, Xiaodong Cun, Ran, Lingmin +13 · 1 voice · 1 citation
    #cs.CV #cs.AI #cs.MM
  26. LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents
    2023/11/09 by Shilong Liu, Liu, Shilong, Hao Cheng +23 · 1 voice · 39 citations
    #cs.CV #cs.AI #cs.CL #cs.LG #cs.MM
  27. From Capture to Display: A Survey on Volumetric Video
    2023/09/11 by Yili Jin, Kaiyuan Hu, Jin, Yili +7 · 1 voice · 2 citations
    Computer Science · Engineering · #FOS: Computer and information sciences #FOS: Electrical engineering #Image and Video Processing (eess.IV) #Multimedia (cs.MM) #Networking and Internet Architecture (cs.NI) #cs.MM #cs.NI #eess.IV #electronic engineering #information engineering
  28. Terrain Diffusion Network: Climatic-Aware Terrain Generation with Geological Sketch Guidance
    2023/08/31 by Zexin Hu, Kun Hu, Hu, Zexin +7 · 1 voice · 1 citation
    #cs.CV #cs.AI #cs.MM
  29. ImageBind: One Embedding Space To Bind Them All
    2023/05/09 by Rohit Girdhar, Girdhar, Rohit, Alaaeldin El-Nouby +11 · 2 voices · 183 citations
    Computer Science · #cs.CV #cs.AI #cs.LG #cs.MM
  30. Steps towards prompt-based creation of virtual worlds
    2022/11/10 by Jasmine Roberts, Andrzej Banburski-Fahey, Roberts, Jasmine +3 · 3 voices
    #cs.HC #cs.AI #cs.LG #cs.MM

more