vix.ing · top · new · best · stats · spec

Hu, Wenze

  1. Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing
    2025/10/22 by Yusu Qian, Eli Bocek-Rivele, Qian, Yusu +14 · 5 voices · 8 citations
    Arts and Humanities · Computer Science · #Digital Humanities and Scholarship #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications #cs.CL #cs.CV #cs.LG
  2. Guiding Instruction-based Image Editing via Multimodal Large Language Models
    2023/09/29 by Tsu-Jui Fu, Fu, Tsu-Jui, Wenze Hu +9 · 3 voices · 25 citations
    Computer Science · #Advanced Image and Video Retrieval Techniques #Domain Adaptation and Few-Shot Learning #Multimodal Machine Learning Applications #cs.CV
  3. STIV: Scalable Text and Image Conditioned Video Generation
    2024/12/10 by Lin, Zongyu, Liu, Wei, Chen, Chen +13 · 8 citations
    #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimedia (cs.MM)
  4. NAR-Former: Neural Architecture Representation Learning towards Holistic Attributes Prediction
    2022/11/15 by Yi, Yun, Zhang, Haokui, Hu, Wenze +2 · 3 citations
    #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  5. ParC-Net: Position Aware Circular Convolution with Merits from ConvNets and Transformer
    2022/03/08 by Zhang, Haokui, Hu, Wenze, Wang, Xiaoyu · 2 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  6. UniVG: A Generalist Diffusion Model for Unified Image Generation and Editing
    2025/03/16 by Fu, Tsu-Jui, Qian, Yusu, Chen, Chen +3 · 5 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  7. Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models
    2024/10/03 by Zhengfeng Lai, Vasileios Saveris, Lai, Zhengfeng +21 · 3 citations
    Computer Science · #Advanced Image and Video Retrieval Techniques #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
  8. DiT-Air: Revisiting the Efficiency of Diffusion Model Architecture Design in Text to Image Generation
    2025/03/13 by Chen, Chen, Qian, Rui, Hu, Wenze +8 · 5 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  9. ALBench: A Framework for Evaluating Active Learning in Object Detection
    2022/07/27 by Zhanpeng Feng, Shiliang Zhang, Feng, Zhanpeng +13 · 1 citation
    Computer Science · #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning and Algorithms #Machine Learning and Data Classification
  10. GIE-Bench: Towards Grounded Evaluation for Text-Guided Image Editing
    2025/05/16 by Yusu Qian, Qian, Yusu, Jiasen Lü +12 · 2 citations
    Arts and Humanities · Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Digital Humanities and Scholarship #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications
  11. ParCNetV2: Oversized Kernel with Enhanced Attention
    2022/11/14 by Rui‐Hua Xu, Xu, Ruihan, Haokui Zhang +7 · 1 citation
    Computer Science · Engineering · #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Industrial Vision Systems and Defect Detection