vix.ing · top · new · best · stats · spec

Song, Guanglu

  1. Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
    2024/03/25 by Hao Shao, Shao, Hao, Shengju Qian +13 · 127 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Natural Language Processing Techniques #Semantic Web and Ontologies #Topic Modeling
  2. Phased Consistency Models
    2024/05/28 by Fu-Yun Wang, Wang, Fu-Yun, Zhaoyang Huang +21 · 1 voice · 24 citations
    #cs.LG #cs.CV
  3. DETRs with Collaborative Hybrid Assignments Training
    2022/11/22 by Zhuofan Zong, Zong, Zhuofan, Guanglu Song +3 · 32 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Machine Learning and Data Classification #Music and Audio Processing
  4. UniFormer: Unifying Convolution and Self-attention for Visual Recognition
    2022/01/24 by Kunchang Li, Li, Kunchang, Yali Wang +13 · 20 citations
    Computer Science · #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications
  5. UniFormer: Unified Transformer for Efficient Spatiotemporal Representation Learning
    2022/01/12 by Kunchang Li, Li, Kunchang, Yali Wang +11 · 13 citations
    Computer Science · #Human Pose and Action Recognition #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning
  6. RAPHAEL: Text-to-Image Generation via Large Mixture of Diffusion Paths
    2023/05/29 by Zeyue Xue, Xue, Zeyue, Guanglu Song +11 · 16 citations
    Computer Science · Social Sciences · #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications #Computational and Text Analysis Methods
  7. FouriScale: A Frequency Perspective on Training-Free High-Resolution Image Synthesis
    2024/03/19 by Huang, Linjiang, Fang, Rongyao, Zhang, Aiping +4 · 18 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  8. Gen-L-Video: Multi-Text to Long Video Generation via Temporal Co-Denoising
    2023/05/29 by Wang, Fu-Yun, Chen, Wenshuo, Song, Guanglu +3 · 14 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  9. Rethinking the Spatial Inconsistency in Classifier-Free Diffusion Guidance
    2024/04/08 by Shen, Dazhong, Song, Guanglu, Xue, Zeyue +2 · 16 citations
    #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  10. CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching
    2024/04/04 by Jiang, Dongzhi, Song, Guanglu, Wu, Xiaoshi +5 · 16 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  11. MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines
    2024/09/19 by Dongzhi Jiang, Jiang, Dongzhi, Renrui Zhang +23 · 19 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Semantic Web and Ontologies #Web Data Mining and Analysis
  12. MoVA: Adapting Mixture of Vision Experts to Multimodal Context
    2024/04/19 by Zhuofan Zong, Zong, Zhuofan, Bingqi Ma +13 · 16 citations
    Social Sciences · Engineering · #Geographic Information Systems Studies #Spatial Cognition and Navigation
  13. AnimateLCM: Computation-Efficient Personalized Style Video Generation without Personalized Video Data
    2024/02/01 by Fu-Yun Wang, Zhaoyang Huang, Wang, Fu-Yun +13 · 1 voice · 5 citations
    #cs.CV #cs.LG
  14. Deep Reward Supervisions for Tuning Text-to-Image Diffusion Models
    2024/05/01 by Xiaoshi Wu, Yiming Hao, Wu, Xiaoshi +13 · 13 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Speech Recognition and Synthesis
  15. Temporal Enhanced Training of Multi-view 3D Object Detector via Historical Object Prediction
    2023/04/03 by Zhuofan Zong, Zong, Zhuofan, Dongzhi Jiang +11 · 7 citations
    Computer Science · #Advanced Image and Video Retrieval Techniques #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Video Surveillance and Tracking Methods
  16. Be-Your-Outpainter: Mastering Video Outpainting through Input-Specific Adaptation
    2024/03/20 by Fuyun Wang, Xiaoshi Wu, Wang, Fu-Yun +13 · 7 citations
    Computer Science · Economics, Econometrics and Finance · Neuroscience · #Generative Adversarial Networks and Image Synthesis #Cinema and Media Studies #Aesthetic Perception and Analysis
  17. Exploring the Role of Large Language Models in Prompt Encoding for Diffusion Models
    2024/06/17 by Bingqi Ma, Ma, Bingqi, Zhuofan Zong +7 · 8 citations
    Computer Science · #Topic Modeling
  18. Revisiting the Sibling Head in Object Detector
    2020/03/17 by Guanglu Song, Song, Guanglu, Yu Liu +3 · 3 citations
    Computer Science · Medicine · #Advanced Neural Network Applications #COVID-19 diagnosis using AI #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences
  19. EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM
    2024/12/12 by Zhuofan Zong, Zong, Zhuofan, Dongzhi Jiang +13 · 8 citations
    Medicine · #Radiomics and Machine Learning in Medical Imaging
  20. Self-slimmed Vision Transformer
    2021/11/24 by Zhuofan Zong, Zong, Zhuofan, Kunchang Li +11 · 2 citations
    Computer Science · Engineering · #Advanced Neural Network Applications #CCD and CMOS Imaging Sensors #Visual Attention and Saliency Detection
  21. Region-based Quality Estimation Network for Large-scale Person Re-identification
    2017/11/23 by Song, Guanglu, Leng, Biao, Liu, Yu +2 · 1 citation
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  22. Decoupled DETR: Spatially Disentangling Localization and Classification for Improved End-to-End Object Detection
    2023/10/24 by Manyuan Zhang, Guanglu Song, Zhang, Manyuan +5 · 2 citations
    Computer Science · #Advanced Neural Network Applications #Domain Adaptation and Few-Shot Learning #Advanced Image and Video Retrieval Techniques
  23. Robo-MUTUAL: Robotic Multimodal Task Specification via Unimodal Learning
    2024/10/02 by Jianxiong Li, Zhihao Wang, Li, Jianxiong +18 · 3 citations
    Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Robot Manipulation and Learning #Robotics (cs.RO) #Robotics and Automated Systems #Speech and dialogue systems
  24. INTERN: A New Learning Paradigm Towards General Vision
    2021/11/16 by Jing Shao, Shao, Jing, Siyu Chen +51 · 1 citation
    Computer Science · #Domain Adaptation and Few-Shot Learning #Advanced Neural Network Applications #Multimodal Machine Learning Applications
  25. See Further When Clear: Curriculum Consistency Model
    2024/12/09 by Yunpeng Liu, Boxiao Liu, Liu, Yunpeng +11 · 2 citations
    Social Sciences · #Higher Education Learning Practices