vix.ing · top · new · best · stats · spec

Yu, Zhuoran

  1. Diversify, Don't Fine-Tune: Scaling Up Visual Recognition Training with Synthetic Images
    2023/12/04 by Zhuoran Yu, Chenchen Zhu, Yu, Zhuoran +9 · 2 citations
    Computer Science · #Domain Adaptation and Few-Shot Learning #Multimodal Machine Learning Applications #Generative Adversarial Networks and Image Synthesis
  2. Scale-Equalizing Pyramid Convolution for Object Detection
    2020/05/06 by Xinjiang Wang, Wang, Xinjiang, Shilong Zhang +7 · 1 citation
    Computer Science · Engineering · #Advanced Image and Video Retrieval Techniques #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Remote-Sensing Image Classification
  3. InPL: Pseudo-labeling the Inliers First for Imbalanced Semi-supervised Learning
    2023/03/13 by Yu, Zhuoran, Li, Yin, Lee, Yong Jae · 1 citation
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  4. LASER: Lip Landmark Assisted Speaker Detection for Robustness
    2025/01/21 by Nguyen, Le Thien Phuc, Yu, Zhuoran, Lee, Yong Jae · 3 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  5. See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models
    2025/12/01 by Le T. Nguyen, Nguyen, Le Thien Phuc, Zhuoran Yu +16 · 2 citations
    Computer Science · #Multimodal Machine Learning Applications #Speech and Audio Processing #Generative Adversarial Networks and Image Synthesis
  6. CuRe: Cultural Gaps in the Long Tail of Text-to-Image Systems
    2025/06/09 by Rege, Aniket, Nie, Zinnia, Ramesh, Mahesh +5 · 2 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  7. Revisiting Active Speaker Detection: An In-the-Wild Benchmark for Generalization and Robustness
    2025/05/28 by Le T. Nguyen, Nguyen, Le Thien Phuc, Zhuoran Yu +16 · 2 citations
    Computer Science · #Speech and Audio Processing #Speech Recognition and Synthesis #Face recognition and analysis
  8. How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding
    2025/08/27 by Zhuoran Yu, Yu, Zhuoran, Yong Jae Lee +1 · 1 citation
    Computer Science · #Semantic Web and Ontologies #Topic Modeling