Yu, Zhuoran
- Diversify, Don't Fine-Tune: Scaling Up Visual Recognition Training with Synthetic Images
2023/12/04 by Zhuoran Yu, Chenchen Zhu, Yu, Zhuoran +9 · 2 citations
Computer Science · #Domain Adaptation and Few-Shot Learning #Multimodal Machine Learning Applications #Generative Adversarial Networks and Image Synthesis
- Scale-Equalizing Pyramid Convolution for Object Detection
2020/05/06 by Xinjiang Wang, Wang, Xinjiang, Shilong Zhang +7 · 1 citation
Computer Science · Engineering · #Advanced Image and Video Retrieval Techniques #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Remote-Sensing Image Classification
- InPL: Pseudo-labeling the Inliers First for Imbalanced Semi-supervised Learning
2023/03/13 by Yu, Zhuoran, Li, Yin, Lee, Yong Jae · 1 citation
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- LASER: Lip Landmark Assisted Speaker Detection for Robustness
2025/01/21 by Nguyen, Le Thien Phuc, Yu, Zhuoran, Lee, Yong Jae · 3 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models
2025/12/01 by Le T. Nguyen, Nguyen, Le Thien Phuc, Zhuoran Yu +16 · 2 citations
Computer Science · #Multimodal Machine Learning Applications #Speech and Audio Processing #Generative Adversarial Networks and Image Synthesis
- CuRe: Cultural Gaps in the Long Tail of Text-to-Image Systems
2025/06/09 by Rege, Aniket, Nie, Zinnia, Ramesh, Mahesh +5 · 2 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- Revisiting Active Speaker Detection: An In-the-Wild Benchmark for Generalization and Robustness
2025/05/28 by Le T. Nguyen, Nguyen, Le Thien Phuc, Zhuoran Yu +16 · 2 citations
Computer Science · #Speech and Audio Processing #Speech Recognition and Synthesis #Face recognition and analysis
- How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding
2025/08/27 by Zhuoran Yu, Yu, Zhuoran, Yong Jae Lee +1 · 1 citation
Computer Science · #Semantic Web and Ontologies #Topic Modeling