vix.ing · top · new · best · stats · spec

Yu, Licheng

  1. The Llama 3 Herd of Models
    2024/07/31 by Grattafiori, Aaron, Dubey, Abhimanyu, Jauhri, Abhinav +556 · 2825 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  2. Modeling Context in Referring Expressions
    2016/07/31 by Yu, Licheng, Poirson, Patrick, Yang, Shan +2 · 118 citations
    #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  3. TVQA: Localized, Compositional Video Question Answering
    2018/09/05 by Jie Lei, Licheng Yu, Lei, Jie +5 · 44 citations
    Computer Science · #Advanced Image and Video Retrieval Techniques #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications
  4. Movie Gen: A Cast of Media Foundation Models
    2024/10/17 by Adam Polyak, Polyak, Adam, Amit Zohar +164 · 126 citations
    Economics, Econometrics and Finance · #Cinema and Media Studies
  5. Apollo: An Exploration of Video Understanding in Large Multimodal Models
    2024/12/13 by Orr Zohar, Xiaohan Wang, Zohar, Orr +21 · 3 voices · 23 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Semantic Web and Ontologies
  6. MAttNet: Modular Attention Network for Referring Expression Comprehension
    2018/01/24 by Yu, Licheng, Lin, Zhe, Shen, Xiaohui +4 · 26 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  7. Learning to Navigate Unseen Environments: Back Translation with Environmental Dropout
    2019/04/08 by Hao Tan, Tan, Hao, Licheng Yu +3 · 18 citations
    Computer Science · #Advanced Neural Network Applications #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
  8. TVQA+: Spatio-Temporal Grounding for Video Question Answering
    2019/04/25 by Lei, Jie, Yu, Licheng, Berg, Tamara L. +1 · 16 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  9. HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training
    2020/05/01 by Linjie Li, Yen-Chun Chen, Li, Linjie +9 · 14 citations
    Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Human Pose and Action Recognition #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
  10. VideoSwap: Customized Video Subject Swapping with Interactive Semantic Point Correspondence
    2023/12/04 by Yuchao Gu, Gu, Yuchao, Yipin Zhou +18 · 1 voice · 11 citations
    Computer Science · #Advanced Vision and Imaging #Generative Adversarial Networks and Image Synthesis #Video Analysis and Summarization #cs.CV
  11. Multi-Target Embodied Question Answering
    2019/04/09 by Licheng Yu, Yu, Licheng, Xinlei Chen +9 · 10 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Topic Modeling
  12. TVR: A Large-Scale Dataset for Video-Subtitle Moment Retrieval
    2020/01/24 by Jie Lei, Licheng Yu, Lei, Jie +5 · 10 citations
    Computer Science · #Advanced Image and Video Retrieval Techniques #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Multimodal Machine Learning Applications #Video Analysis and Summarization
  13. FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis
    2023/12/29 by Feng Liang, Liang, Feng, Bichen Wu +19 · 16 citations
    Computer Science · #Advanced Vision and Imaging #Video Coding and Compression Technologies #Computer Graphics and Visualization Techniques
  14. AVID: Any-Length Video Inpainting with Diffusion Model
    2023/12/06 by Zhixing Zhang, Bichen Wu, Zhang, Zhixing +15 · 14 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis
  15. A Joint Speaker-Listener-Reinforcer Model for Referring Expressions
    2016/12/30 by Licheng Yu, Yu, Licheng, Hao Tan +5 · 5 citations
    Computer Science · #Multimodal Machine Learning Applications #Speech and dialogue systems #Topic Modeling
  16. UNITER: UNiversal Image-TExt Representation Learning
    2019/09/25 by Chen, Yen-Chun, Li, Linjie, Yu, Licheng +5 · 5 citations
    #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  17. VALUE: A Multi-Task Benchmark for Video-and-Language Understanding Evaluation
    2021/06/08 by Linjie Li, Li, Linjie, Jie Lei +27 · 6 citations
    Computer Science · #Multimodal Machine Learning Applications #Topic Modeling #Domain Adaptation and Few-Shot Learning
  18. Behind the Scene: Revealing the Secrets of Pre-trained Vision-and-Language Models
    2020/05/15 by Cao, Jize, Gan, Zhe, Cheng, Yu +3 · 4 citations
    #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  19. Fairy: Fast Parallelized Instruction-Guided Video-to-Video Synthesis
    2023/12/20 by Wu, Bichen, Chuang, Ching-Yao, Wang, Xiaoyan +6 · 7 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  20. Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction
    2024/11/30 by Shiyu Zhao, Zhenting Wang, Zhao, Shiyu +17 · 10 citations
    Computer Science · #Multimodal Machine Learning Applications #Advanced Image and Video Retrieval Techniques #Domain Adaptation and Few-Shot Learning
  21. Learning and Verification of Task Structure in Instructional Videos
    2023/03/23 by Medhini Narasimhan, Licheng Yu, Narasimhan, Medhini +7 · 4 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Human Pose and Action Recognition #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
  22. Learning Procedure-aware Video Representation from Instructional Videos and Their Narrations
    2023/03/31 by Zhong, Yiwu, Yu, Licheng, Bai, Yang +3 · 4 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimedia (cs.MM)
  23. FAME-ViL: Multi-Tasking Vision-Language Model for Heterogeneous Fashion Tasks
    2023/03/04 by Xiao Han, Xiatian Zhu, Han, Xiao +9 · 4 citations
    Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications
  24. Visual Madlibs: Fill in the blank Image Generation and Question Answering
    2015/05/31 by Licheng Yu, Eunbyung Park, Yu, Licheng +5 · 5 citations
    Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  25. CommerceMM: Large-Scale Commerce MultiModal Representation Learning with Omni Retrieval
    2022/02/15 by Yu, Licheng, Chen, Jun, Sinha, Animesh +4 · 3 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimedia (cs.MM) #Social and Information Networks (cs.SI)
  26. VIOLIN: A Large-Scale Dataset for Video-and-Language Inference
    2020/03/25 by Jingzhou Liu, Liu, Jingzhou, Wenhu Chen +11 · 3 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  27. What is More Likely to Happen Next? Video-and-Language Future Event Prediction
    2020/10/15 by Lei, Jie, Yu, Licheng, Berg, Tamara L. +1 · 2 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  28. RoPAWS: Robust Semi-supervised Representation Learning from Uncurated Data
    2023/02/28 by Sangwoo Mo, Mo, Sangwoo, Jong-Chyi Su +10 · 3 citations
    Biochemistry, Genetics and Molecular Biology · Computer Science · #Cancer-related molecular mechanisms research #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Multimodal Machine Learning Applications
  29. CiT: Curation in Training for Effective Vision-Language Data
    2023/01/05 by Xu, Hu, Xie, Saining, Huang, Po-Yao +5 · 2 citations
    #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  30. Hierarchically-Attentive RNN for Album Summarization and Storytelling
    2017/08/09 by Licheng Yu, Mohit Bansal, Yu, Licheng +3 · 1 citation
    Computer Science · #Advanced Image and Video Retrieval Techniques #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Video Analysis and Summarization
  31. Connecting What to Say With Where to Look by Modeling Human Attention Traces
    2021/05/12 by Meng, Zihang, Yu, Licheng, Zhang, Ning +4 · 1 citation
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  32. GEB+: A Benchmark for Generic Event Boundary Captioning, Grounding and Retrieval
    2022/04/01 by Yuxuan Wang, Wang, Yuxuan, Difei Gao +9 · 1 citation
    Computer Science · #Human Pose and Action Recognition #Multimodal Machine Learning Applications #Machine Learning in Healthcare
  33. FashionViL: Fashion-Focused Vision-and-Language Representation Learning
    2022/07/17 by Han, Xiao, Yu, Licheng, Zhu, Xiatian +3 · 1 citation
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  34. Building a Mind Palace: Structuring Environment-Grounded Semantic Graphs for Effective Long Video Analysis with LLMs
    2025/01/08 by Zeyi Huang, Yuyang Ji, Huang, Zeyi +23 · 3 citations
    Computer Science · #Topic Modeling #Explainable Artificial Intelligence (XAI) #Semantic Web and Ontologies
  35. Tell Me What Happened: Unifying Text-guided Video Completion via Multimodal Masked Video Generation
    2022/11/23 by Fu, Tsu-Jui, Yu, Licheng, Zhang, Ning +4 · 1 citation
    #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  36. ROICtrl: Boosting Instance Control for Visual Generation
    2024/11/27 by Gu, Yuchao, Zhou, Yipin, Ye, Yunfan +5 · 2 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  37. AMELI: Enhancing Multimodal Entity Linking with Fine-Grained Attributes
    2023/05/24 by Barry Menglong Yao, Sijia Wang, Yao, Barry Menglong +13 · 1 citation
    Computer Science · #Topic Modeling #Text and Document Classification Technologies #Natural Language Processing Techniques
  38. AdvancedIF: Rubric-Based Benchmarking and Reinforcement Learning for Advancing LLM Instruction Following
    2025/11/13 by Yun He, He, Yun, Wenzhe Li +46 · 2 citations
    Computer Science · #Topic Modeling #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning