vix.ing · top · new · best · stats · spec

Wang, Lijuan

  1. The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)
    2023/09/29 by Zhengyuan Yang, Linjie Li, Yang, Zhengyuan +11 · 11 voices · 71 citations
    Computer Science · #Multimodal Machine Learning Applications #Topic Modeling #Natural Language Processing Techniques
  2. Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
    2024/04/22 by Marah Abdin, Abdin, Marah, Jyoti Aneja +262 · 8 voices · 316 citations
    Computer Science · #Computational Physics and Python Applications
  3. GIT: A Generative Image-to-text Transformer for Vision and Language
    2022/05/27 by Jianfeng Wang, Zhengyuan Yang, Wang, Jianfeng +15 · 2 voices · 47 citations
    Computer Science · #Handwritten Text Recognition Techniques #Multimodal Machine Learning Applications #Natural Language Processing Techniques #cs.CV
  4. Segment Everything Everywhere All at Once
    2023/04/13 by Xueyan Zou, Zou, Xueyan, Jianwei Yang +15 · 1 voice · 66 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Topic Modeling #cs.CV
  5. MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
    2023/08/04 by Weihao Yu, Zhengyuan Yang, Yu, Weihao +13 · 184 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Multimodal Machine Learning Applications
  6. Grounded Language-Image Pre-training
    2021/12/07 by Li, Liunian Harold, Zhang, Pengchuan, Zhang, Haotian +9 · 130 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimedia (cs.MM)
  7. Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
    2023/06/26 by Fuxiao Liu, Liu, Fuxiao, Kevin Lin +9 · 69 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computational Engineering #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Finance #Multimedia (cs.MM) #Multimodal Machine Learning Applications #and Science (cs.CE)
  8. Large Scale Incremental Learning
    2019/05/30 by Yue Wu, Yinpeng Chen, Wu, Yue +11 · 39 citations
    Computer Science · #Domain Adaptation and Few-Shot Learning #Machine Learning and ELM #Multimodal Machine Learning Applications
  9. MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action
    2023/03/20 by Zhengyuan Yang, Yang, Zhengyuan, Linjie Li +16 · 60 citations
    Computer Science · #Multimodal Machine Learning Applications #Topic Modeling #Natural Language Processing Techniques
  10. Florence: A New Foundation Model for Computer Vision
    2021/11/22 by Lu Yuan, Yuan, Lu, Dongdong Chen +43 · 47 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Human Pose and Action Recognition #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
  11. Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks
    2020/04/13 by Li, Xiujun, Yin, Xi, Li, Chunyuan +9 · 34 citations
    #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Machine Learning (cs.LG)
  12. GenXD: Generating Any 3D and 4D Scenes
    2024/11/04 by Yuyang Zhao, Zhao, Yuyang, Chung-Ching Lin +15 · 1 voice · 17 citations
    #cs.CV #cs.AI
  13. Generalized Decoding for Pixel, Image, and Language
    2022/12/21 by Xueyan Zou, Zou, Xueyan, Zi-Yi Dou +25 · 33 citations
    Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  14. ShowUI: One Vision-Language-Action Model for GUI Visual Agent
    2024/11/26 by Kevin Qinghong Lin, Linjie Li, Lin, Kevin Qinghong +15 · 54 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Multimodal Machine Learning Applications
  15. RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
    2025/04/24 by Zihan Wang, Kangrui Wang, Wang, Zihan +33 · 77 citations
    Computer Science · #Reinforcement Learning in Robotics #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning
  16. Multimodal Foundation Models: From Specialists to General-Purpose Assistants
    2023/09/18 by Chunyuan Li, Zhe Gan, Li, Chunyuan +11 · 28 citations
    Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  17. End-to-End Human Pose and Mesh Reconstruction with Transformers
    2020/12/17 by Kevin Lin, Lijuan Wang, Lin, Kevin +3 · 19 citations
    Computer Science · Engineering · #3D Shape Modeling and Analysis #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition
  18. GLIPv2: Unifying Localization and Vision-Language Understanding
    2022/06/12 by Haotian Zhang, Pengchuan Zhang, Zhang, Haotian +17 · 21 citations
    Computer Science · #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning #Topic Modeling
  19. End-to-End Semi-Supervised Object Detection with Soft Teacher
    2021/06/16 by Mengde Xu, Xu, Mengde, Zheng Zhang +13 · 16 citations
    Computer Science · #Advanced Image and Video Retrieval Techniques #Advanced Neural Network Applications #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications
  20. DisCo: Disentangled Control for Realistic Human Dance Generation
    2023/06/30 by Wang, Tan, Li, Linjie, Lin, Kevin +6 · 22 citations
    #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  21. ReCo: Region-Controlled Text-to-Image Generation
    2022/11/23 by Zhengyuan Yang, Jianfeng Wang, Yang, Zhengyuan +19 · 19 citations
    Computer Science · #Multimodal Machine Learning Applications #Image Retrieval and Classification Techniques #Generative Adversarial Networks and Image Synthesis
  22. Mesh Graphormer
    2021/04/01 by Kevin Lin, Lijuan Wang, Lin, Kevin +3 · 15 citations
    Computer Science · Engineering · #Human Pose and Action Recognition #3D Shape Modeling and Analysis #Video Surveillance and Tracking Methods
  23. ART: Anonymous Region Transformer for Variable Multi-Layer Transparent Image Generation
    2025/02/25 by Yunfei Pu, Yifan Pu, Pu, Yifan +35 · 2 voices · 8 citations
    Computer Science · #Advanced Steganography and Watermarking Techniques #Chaos-based Image/Signal Encryption #Computer Vision and Pattern Recognition (cs.CV) #Digital Media Forensic Detection #FOS: Computer and information sciences #cs.CV
  24. An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQA
    2021/09/10 by Zhengyuan Yang, Zhe Gan, Yang, Zhengyuan +11 · 15 citations
    Computer Science · #Multimodal Machine Learning Applications #Advanced Image and Video Retrieval Techniques #Domain Adaptation and Few-Shot Learning
  25. NUWA-XL: Diffusion over Diffusion for eXtremely Long Video Generation
    2023/03/22 by Shengming Yin, Yin, Shengming, Chenfei Wu +28 · 19 citations
    Computer Science · #Generative Adversarial Networks and Image Synthesis #Video Analysis and Summarization #Video Coding and Compression Technologies
  26. Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
    2025/01/09 by Yi Hao, Hao, Yunzhuo, Jiawei Gu +11 · 37 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Natural Language Processing Techniques #Semantic Web and Ontologies #Speech and dialogue systems
  27. Prompting GPT-3 To Be Reliable
    2022/10/17 by Si, Chenglei, Gan, Zhe, Yang, Zhengyuan +4 · 15 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences
  28. GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation
    2023/11/13 by An Yan, Zhengyuan Yang, Yan, An +21 · 18 citations
    Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications #Social Robot Interaction and HRI
  29. Scaling Up Vision-Language Pre-training for Image Captioning
    2021/11/24 by Xiaowei Hu, Hu, Xiaowei, Zhe Gan +11 · 12 citations
    Computer Science · #Advanced Image and Video Retrieval Techniques #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications
  30. SwinBERT: End-to-End Transformers with Sparse Attention for Video Captioning
    2021/11/25 by Lin, Kevin, Li, Linjie, Lin, Chung-Ching +5 · 11 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  31. An Empirical Study of Training End-to-End Vision-and-Language Transformers
    2021/11/03 by Zi-Yi Dou, Yichong Xu, Dou, Zi-Yi +21 · 12 citations
    Computer Science · #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning #Advanced Neural Network Applications
  32. SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement
    2025/04/10 by Xiyao Wang, Zhengyuan Yang, Wang, Xiyao +15 · 40 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning and Data Classification #Multimodal Machine Learning Applications
  33. MM-VID: Advancing Video Understanding with GPT-4V(ision)
    2023/10/30 by Lin, Kevin, Ahmed, Faisal, Li, Linjie +9 · 13 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  34. MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities
    2024/08/01 by Weihao Yu, Zhengyuan Yang, Yu, Weihao +17 · 17 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Speech and dialogue systems
  35. NUWA-Infinity: Autoregressive over Autoregressive Generation for Infinite Visual Synthesis
    2022/07/20 by Chenfei Wu, Wu, Chenfei, Jian Liang +15 · 8 citations
    Computer Science · #Advanced Image and Video Retrieval Techniques #Advanced Vision and Imaging #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis
  36. UniTAB: Unifying Text and Box Outputs for Grounded Vision-Language Modeling
    2021/11/23 by Yang, Zhengyuan, Gan, Zhe, Wang, Jianfeng +5 · 7 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  37. GRiT: A Generative Region-to-text Transformer for Object Understanding
    2022/12/01 by Jialian Wu, Jianfeng Wang, Wu, Jialian +11 · 8 citations
    Computer Science · #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  38. SEED: Self-supervised Distillation For Visual Representation
    2021/01/12 by Fang, Zhiyuan, Wang, Jianfeng, Wang, Lijuan +3 · 6 citations
    #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  39. VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
    2021/11/24 by Fu, Tsu-Jui, Li, Linjie, Gan, Zhe +4 · 6 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  40. Motion Consistency Model: Accelerating Video Diffusion with Disentangled Motion-Appearance Distillation
    2024/06/11 by Yuanhao Zhai, Kevin Lin, Zhai, Yuanhao +15 · 11 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Image and Video Quality Assessment
  41. K-LITE: Learning Transferable Visual Models with External Knowledge
    2022/04/20 by Shen, Sheng, Li, Chunyuan, Hu, Xiaowei +11 · 6 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  42. VALUE: A Multi-Task Benchmark for Video-and-Language Understanding Evaluation
    2021/06/08 by Linjie Li, Li, Linjie, Jie Lei +27 · 6 citations
    Computer Science · #Multimodal Machine Learning Applications #Topic Modeling #Domain Adaptation and Few-Shot Learning
  43. Rethinking Classification and Localization for Object Detection
    2019/04/13 by Yue Wu, Yinpeng Chen, Wu, Yue +11 · 5 citations
    Computer Science · Engineering · #Advanced Neural Network Applications #Domain Adaptation and Few-Shot Learning #Medical Imaging and Analysis
  44. TAP: Text-Aware Pre-training for Text-VQA and Text-Caption
    2020/12/08 by Zhengyuan Yang, Yang, Zhengyuan, Yijuan Lu +15 · 5 citations
    Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications
  45. An Empirical Study of Multimodal Model Merging
    2023/04/28 by Yi-Lin Sung, Linjie Li, Sung, Yi-Lin +9 · 6 citations
    Computer Science · #Multimodal Machine Learning Applications #Topic Modeling #Domain Adaptation and Few-Shot Learning
  46. MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos
    2024/06/12 by Xuehai He, He, Xuehai, Weixi Feng +25 · 9 citations
    Computer Science · #Advanced Vision and Imaging #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Video Surveillance and Tracking Methods
  47. MM-Narrator: Narrating Long-form Videos with Multimodal In-Context Learning
    2023/11/29 by Chaoyi Zhang, Zhang, Chaoyi, Kevin Lin +13 · 7 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Music and Audio Processing #Natural Language Processing Techniques #Video Analysis and Summarization
  48. Compressing Visual-linguistic Model via Knowledge Distillation
    2021/04/05 by Fang, Zhiyuan, Wang, Jianfeng, Hu, Xiaowei +3 · 4 citations
    #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  49. List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs
    2024/04/25 by Yan An, Yan, An, Zhengyuan Yang +19 · 8 citations
    Arts and Humanities · Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Lexicography and Language Studies #Natural Language Processing Techniques #Semantic Web and Ontologies
  50. Equivariant Similarity for Vision-Language Foundation Models
    2023/03/25 by Wang, Tan, Lin, Kevin, Li, Linjie +5 · 5 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  51. Glyph-ByT5-v2: A Strong Aesthetic Baseline for Accurate Multilingual Visual Text Rendering
    2024/06/14 by Zeyu Liu, Liu, Zeyu, Weicong Liang +9 · 7 citations
    Engineering · Computer Science · #Human Motion and Animation #Handwritten Text Recognition Techniques #Advanced Image and Video Retrieval Techniques
  52. ImageGen-CoT: Enhancing Text-to-Image In-context Learning with Chain-of-Thought Reasoning
    2025/03/25 by Liao, Jiaqi, Yang, Zhengyuan, Li, Linjie +4 · 13 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  53. LAVENDER: Unifying Video-Language Understanding as Masked Language Modeling
    2022/06/14 by Li, Linjie, Gan, Zhe, Lin, Kevin +4 · 4 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  54. An Empirical Study of End-to-End Video-Language Transformers with Masked Visual Modeling
    2022/09/04 by Tsu-Jui Fu, Linjie Li, Fu, Tsu-Jui +11 · 4 citations
    Computer Science · #Multimodal Machine Learning Applications #Advanced Neural Network Applications #Domain Adaptation and Few-Shot Learning
  55. Vision-Language Pre-training: Basics, Recent Advances, and Future Trends
    2022/10/17 by Gan, Zhe, Li, Linjie, Li, Chunyuan +3 · 4 citations
    #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  56. Non-Contrastive Learning Meets Language-Image Pre-Training
    2022/10/17 by Zhou, Jinghao, Dong, Li, Gan, Zhe +2 · 4 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  57. SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation
    2024/10/30 by Yining Hong, Hong, Yining, Baiyu Liu +21 · 7 citations
    Biochemistry, Genetics and Molecular Biology · Computer Science · #Advanced Vision and Imaging #Artificial Intelligence (cs.AI) #Cell Image Analysis Techniques #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Machine Learning (cs.LG) #Robotics (cs.RO)
  58. EditRoom: LLM-parameterized Graph Diffusion for Composable 3D Room Layout Editing
    2024/10/03 by Zheng, Kaizhi, Xiaotong Chen, Chen, Xiaotong +16 · 7 citations
    Computer Science · Engineering · #3D Shape Modeling and Analysis #Artificial Intelligence (cs.AI) #Computer Graphics and Visualization Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Graph Theory and Algorithms #Graphics (cs.GR) #Human-Computer Interaction (cs.HC)
  59. MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models
    2024/10/14 by Peng Xia, Xia, Peng, Siwei Han +21 · 7 citations
    Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  60. Coarse-to-Fine Vision-Language Pre-training with Fusion in the Backbone
    2022/06/15 by Dou, Zi-Yi, Kamath, Aishwarya, Gan, Zhe +9 · 3 citations
    #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  61. Idea2Img: Iterative Self-Refinement with GPT-4V(ision) for Automatic Image Design and Generation
    2023/10/12 by Zhengyuan Yang, Yang, Zhengyuan, Jianfeng Wang +11 · 4 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  62. Segment and Caption Anything
    2023/12/01 by Huang, Xiaoke, Wang, Jianfeng, Tang, Yansong +5 · 4 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  63. Tuning Timestep-Distilled Diffusion Model Using Pairwise Sample Optimization
    2024/10/04 by Zichen Miao, Zhengyuan Yang, Miao, Zichen +11 · 6 citations
    Computer Science · #Neural Networks and Applications
  64. IDOL: Unified Dual-Modal Latent Diffusion for Human-Centric Joint Video-Depth Generation
    2024/07/15 by Zhai, Yuanhao, Lin, Kevin, Li, Linjie +7 · 5 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  65. Audio-Aware Large Language Models as Judges for Speaking Styles
    2025/06/06 by Cheng-Han Chiang, Xiaofei Wang, Chiang, Cheng-Han +18 · 10 citations
    Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #electronic engineering #information engineering
  66. MMSum: A Dataset for Multimodal Summarization and Thumbnail Generation of Videos
    2023/06/07 by Jielin Qiu, Jiacheng Zhu, Qiu, Jielin +21 · 3 citations
    Computer Science · #Video Analysis and Summarization #Natural Language Processing Techniques #Music and Audio Processing
  67. TextAtlas5M: A Large-scale Dataset for Dense Text Image Generation
    2025/02/11 by Wang, Alex Jinpeng, Mao, Dongxing, Zhang, Jiawei +9 · 7 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  68. ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs
    2025/06/11 by Xiyao Wang, Zhengyuan Yang, Wang, Xiyao +22 · 11 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
  69. Playing Lottery Tickets with Vision and Language
    2021/04/23 by Gan, Zhe, Chen, Yen-Chun, Li, Linjie +6 · 2 citations
    #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  70. Zero-Shot Audio-Visual Editing via Cross-Modal Delta Denoising
    2025/03/26 by Lin, Yan-Bo, Lin, Kevin, Yang, Zhengyuan +6 · 6 citations
    #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
  71. COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training
    2024/01/01 by Alex Jinpeng Wang, Wang, Alex Jinpeng, Linjie Li +13 · 3 citations
    Computer Science · #Multimodal Machine Learning Applications #Topic Modeling #Natural Language Processing Techniques
  72. Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning
    2025/05/26 by Minheng Ni, Ni, Minheng, Zhengyuan Yang +11 · 10 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Natural Language Processing Techniques #Semantic Web and Ontologies #Speech and dialogue systems
  73. StrokeNUWA: Tokenizing Strokes for Vector Graphic Synthesis
    2024/01/30 by Zecheng Tang, Tang, Zecheng, Chenfei Wu +19 · 3 citations
    Engineering · #Additive Manufacturing and 3D Printing Technologies #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Manufacturing Process and Optimization
  74. VideoGUI: A Benchmark for GUI Automation from Instructional Videos
    2024/06/14 by Kevin Qinghong Lin, Lin, Kevin Qinghong, Linjie Li +13 · 5 citations
    Computer Science · #Video Analysis and Summarization
  75. Leveraging Visual Tokens for Extended Text Contexts in Multi-Modal Learning
    2024/06/04 by Alex Jinpeng Wang, Linjie Li, Wang, Alex Jinpeng +9 · 4 citations
    Computer Science · #Multimodal Machine Learning Applications #Open Education and E-Learning #Speech and dialogue systems
  76. VAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM Agents
    2025/10/19 by Kangrui Wang, Pingyue Zhang, Wang, Kangrui +27 · 10 citations
    Computer Science · #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning #Topic Modeling
  77. Diagnostic Benchmark and Iterative Inpainting for Layout-Guided Image Generation
    2023/04/13 by Cho, Jaemin, Li, Linjie, Yang, Zhengyuan +3 · 2 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  78. Impulse and sampled-data optimal control of heat equations, and error estimates
    2015/09/19 by Emmanuel Trélat, Trélat, Emmanuel, Lijuan Wang +3 · 1 citation
    Computer Science · Engineering · Mathematics · #Advanced Mathematical Modeling in Engineering #FOS: Mathematics #Numerical methods in inverse problems #Optimization and Control (math.OC) #Stability and Controllability of Differential Equations
  79. ORES: Open-vocabulary Responsible Visual Synthesis
    2023/08/26 by Ni, Minheng, Wu, Chenfei, Wang, Xiaodong +4 · 2 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  80. Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension
    2024/12/04 by Wang, Xiyao, Yang, Zhengyuan, Li, Linjie +6 · 4 citations
    #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  81. Incremental Classifier Learning with Generative Adversarial Networks
    2018/02/02 by Wu, Yue, Chen, Yinpeng, Wang, Lijuan +5 · 1 citation
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  82. STITCH: Simultaneous Thinking and Talking with Chunked Reasoning for Spoken Language Models
    2025/07/21 by Chiang, Cheng-Han, Wang, Xiaofei, Li, Linjie +7 · 10 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
  83. DEsignBench: Exploring and Benchmarking DALL-E 3 for Imagining Visual Design
    2023/10/23 by Lin, Kevin, Yang, Zhengyuan, Li, Linjie +2 · 2 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  84. MiniVLM: A Smaller and Faster Vision-Language Model
    2020/12/13 by Wang, Jianfeng, Hu, Xiaowei, Zhang, Pengchuan +5 · 1 citation
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  85. Adversarial Feature Augmentation and Normalization for Visual Recognition
    2021/03/22 by Chen, Tianlong, Cheng, Yu, Gan, Zhe +4 · 1 citation
    #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  86. SITE: towards Spatial Intelligence Thorough Evaluation
    2025/05/08 by Wang, Wenqi, Tan, Reuben, Zhu, Pengyue +6 · 4 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  87. Exploring Discrete Diffusion Models for Image Captioning
    2022/11/21 by Zixin Zhu, Yixuan Wei, Zhu, Zixin +17 · 1 citation
    Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Video Analysis and Summarization
  88. MPT: Mesh Pre-Training with Transformers for Human Pose and Mesh Reconstruction
    2022/11/24 by Kevin Lin, Lin, Kevin, Chung-Ching Lin +6 · 1 citation
    Computer Science · Medicine · #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #Diabetic Foot Ulcer Assessment and Management #FOS: Computer and information sciences #Human Pose and Action Recognition
  89. Adaptive Human Matting for Dynamic Videos
    2023/04/12 by Chung-Ching Lin, Lin, Chung-Ching, Jiang Wang +11 · 1 citation
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Image Enhancement Techniques #Visual Attention and Saliency Detection
  90. Learning 3D Photography Videos via Self-supervised Diffusion on Single Images
    2023/02/21 by Wang, Xiaodong, Wu, Chenfei, Yin, Shengming +9 · 1 citation
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  91. Cross-border Commodity Pricing Strategy Optimization via Mixed Neural Network for Time Series Analysis
    2024/08/22 by Wang, Lijuan, Hu, Yijia, Zhou, Yan · 2 citations
    #Computational Engineering #FOS: Computer and information sciences #FOS: Economics and business #Finance #General Economics (econ.GN) #Machine Learning (cs.LG) #and Science (cs.CE)
  92. Spatial-Frequency U-Net for Denoising Diffusion Probabilistic Models
    2023/07/27 by Yuan, Xin, Li, Linjie, Wang, Jianfeng +4 · 1 citation
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  93. What makes Reasoning Models Different? Follow the Reasoning Leader for Efficient Decoding
    2025/06/08 by Ming Li, Li, Ming, Zhengyuan Yang +11 · 5 citations
    Computer Science · #Topic Modeling #Multimodal Machine Learning Applications #Constraint Satisfaction and Optimization
  94. OpenLEAF: Open-Domain Interleaved Image-Text Generation and Evaluation
    2023/10/11 by An, Jie, Yang, Zhengyuan, Li, Linjie +5 · 1 citation
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  95. Interfacing Foundation Models' Embeddings
    2023/12/12 by Zou, Xueyan, Li, Linjie, Wang, Jianfeng +10 · 1 citation
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  96. Constrained approximate null controllability of coupled heat equation with periodic impulse controls
    2020/05/15 by Lijuan Wang, Wang, Lijuan, Qishu Yan +3 · 1 citation
    Computer Science · Engineering · Mathematics · #35K40 #93B05 #93C20 #Advanced Mathematical Modeling in Engineering #FOS: Mathematics #Nonlinear Differential Equations Analysis #Optimization and Control (math.OC) #Stability and Controllability of Differential Equations
  97. Entity6K: A Large Open-Domain Evaluation Dataset for Real-World Entity Recognition
    2024/03/19 by Qiu, Jielin, Han, William, Wang, Winfred +6 · 1 citation
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  98. Certainly Uncertain: A Benchmark and Metric for Multimodal Epistemic and Aleatoric Awareness
    2024/07/02 by Chandu, Khyathi Raghavi, Li, Linjie, Awadalla, Anas +5 · 1 citation
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  99. GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?
    2025/07/13 by Yiyang Zhou, Linjie Li, Zhou, Yiyang +21 · 6 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications
  100. Conditional Text-to-Image Generation with Reference Guidance
    2024/11/22 by Kim, Taewook, Wang, Ze, Yang, Zhengyuan +4 · 1 citation
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  101. Beyond Words: Advancing Long-Text Image Generation via Multimodal Autoregressive Models
    2025/03/26 by Wang, Alex Jinpeng, Li, Linjie, Yang, Zhengyuan +2 · 1 citation
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences