vix.ing · top · new · best · stats · spec

Yaohui Wang

  1. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
    DeepSeek-R1 shows an LLM can learn strong step-by-step reasoning from pure reinforcement learning, with no human-labeled reasoning examples.
    2025/01/22 by DeepSeek-AI, Daya Guo, Dejian Yang +404 · 93 voices · 1687 citations
    Computer Science · #Reinforcement Learning in Robotics #Data Stream Mining Techniques #Explainable Artificial Intelligence (XAI)
  2. DeepSeek-V3 Technical Report
    2024/12/27 by DeepSeek-AI, Aixin Liu, Liu, Aixin +404 · 39 voices · 6 citations
    Computer Science · Engineering · #Distributed and Parallel Computing Systems #Robotics and Automated Systems #cs.AI #cs.CL
  3. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
    2024/05/07 by DeepSeek-AI, Aixin Liu, Liu, Aixin +310 · 5 voices · 244 citations
    Computer Science · #Expert finding and Q&A systems #Topic Modeling #Speech and dialogue systems
  4. VBench: Comprehensive Benchmark Suite for Video Generative Models
    2023/11/29 by Ziqi Huang, Huang, Ziqi, Yinan He +29 · 347 citations
    Computer Science · #Generative Adversarial Networks and Image Synthesis #Visual Attention and Saliency Detection #Human Pose and Action Recognition
  5. AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning
    2023/07/10 by Yuwei Guo, Ceyuan Yang, Guo, Yuwei +13 · 211 citations
    Computer Science · Engineering · #Image Retrieval and Classification Techniques #Human Motion and Animation #Music and Audio Processing
  6. DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence
    2024/06/17 by DeepSeek-AI, Qihao Zhu, Daya Guo +79 · 1 voice · 70 citations
    Biochemistry, Genetics and Molecular Biology · Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Software Engineering (cs.SE) #cs.AI #cs.LG #cs.SE #vaccines and immunoinformatics approaches
  7. DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
    2024/01/05 by DeepSeek-AI, Xiao Guo Bi, : +170 · 110 citations
    Computer Science · #Natural Language Processing Techniques #Topic Modeling #Text Readability and Simplification
  8. InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
    2023/07/13 by Yi Wang, Yinan He, Wang, Yi +26 · 84 citations
    Computer Science · #Multimodal Machine Learning Applications #Human Pose and Action Recognition #Video Analysis and Summarization
  9. Latte: Latent Diffusion Transformer for Video Generation
    2024/01/05 by Xin Ma, Ma, Xin, Yaohui Wang +13 · 78 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis
  10. LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models
    2023/09/26 by Yaohui Wang, Xinyuan Chen, Wang, Yaohui +37 · 41 citations
    Computer Science · #Generative Adversarial Networks and Image Synthesis #Advanced Vision and Imaging
  11. SEINE: Short-to-Long Video Diffusion Model for Generative Transition and Prediction
    2023/10/31 by Xinyuan Chen, Yaohui Wang, Chen, Xinyuan +17 · 35 citations
    Computer Science · Economics, Econometrics and Finance · #Generative Adversarial Networks and Image Synthesis #Cinema and Media Studies
  12. 4Diffusion: Multi-view Video Diffusion Model for 4D Generation
    2024/05/31 by Haiyu Zhang, Zhang, Haiyu, Xinyuan Chen +9 · 33 citations
    Social Sciences · #Multimedia Communication and Technology
  13. Vlogger: Make Your Dream A Vlog
    2024/01/17 by Shaobin Zhuang, Kunchang Li, Zhuang, Shaobin +11 · 1 voice · 15 citations
    Computer Science · #Human Pose and Action Recognition #Multimodal Machine Learning Applications #Video Analysis and Summarization #cs.AI #cs.CV #cs.LG #cs.MM
  14. In situ Raman spectroscopy reveals the structure and dissociation of interfacial water
    2021/12/01 by Yaohui Wang, Yao-Hui Wang, Shisheng Zheng +20 · 15 citations
    Chemistry · Energy · Physics and Astronomy · #Electrocatalysts for Energy Conversion #Electrochemical Analysis and Applications #Spectroscopy and Quantum Chemical Studies
  15. Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models
    2025/01/14 by Chenyang Si, Fan, Weichen, Si, Chenyang +31 · 19 citations
    Computer Science · #Advanced Vision and Imaging #Computer Graphics and Visualization Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Image and Signal Denoising Methods #Machine Learning (cs.LG)
  16. UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
    2025/09/02 by Haoming Wang, Wang, Haoming, Haoyang Zou +201 · 50 citations
    Computer Science · Psychology · Engineering · #Context-Aware Activity Recognition Systems #Social Robot Interaction and HRI #Robotics and Automated Systems
  17. ViA: View-invariant Skeleton Action Representation Learning via Motion Retargeting
    2022/08/31 by Di Yang, Yaohui Wang, Yang, Di +9 · 5 citations
    Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Gait Recognition and Analysis #Human Pose and Action Recognition
  18. LAC: Latent Action Composition for Skeleton-based Action Segmentation
    2023/08/28 by Di Yang, Yang, Di, Yaohui Wang +11 · 5 citations
    Computer Science · #Human Pose and Action Recognition #Multimodal Machine Learning Applications #Anomaly Detection Techniques and Applications
  19. Self-Supervised Video Representation Learning via Latent Time Navigation
    2023/05/10 by Di Yang, Yaohui Wang, Yang, Di +11 · 4 citations
    Computer Science · Engineering · #Anomaly Detection Techniques and Applications #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Gait Recognition and Analysis #Human Pose and Action Recognition
  20. Hierarchical Diffusion Autoencoders and Disentangled Image Manipulation
    2023/04/24 by Zeyu Lu, Lu, Zeyu, Chengyue Wu +11 · 3 citations
    Computer Science · #Generative Adversarial Networks and Image Synthesis #Advanced Image Processing Techniques #Digital Media Forensic Detection
  21. UNIK: A Unified Framework for Real-world Skeleton-based Action Recognition
    2021/07/19 by Di Yang, Yaohui Wang, Yang, Di +9 · 2 citations
    Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Gait Recognition and Analysis #Hand Gesture Recognition Systems #Human Pose and Action Recognition
  22. AccVideo: Accelerating Video Diffusion Model with Synthetic Dataset
    2025/03/25 by Haiyu Zhang, Xinyuan Chen, Zhang, Haiyu +9 · 1 voice · 3 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #cs.CV
  23. VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models
    2024/11/20 by Ziqi Huang, Huang, Ziqi, Fan Zhang +31 · 4 citations
    Computer Science · Engineering · #Advanced Vision and Imaging #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Human Motion and Animation
  24. CineTrans: Learning to Generate Videos with Cinematic Transitions via Masked Diffusion Models
    2025/08/15 by Xiaoxue Wu, Bin Gao, Wu, Xiaoxue +7 · 7 citations
    Computer Science · #Video Analysis and Summarization #Generative Adversarial Networks and Image Synthesis #Human Pose and Action Recognition
  25. The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video Generation
    2025/04/16 by Bin Gao, Xinyu Gao, Gao, Bingjie +13 · 5 citations
    Computer Science · Engineering · Social Sciences · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Motion and Animation #Multimedia Communication and Technology #Video Analysis and Summarization
  26. Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model
    2024/12/30 by Yifei Huang, Huang, Yifei, Jilan Xu +33 · 3 citations
    Engineering · #Robotics and Automated Systems
  27. MuLan: Adapting Multilingual Diffusion Models for Hundreds of Languages with Negligible Cost
    2024/12/02 by Sen Xing, Xing, Sen, Zeqiang Lai +12 · 2 citations
    Computer Science · #Natural Language Processing Techniques #Text Readability and Simplification
  28. An Egocentric Vision-Language Model based Portable Real-time Smart Assistant
    2025/03/06 by Yifei Huang, Huang, Yifei, Jilan Xu +35 · 2 citations
    Computer Science · Psychology · #Multimodal Machine Learning Applications #Social Robot Interaction and HRI #Advanced Neural Network Applications
  29. DeforM: Reasoning-Guided Physics-Aware Video Generation via Spatial-Temporal Masking
    2026/07/21 by Yunyi Li, Yu Qiao, Yaohui Wang +1
    #cs.CV