vix.ing · top · new · best · stats · spec

Rui Qian

  1. Attentive Generative Adversarial Network for Raindrop Removal from a Single Image
    2017/11/28 by Rui Qian, Qian, Rui, Robby T. Tan +7 · 2 voices · 21 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #cs.CV
  2. Simple Copy-Paste is a Strong Data Augmentation Method for Instance\n Segmentation
    2020/12/13 by Golnaz Ghiasi, Yin Cui, Ghiasi, Golnaz +13 · 36 citations
    Computer Science · #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Medical Image Segmentation Techniques
  3. Seed1.5-VL Technical Report
    2025/05/11 by Dong Guo, Guo, Dong, Faming Wu +287 · 112 citations
    Computer Science · #Advanced Neural Network Applications #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Multimodal Machine Learning Applications
  4. Spatiotemporal Contrastive Video Representation Learning
    2020/08/09 by Rui Qian, Qian, Rui, Tianjian Meng +11 · 22 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Human Pose and Action Recognition #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
  5. Streaming Long Video Understanding with Large Language Models
    2024/05/25 by Rui Qian, Xiaoyi Dong, Qian, Rui +11 · 38 citations
    Computer Science · Medicine · #COVID-19 diagnosis using AI #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning in Healthcare #Multimodal Machine Learning Applications
  6. VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text
    2021/04/22 by Hassan Akbari, Akbari, Hassan, Liangzhe Yuan +11 · 21 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Human Pose and Action Recognition #Image and Video Processing (eess.IV) #Machine Learning (cs.LG) #Multimedia (cs.MM) #Multimodal Machine Learning Applications #Music and Audio Processing #electronic engineering #information engineering
  7. Learning Hierarchical Cross-Modal Association for Co-Speech Gesture Generation
    2022/03/24 by Xian Liu, Qianyi Wu, Liu, Xian +17 · 11 citations
    Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Hand Gesture Recognition Systems #Human Motion and Animation #Human Pose and Action Recognition
  8. OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?
    2025/01/09 by Yifei Li, Li, Yifei, Junbo Niu +26 · 21 citations
    Computer Science · #Video Analysis and Summarization
  9. Multiple Sound Sources Localization from Coarse to Fine
    2020/07/13 by Rui Qian, Di Hu, Qian, Rui +9 · 7 citations
    Computer Science · Neuroscience · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Hearing Loss and Rehabilitation #Music and Audio Processing #Speech and Audio Processing
  10. Imagen 3
    2024/08/13 by Imagen-Team-Google, :, Jason Baldridge +526 · 1 voice · 3 citations
    Computer Science · Medicine · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Radiomics and Machine Learning in Medical Imaging #cs.CV
  11. InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
    2024/12/12 by Pan Zhang, Xiaoyi Dong, Zhang, Pan +53 · 14 citations
    Computer Science · Social Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimedia Communication and Technology #Speech and dialogue systems #Video Analysis and Summarization
  12. Prune Spatio-temporal Tokens by Semantic-aware Temporal Accumulation
    2023/08/08 by Shuangrui Ding, Ding, Shuangrui, Peisen Zhao +9 · 5 citations
    Computer Science · #Human Pose and Action Recognition #Video Surveillance and Tracking Methods #Generative Adversarial Networks and Image Synthesis
  13. Semantics Meets Temporal Correspondence: Self-supervised Object-centric Learning in Videos
    2023/08/19 by Rui Qian, Qian, Rui, Shuangrui Ding +5 · 3 citations
    Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications
  14. SongComposer: A Large Language Model for Lyric and Melody Generation in Song Composition
    2024/02/27 by Shuangrui Ding, Ding, Shuangrui, Zihan Liu +14 · 3 citations
    Computer Science · #Music and Audio Processing
  15. Betrayed by Attention: A Simple yet Effective Approach for Self-supervised Video Object Segmentation
    2023/11/29 by Shuangrui Ding, Ding, Shuangrui, Rui Qian +7 · 2 citations
    Computer Science · #Advanced Image and Video Retrieval Techniques #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Visual Attention and Saliency Detection
  16. Static and Dynamic Concepts for Self-supervised Video Representation Learning
    2022/07/26 by Rui Qian, Shuangrui Ding, Qian, Rui +5 · 1 citation
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Human Pose and Action Recognition #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
  17. CAR-Flow: Condition-Aware Reparameterization Aligns Source and Target for Better Flow Matching
    2025/09/23 by Chen Chen, Pengsheng Guo, Chen, Chen +17 · 1 citation
    Computer Science · #Generative Adversarial Networks and Image Synthesis #Advanced Neural Network Applications #Domain Adaptation and Few-Shot Learning
  18. Why Are GUI Agents Correct but Late? Decode on the Decision-Time Critical Path, Tested with Pre-Compiled Policy Trees
    2026/07/30 by Zihan Dong, Rui Qian, Qishi Zhan +3
    Computer Science · #cs.LG
  19. How Benchmarks Mis-Score Computer-Use Agents
    2026/07/30 by Zihan Dong, Zhiyuan Ma, Zekun Wang +5
    Computer Science · #cs.AI
  20. Agents Don't Just Agree, They Remember: Benchmarking Persistent Sycophancy in Stateful Personal Agents
    2026/07/27 by Xutao Mao, Liangjie Zhao, Leyao Wang +6
    #cs.AI
  21. Seeing What Is Actually There: PriVE-Bench and PriVE-Tools for Counterfactual Evaluation of Agentic Visual Evidence in VLMs
    2026/07/14 by Jingyu Sun, Jiachen Tu, Yuyang Xue +8
    #cs.CV #cs.AI