vix.ing · top · new · best · stats · spec

Zheng, Sipeng

  1. Steve-Eye: Equipping LLM-based Embodied Agents with Visual Perception in Open Worlds
    2023/10/20 by Sipeng Zheng, Zheng, Sipeng, Jiazheng Liu +5 · 8 citations
    Computer Science · #Multimodal Machine Learning Applications #Topic Modeling #Natural Language Processing Techniques
  2. Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
    2025/07/21 by Luo, Hao, Feng, Yicheng, Zhang, Wanpeng +7 · 25 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Robotics (cs.RO)
  3. Scaling Large Motion Models with Million-Level Human Motions
    2024/10/04 by Ye Wang, Wang, Ye, Bin Cao +7 · 8 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques
  4. LLaMA Rider: Spurring Large Language Models to Explore the Open World
    2023/10/13 by Feng, Yicheng, Wang, Yuxuan, Liu, Jiazheng +2 · 3 citations
    #FOS: Computer and information sciences #Machine Learning (cs.LG)
  5. From Pixels to Tokens: Byte-Pair Encoding on Quantized Visual Modalities
    2024/10/03 by Zhang, Wanpeng, Xie, Zilong, Feng, Yicheng +4 · 4 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  6. VideoOrion: Tokenizing Object Dynamics in Videos
    2024/11/25 by Feng, Yicheng, Li, Yijiang, Zhang, Wanpeng +4 · 4 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  7. Accommodating Audio Modality in CLIP for Multimodal Processing
    2023/03/12 by Ludan Ruan, Anwen Hu, Ruan, Ludan +9 · 2 citations
    Arts and Humanities · Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Music and Audio Processing #Subtitles and Audiovisual Media
  8. SPAFormer: Sequential 3D Part Assembly with Transformers
    2024/03/09 by Xu, Boshen, Zheng, Sipeng, Jin, Qin · 2 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Robotics (cs.RO)
  9. QuadrupedGPT: Towards a Versatile Quadruped Agent in Open-ended Worlds
    2024/06/24 by Yuting Mei, Ye Wang, Mei, Yuting +4 · 2 citations
    Computer Science · #Algorithms and Data Compression #Artificial Intelligence (cs.AI) #Distributed and Parallel Computing Systems #FOS: Computer and information sciences #Mobile Agent-Based Network Management #Robotics (cs.RO)
  10. RL from Physical Feedback: Aligning Large Motion Models with Humanoid Control
    2025/06/15 by Junpeng Yue, Yue, Junpeng, Zepeng Wang +15 · 4 citations
    Computer Science · Engineering · #Advanced Vision and Imaging #Autonomous Vehicle Technology and Safety #FOS: Computer and information sciences #Machine Learning (cs.LG) #Robotics (cs.RO)
  11. No-frills Temporal Video Grounding: Multi-Scale Neighboring Attention and Zoom-in Boundary Detection
    2023/07/20 by Qi Zhang, Zhang, Qi, Sipeng Zheng +3 · 1 citation
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications #Video Analysis and Summarization
  12. UniCode: Learning a Unified Codebook for Multimodal Large Language Models
    2024/03/14 by Sipeng Zheng, Bohan Zhou, Zheng, Sipeng +7 · 1 citation
    Computer Science · #Natural Language Processing Techniques #Topic Modeling #Speech and dialogue systems
  13. Do Egocentric Video-Language Models Truly Understand Hand-Object Interactions?
    2024/05/28 by Boshen Xu, Xu, Boshen, Ziheng Wang +9 · 1 citation
    Psychology · Social Sciences · #Action Observation and Synchronization #Social and Intergroup Psychology
  14. Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model
    2025/08/11 by Cao, Bin, Zheng, Sipeng, Wang, Ye +5 · 3 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  15. Unified Multimodal Understanding via Byte-Pair Visual Encoding
    2025/06/30 by Zhang, Wanpeng, Feng, Yicheng, Luo, Hao +4 · 2 citations
    #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences