vix.ing · top · new · best · stats · spec

Hu, Hanpeng

  1. Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
    2025/02/14 by Guoqing Ma, Haoyang Huang, Ma, Guoqing +235 · 4 voices · 72 citations
    Computer Science · #Video Coding and Compression Technologies #cs.CL #cs.CV
  2. Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
    2025/02/17 by Ailin Huang, Boyong Wu, Huang, Ailin +299 · 1 voice · 46 citations
    Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Sound (cs.SD) #Speech and dialogue systems #cs.AI #cs.CL #cs.HC #cs.SD #eess.AS #electronic engineering #information engineering
  3. Optimizing RLHF Training for Large Language Models with Stage Fusion
    2024/09/20 by Yinmin Zhong, Zili Zhang, Zhong, Yinmin +18 · 23 citations
    Computer Science · #Computation and Language (cs.CL) #Distributed #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Parallel #Speech Recognition and Synthesis #Topic Modeling #and Cluster Computing (cs.DC)
  4. Step-Audio 2 Technical Report
    2025/07/22 by Boyong Wu, Chao Yan, Wu, Boyong +194 · 35 citations
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  5. StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation
    2025/04/22 by Yinmin Zhong, Zhong, Yinmin, Zili Zhang +25 · 19 citations
    Engineering · Computer Science · #VLSI and FPGA Design Techniques #Iterative Learning Control Systems #Algorithms and Data Compression
  6. NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
    2025/08/14 by NextStep Team, Han, Chunrui, Li, Guopeng +47 · 21 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  7. Step-3 is Large yet Affordable: Model-system Co-design for Cost-effective Decoding
    2025/07/25 by StepFun, :, Bin Wang +284 · 18 citations
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
  8. DistTrain: Addressing Model and Data Heterogeneity with Disaggregated Training for Multimodal Large Language Models
    2024/08/08 by Zili Zhang, Yinmin Zhong, Zhang, Zili +12 · 5 citations
    Computer Science · #Distributed #FOS: Computer and information sciences #Natural Language Processing Techniques #Parallel #Topic Modeling #and Cluster Computing (cs.DC)
  9. Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model
    2025/06/10 by Ailin Huang, Huang, Ailin, Bingxin Li +143 · 6 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  10. PipeWeaver: Addressing Data Dynamicity in Large Multimodal Model Training with Dynamic Interleaved Pipeline
    2025/04/19 by Xue, Zhenliang, Hu, Hanpeng, Chen, Xing +7 · 2 citations
    #Artificial Intelligence (cs.AI) #Distributed #FOS: Computer and information sciences #Parallel #and Cluster Computing (cs.DC)