Hu, Hanpeng
- Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
2025/02/14 by Guoqing Ma, Haoyang Huang, Ma, Guoqing +235 · 4 voices · 72 citations
Computer Science · #Video Coding and Compression Technologies #cs.CL #cs.CV
- Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
2025/02/17 by Ailin Huang, Boyong Wu, Huang, Ailin +299 · 1 voice · 46 citations
Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Sound (cs.SD) #Speech and dialogue systems #cs.AI #cs.CL #cs.HC #cs.SD #eess.AS #electronic engineering #information engineering
- Optimizing RLHF Training for Large Language Models with Stage Fusion
2024/09/20 by Yinmin Zhong, Zili Zhang, Zhong, Yinmin +18 · 23 citations
Computer Science · #Computation and Language (cs.CL) #Distributed #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Parallel #Speech Recognition and Synthesis #Topic Modeling #and Cluster Computing (cs.DC)
- Step-Audio 2 Technical Report
2025/07/22 by Boyong Wu, Chao Yan, Wu, Boyong +194 · 35 citations
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation
2025/04/22 by Yinmin Zhong, Zhong, Yinmin, Zili Zhang +25 · 19 citations
Engineering · Computer Science · #VLSI and FPGA Design Techniques #Iterative Learning Control Systems #Algorithms and Data Compression
- NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
2025/08/14 by NextStep Team, Han, Chunrui, Li, Guopeng +47 · 21 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- Step-3 is Large yet Affordable: Model-system Co-design for Cost-effective Decoding
2025/07/25 by StepFun, :, Bin Wang +284 · 18 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
- DistTrain: Addressing Model and Data Heterogeneity with Disaggregated Training for Multimodal Large Language Models
2024/08/08 by Zili Zhang, Yinmin Zhong, Zhang, Zili +12 · 5 citations
Computer Science · #Distributed #FOS: Computer and information sciences #Natural Language Processing Techniques #Parallel #Topic Modeling #and Cluster Computing (cs.DC)
- Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model
2025/06/10 by Ailin Huang, Huang, Ailin, Bingxin Li +143 · 6 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- PipeWeaver: Addressing Data Dynamicity in Large Multimodal Model Training with Dynamic Interleaved Pipeline
2025/04/19 by Xue, Zhenliang, Hu, Hanpeng, Chen, Xing +7 · 2 citations
#Artificial Intelligence (cs.AI) #Distributed #FOS: Computer and information sciences #Parallel #and Cluster Computing (cs.DC)