Jiansheng Chen
- Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
2025/02/14 by Guoqing Ma, Haoyang Huang, Ma, Guoqing +235 · 4 voices · 45 citations
Computer Science · #Video Coding and Compression Technologies #cs.CL #cs.CV
- Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
2025/02/17 by Ailin Huang, Boyong Wu, Huang, Ailin +299 · 1 voice · 38 citations
Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Sound (cs.SD) #Speech and dialogue systems #cs.AI #cs.CL #cs.HC #cs.SD #eess.AS #electronic engineering #information engineering
- Learning a Graph Neural Network with Cross Modality Interaction for Image Fusion
2023/08/07 by Jiawei Li, Jiansheng Chen, Li, Jiawei +5 · 8 citations
Computer Science · Engineering · #Advanced Image Fusion Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #I.2 #I.4 #Remote-Sensing Image Classification #Visual Attention and Saliency Detection
- RhythmFormer: Extracting Patterned rPPG Signals based on Periodic Sparse Attention
2024/02/20 by Bochao Zou, Zizheng Guo, Zou, Bochao +9 · 9 citations
Engineering · Computer Science · Neuroscience · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- Isolated Diffusion: Optimizing Multi-Concept Text-to-Image Generation Training-Freely with Isolated Diffusion Guidance
2024/03/25 by Jingyuan Zhu, Zhu, Jingyuan, Huimin Ma +5 · 1 voice · 2 citations
#cs.CV
- Few-shot Image Generation with Diffusion Models
2022/11/07 by Jingyuan Zhu, Huimin Ma, Zhu, Jingyuan +5 · 5 citations
Computer Science · #Advanced Image Processing Techniques #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis
- Enhancing LLM Reasoning with Multi-Path Collaborative Reactive and Reflection agents
2024/12/31 by Cunfu He, He, Chengbo, Zou, Bochao +7 · 7 citations
Business, Management and Accounting · Computer Science · #Business Process Modeling and Analysis #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multi-Agent Systems and Negotiation #Semantic Web and Ontologies
- Step-Video-TI2V Technical Report: A State-of-the-Art Text-Driven Image-to-Video Generation Model
2025/03/14 by Haoyang Huang, Guoqing Ma, Huang, Haoyang +100 · 4 citations
Computer Science · #Advanced Vision and Imaging #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications
- Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model
2025/06/10 by Ailin Huang, Bingxin Li, Huang, Ailin +143 · 5 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models
2024/11/27 by Yudong Zhang, Zhang, Yudong, Ruobing Xie +10 · 2 citations
Neuroscience · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #EEG and Brain-Computer Interfaces #FOS: Computer and information sciences #Functional Brain Connectivity Studies
- Decoupling Cross-Modality Manifold Discrepancy: Leveraging Visible Diffusion Priors for Infrared Super-Resolution
2026/07/23 by Yunpeng Hua, Hongwei Yu, Jiawei Li +3
#cs.CV