Shuiping Yu
- DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
DeepSeek-R1 shows an LLM can learn strong step-by-step reasoning from pure reinforcement learning, with no human-labeled reasoning examples.
2025/01/22 by DeepSeek-AI, Daya Guo, Guo, Daya +404 · 93 voices · 2191 citations
Computer Science · #Reinforcement Learning in Robotics #Data Stream Mining Techniques #Explainable Artificial Intelligence (XAI)
- DeepSeek-V3 Technical Report
2024/12/27 by DeepSeek-AI, Aixin Liu, Liu, Aixin +404 · 39 voices · 7 citations
Computer Science · Engineering · #Distributed and Parallel Computing Systems #Robotics and Automated Systems #cs.AI #cs.CL
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
2024/05/07 by DeepSeek-AI, Aixin Liu, Liu, Aixin +310 · 5 voices · 306 citations
Computer Science · #Expert finding and Q&A systems #Topic Modeling #Speech and dialogue systems
- DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
2024/01/05 by DeepSeek-AI, Xiao Bi, : +178 · 2 voices · 155 citations
Computer Science · #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling #cs.AI #cs.CL #cs.LG
- Fire-Flyer AI-HPC: A Cost-Effective Software-Hardware Co-Design for Deep Learning
2024/08/26 by Wei An, Xiao Bi, An, Wei +107 · 2 voices · 5 citations
Computer Science · #Artificial Intelligence (cs.AI) #Distributed #FOS: Computer and information sciences #Parallel #Parallel Computing and Optimization Techniques #and Cluster Computing (cs.DC) #cs.AI #cs.DC