Zhihong Shao
- DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
DeepSeek-R1 shows an LLM can learn strong step-by-step reasoning from pure reinforcement learning, with no human-labeled reasoning examples.
2025/01/22 by DeepSeek-AI, Daya Guo, Guo, Daya +404 · 93 voices · 1664 citations
Computer Science · #Reinforcement Learning in Robotics #Data Stream Mining Techniques #Explainable Artificial Intelligence (XAI)
- BioXP-0.5B: Explainable Medical-AI via RL-GRPO
2024/02/05 by Zhihong Shao, Peiyi Wang, Shao, Zhihong +20 · 17 voices · 2024 citations
Computer Science · #Mathematics, Computing, and Information Processing #Natural Language Processing Techniques #cs.AI #cs.CL #cs.LG
- DeepSeek-V3 Technical Report
2024/12/27 by DeepSeek-AI, Aixin Liu, Bei Feng +404 · 39 voices · 6 citations
Computer Science · Engineering · #Distributed and Parallel Computing Systems #Robotics and Automated Systems #cs.AI #cs.CL
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
2024/05/07 by Aixin Liu, DeepSeek-AI, Liu, Aixin +310 · 5 voices · 237 citations
Computer Science · #Expert finding and Q&A systems #Topic Modeling #Speech and dialogue systems
- DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data
2024/05/23 by Huajian Xin, Xin, Huajian, Daya Guo +16 · 2 voices · 50 citations
Decision Sciences · #Scientific Computing and Data Management
- CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
2023/05/19 by Zhibin Gou, Zhihong Shao, Gou, Zhibin +11 · 1 voice · 109 citations
Computer Science · #cs.CL #cs.AI
- DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence
2024/06/17 by Qihao Zhu, DeepSeek-AI, Daya Guo +79 · 1 voice · 70 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Software Engineering (cs.SE) #cs.AI #cs.LG #cs.SE #vaccines and immunoinformatics approaches
- DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
2024/01/05 by DeepSeek-AI, Xiao Guo Bi, : +170 · 109 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling #Text Readability and Simplification
- DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search
2024/08/15 by Huajian Xin, Z. Z. Ren, Xin, Huajian +32 · 2 voices · 39 citations
Computer Science · #Reinforcement Learning in Robotics #cs.AI #cs.CL #cs.LG #cs.LO
- DeepSeek-Prover-V2: Advancing Formal Mathematical Reasoning via Reinforcement Learning for Subgoal Decomposition
2025/04/30 by Z. Z. Ren, Zhihong Shao, Ren, Z. Z. +33 · 4 voices · 57 citations
#cs.CL #cs.AI
- DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning
2025/11/27 by Zhihong Shao, Shao, Zhihong, Yuxiang Luo +15 · 7 citations
Computer Science · Materials Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning in Materials Science #Mathematics, Computing, and Information Processing #Topic Modeling
- Learning Task Decomposition to Assist Humans in Competitive Programming
2024/06/07 by Jiaxin Wen, Wen, Jiaxin, Ruiqi Zhong +9 · 1 citation
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Programming Languages (cs.PL) #Reinforcement Learning in Robotics