Lu, Chengda
- DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
DeepSeek-R1 shows an LLM can learn strong step-by-step reasoning from pure reinforcement learning, with no human-labeled reasoning examples.
2025/01/22 by DeepSeek-AI, Daya Guo, Guo, Daya +404 · 93 voices · 2169 citations
Computer Science · #Reinforcement Learning in Robotics #Data Stream Mining Techniques #Explainable Artificial Intelligence (XAI)
- DeepSeek-V3 Technical Report
2024/12/27 by DeepSeek-AI, Aixin Liu, Bei Feng +404 · 39 voices · 7 citations
Computer Science · Engineering · #Distributed and Parallel Computing Systems #Robotics and Automated Systems #cs.AI #cs.CL
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
2025/12/02 by DeepSeek-AI, Liu, Aixin, Mei, Aoxue +260 · 46 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences
- Does Chain-of-Thought Reasoning Really Reduce Harmfulness from Jailbreaking?
2025/05/23 by Lu, Chengda, Fan, Xiaoyu, Huang, Yu +3 · 5 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences
- A Generalized Probabilistic Monitoring Model with Both Random and Sequential Data
2022/06/27 by Yu, Wanke, Wu, Min, Huang, Biao +1 · 1 citation
#FOS: Electrical engineering #Systems and Control (eess.SY) #electronic engineering #information engineering
- DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning
2025/11/27 by Zhihong Shao, Shao, Zhihong, Yuxiang Luo +15 · 7 citations
Computer Science · Materials Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning in Materials Science #Mathematics, Computing, and Information Processing #Topic Modeling