vix.ing · top · new · best · stats · spec

Yunfan Xiong

  1. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning
    DeepSeek-R1 shows an LLM can learn strong step-by-step reasoning from pure reinforcement learning, with no human-labeled reasoning examples.
    2025/01/22 by DeepSeek-AI, Daya Guo, Guo, Daya +404 · 93 voices · 1689 citations
    Computer Science · #Reinforcement Learning in Robotics #Data Stream Mining Techniques #Explainable Artificial Intelligence (XAI)
  2. DeepSeek-V3 Technical Report
    2024/12/27 by DeepSeek-AI, Aixin Liu, Bei Feng +404 · 39 voices · 6 citations
    Computer Science · Engineering · #Distributed and Parallel Computing Systems #Robotics and Automated Systems #cs.AI #cs.CL
  3. Exploiting Prefix-Tree in Structured Output Interfaces for Enhancing Jailbreak Attacking
    2025/02/19 by Yanzeng Li, Li, Yanzeng, Yunfan Xiong +8 · 2 citations
    Computer Science · #Advanced Malware Detection Techniques #Artificial Intelligence (cs.AI) #Cryptography and Security (cs.CR) #Digital and Cyber Forensics #FOS: Computer and information sciences #Network Security and Intrusion Detection