vix.ing · top · new · best · stats · spec

Zhang, Hanning

  1. Mitigating the Alignment Tax of RLHF
    2023/09/12 by Yong Lin, Lin, Yong, Hangyu Lin +23 · 27 citations
    Computer Science · #FOS: Computer and information sciences #Image Retrieval and Classification Techniques #Machine Learning (cs.LG)
  2. R-Tuning: Instructing Large Language Models to Say `I Don't Know'
    2023/11/16 by Hanning Zhang, Shizhe Diao, Zhang, Hanning +15 · 15 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
  3. ScaleBiO: Scalable Bilevel Optimization for LLM Data Reweighting
    2024/06/28 by Rui Pan, Pan, Rui, Zhang, Dylan +12 · 11 citations
    Computer Science · Engineering · #Advanced Data Storage Technologies #Algorithms and Data Compression #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Optimization and Control (math.OC) #Reservoir Engineering and Simulation Methods
  4. Self-rewarding correction for mathematical reasoning
    2025/02/26 by Wei Xiong, Hanning Zhang, Xiong, Wei +9 · 15 citations
    Computer Science · #Advanced Graph Neural Networks #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Topic Modeling
  5. Entropy-Regularized Process Reward Model
    2024/12/15 by Zhang, Hanning, Wang, Pengcheng, Diao, Shizhe +6 · 10 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  6. Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL
    2025/05/05 by Jiarui Yao, Yifan Hao, Yao, Jiarui +11 · 8 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Data Stream Mining Techniques #FOS: Computer and information sciences #Machine Learning (cs.LG)
  7. OpenGenAlign: A Preference Dataset and Benchmark for Trustworthy Reward Modeling in Open-Ended, Long-Context Generation
    2025/01/22 by Zhang, Hanning, Song, Juntong, Zhu, Juno +3 · 1 citation
    #Computation and Language (cs.CL) #FOS: Computer and information sciences
  8. Understanding Overadaptation in Supervised Fine-Tuning: The Role of Ensemble Methods
    2025/06/02 by Hao, Yifan, Pan, Xingyuan, Zhang, Hanning +3 · 2 citations
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences