Huang, Audrey
- Self-Improvement in Language Models: The Sharpening Mechanism
2024/12/02 by Audrey Huang, Huang, Audrey, Adam Block +13 · 2 voices · 26 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG #stat.ML
- Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment
2025/03/27 by Audrey Huang, Huang, Audrey, Adam Block +9 · 2 voices · 20 citations
#cs.AI #cs.LG #stat.ML
- Offline Reinforcement Learning with Realizability and Single-policy\n Concentrability
2022/02/09 by Wenhao Zhan, Baihe Huang, Zhan, Wenhao +7 · 7 citations
Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Mobile Crowdsensing and Crowdsourcing #Reinforcement Learning in Robotics
- Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
2024/07/18 by Audrey Huang, Wenhao Zhan, Huang, Audrey +11 · 1 voice · 9 citations
Computer Science · Engineering · #Advanced Multi-Objective Optimization Algorithms #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural Networks and Applications #Topology Optimization in Engineering #cs.AI #cs.CL #cs.LG
- Graph-Structured Visual Imitation
2019/07/11 by Maximilian Sieb, Sieb, Maximilian, Audrey Huang +6 · 4 citations
Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications #Robot Manipulation and Learning #Robotics (cs.RO)
- On the Convergence and Optimality of Policy Gradient for Markov Coherent Risk
2021/03/04 by Huang, Audrey, Leqi, Liu, Lipton, Zachary C. +1 · 3 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Computational-Statistical Tradeoffs at the Next-Token Prediction Barrier: Autoregressive and Imitation Learning under Misspecification
2025/02/18 by Dhruv Rohatgi, Rohatgi, Dhruv, Adam Block +7 · 1 voice · 7 citations
Computer Science · #Natural Language Processing Techniques
- Off-Policy Risk Assessment in Contextual Bandits
2021/04/18 by Audrey Huang, Huang, Audrey, Liu Leqi +5 · 2 citations
Decision Sciences · #Advanced Bandit Algorithms Research #Decision-Making and Behavioral Economics #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Risk and Portfolio Optimization
- Supervised Learning with General Risk Functionals
2022/06/27 by Liu Leqi, Audrey Huang, Leqi, Liu +5 · 2 citations
Computer Science · Mathematics · Medicine · #Domain Adaptation and Few-Shot Learning #Statistical Methods and Inference #Colorectal Cancer Screening and Detection
- Reinforcement Learning in Low-Rank MDPs with Density Features
2023/02/04 by Audrey Huang, Huang, Audrey, Jing‐Lin Chen +3 · 1 citation
Computer Science · Decision Sciences · Engineering · #Advanced Bandit Algorithms Research #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Reinforcement Learning in Robotics #Smart Grid Energy Management
- Model Selection for Off-policy Evaluation: New Algorithms and Experimental Protocol
2025/02/11 by Pai Liu, Lingfeng Zhao, Liu, Pai +11 · 3 citations
Computer Science · Mathematics · Decision Sciences · #Reinforcement Learning in Robotics #Advanced Causal Inference Techniques #Advanced Bandit Algorithms Research