Publications and Preprints
★ Equal contribution
2026
On Repulsive and Attractive Teachers: Separating Correctness from Behavior in Self-Distillation
Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models
Aligning Language Models from User Interactions
MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference
Reinforcement Learning via Self-Distillation
2025
Stackelberg Learning from Human Feedback: Preference Optimization as a Sequential Game
Causal Imitation Learning under Expert-Observable and Expert-Unobservable Confounding
Strategyproof Reinforcement Learning from Human Feedback
Data Source Adaptive Online Learning under Heteroscedastic Noise
A Minimax Approach to Ad Hoc Teamwork
