Vanilla Policy Optimization Is Both Optimal and Differentially Private for Stochastic Contextual Bandits
Preprint,
I am a first-year PhD student at the University of Toronto, grateful to be advised by Murat Erdogdu, Wenlong Mou, and Stanislav Volgushev. My research interests are broadly in the mathematical foundations of machine learning, especially learning and decision-making with limited feedback, memory, or compute.
Before starting my PhD, I studied mathematics and computer science at the University of Toronto and did research at the Vector Institute, where I had the pleasure of working with Daniel Roy and Idan Attias.
New publications: Three papers to appear at NeurIPS 2026.
paperNew preprint: Vanilla Policy Optimization Is Both Optimal and Differentially Private for Stochastic Contextual Bandits.
paperStarted my PhD in Statistical Sciences at the University of Toronto.
lifePresented Reinforcement Learning with Action-Triggered Observations at ICML 2026.
paperFeatured in The Logic’s Top Prospects 2025.
lifePresented my first paper, Capacity-Constrained Online Learning with Delays, at COLT 2025.
paperGraduated from the University of Toronto and received the Governor General’s Silver Medal.
lifePreprint,
Conference on Neural Information Processing Systems (NeurIPS), (to appear)
Conference on Neural Information Processing Systems (NeurIPS), (to appear)
Conference on Neural Information Processing Systems (NeurIPS), (to appear)
International Conference on Machine Learning (ICML),
Conference on Learning Theory (COLT),