Simon Shaolei Du
· Assistant ProfessorUniversity of Washington · Computer Science & Engineering
Active 2013–2026
Academic metrics are sourced from OpenAlex and public funding records; values may differ from Google Scholar.
About
Simon Shaolei Du is an assistant professor in the Paul G. Allen School of Computer Science & Engineering at the University of Washington. His research interests are broadly in machine learning, including deep learning, representation learning, reinforcement learning, and data selection. Prior to his faculty position, Du was a postdoctoral researcher at the Institute for Advanced Study of Princeton, hosted by Sanjeev Arora. He completed his Ph.D. in Machine Learning at Carnegie Mellon University, where he was co-advised by Aarti Singh and Barnabás Póczos. His academic background also includes studies in EECS and EMS at UC Berkeley, and he has spent time at the Simons Institute as well as research labs of Meta, Google, and Microsoft.
Research topics
- Computer Science
- Machine Learning
- Artificial Intelligence
- Mathematics
- Applied mathematics
- Geometry
- Mathematical optimization
- Mathematical analysis
- Theoretical computer science
Selected publications
Understanding the acceleration phenomenon via high-resolution differential equations
Mathematical Programming · 2021 · 124 citations
Abstract Gradient-based optimization algorithms can be studied from the perspective of limiting ordinary differential equations (ODEs). Motivated by the fact that existing ODEs do not distinguish between two fundamentally different algorithms—Nesterov’s accelerated gradient method for strongly convex functions (NAG-) and Polyak’s heavy-ball method—we study an alternative limiting process that yields high-resolution ODEs . We show that these ODEs permit a general Lyapunov function framework for t…
How Neural Networks Extrapolate: From Feedforward to Graph Neural Networks
arXiv (Cornell University) · 2020 · 110 citations
We study how neural networks trained by gradient descent extrapolate, i.e., what they learn outside the support of the training distribution. Previous works report mixed empirical results when extrapolating with neural networks: while feedforward neural networks, a.k.a. multilayer perceptrons (MLPs), do not extrapolate well in certain simple tasks, Graph Neural Networks (GNNs) -- structured networks with MLP modules -- have shown some success in more complex tasks. Working towards a theoretical…
Settling the Sample Complexity of Online Reinforcement Learning
Journal of the ACM · 2025-05-02 · 2 citations
articleOpen accessSenior authorA central issue lying at the heart of online reinforcement learning (RL) is data efficiency. While a number of recent works achieved asymptotically minimal regret in online RL, the optimality of these results is only guaranteed in a “large-sample” regime, imposing enormous burn-in cost in order for their algorithms to operate optimally. How to achieve minimax-optimal regret without incurring any burn-in cost has been an open problem in RL theory. We settle this problem for finite-horizon inhomog…
Reinforcement Learning for Reasoning in Large Language Models with One Training Example
ArXiv.org · 2025-04-29 · 1 citations
preprintOpen accessWe show that reinforcement learning with verifiable reward using one training example (1-shot RLVR) is effective in incentivizing the math reasoning capabilities of large language models (LLMs). Applying RLVR to the base model Qwen2.5-Math-1.5B, we identify a single example that elevates model performance on MATH500 from 36.0% to 73.6% (8.6% improvement beyond format correction), and improves the average performance across six common mathematical reasoning benchmarks from 17.6% to 35.7% (7.0% no…
2025-06-10 · 1 citations
articleThe current state-of-the-art video generative models can produce commercial-grade videos with highly realistic details. However, they still struggle to coherently present multiple sequential events in the stories specified by the prompts, which is foreseeable an essential capability for future long video generation scenarios. For example, top T2V generative models still fail to generate a video of the short simple story "how to put an elephant into a refrigerator." While existing detail-oriented…
Frequent coauthors
- 38 shared
Jason D. Lee
- 33 shared
Ruosong Wang
Beijing Institute of Technology
- 17 shared
Barnabás Póczos
- 17 shared
Yining Wang
The University of Texas at Dallas
- 16 shared
Lin F. Yang
- 16 shared
Kevin Jamieson
- 13 shared
Qiwen Cui
- 12 shared
Sham M. Kakade
Education
Ph.D., Machine Learning
Carnegie Mellon University
B.S., EECS and EMS
UC Berkeley
Similar researchers at University of Washington
- Resume-aware match score
- Save to shortlist
- AI-drafted outreach
See your match with Simon Shaolei Du
PhdFit ranks faculty by your research interests, methods, and publications — grounded in their actual work, not templates.
- Free to start
- No credit card
- 30-second signup
