Resume-aware faculty matching

Find professors who actually fit you

Review faculty evidence in public, then use the workspace to turn your background into a shortlist, outreach, and meeting prep.

Profile-awarePaper evidenceSix agents
Simon Shaolei Du

Simon Shaolei Du

· Assistant Professor

University of Washington · Computer Science & Engineering

Active 2013–2026

h-index37
Citations5.4k
Papers230146 last 5y
Funding

Academic metrics are sourced from OpenAlex and public funding records; values may differ from Google Scholar.

See your match with Simon Shaolei Du — sign in to PhdFit.Sign in

About

Simon Shaolei Du is an assistant professor in the Paul G. Allen School of Computer Science & Engineering at the University of Washington. His research interests are broadly in machine learning, including deep learning, representation learning, reinforcement learning, and data selection. Prior to his faculty position, Du was a postdoctoral researcher at the Institute for Advanced Study of Princeton, hosted by Sanjeev Arora. He completed his Ph.D. in Machine Learning at Carnegie Mellon University, where he was co-advised by Aarti Singh and Barnabás Póczos. His academic background also includes studies in EECS and EMS at UC Berkeley, and he has spent time at the Simons Institute as well as research labs of Meta, Google, and Microsoft.

Research topics

  • Computer Science
  • Machine Learning
  • Artificial Intelligence
  • Mathematics
  • Applied mathematics
  • Geometry
  • Mathematical optimization
  • Mathematical analysis
  • Theoretical computer science

Selected publications

  • Understanding the acceleration phenomenon via high-resolution differential equations

    Mathematical Programming · 2021 · 124 citations

    Abstract Gradient-based optimization algorithms can be studied from the perspective of limiting ordinary differential equations (ODEs). Motivated by the fact that existing ODEs do not distinguish between two fundamentally different algorithms—Nesterov’s accelerated gradient method for strongly convex functions (NAG-) and Polyak’s heavy-ball method—we study an alternative limiting process that yields high-resolution ODEs . We show that these ODEs permit a general Lyapunov function framework for t…

  • How Neural Networks Extrapolate: From Feedforward to Graph Neural Networks

    arXiv (Cornell University) · 2020 · 110 citations

    We study how neural networks trained by gradient descent extrapolate, i.e., what they learn outside the support of the training distribution. Previous works report mixed empirical results when extrapolating with neural networks: while feedforward neural networks, a.k.a. multilayer perceptrons (MLPs), do not extrapolate well in certain simple tasks, Graph Neural Networks (GNNs) -- structured networks with MLP modules -- have shown some success in more complex tasks. Working towards a theoretical…

  • Settling the Sample Complexity of Online Reinforcement Learning

    Journal of the ACM · 2025-05-02 · 2 citations

    articleOpen accessSenior author

    A central issue lying at the heart of online reinforcement learning (RL) is data efficiency. While a number of recent works achieved asymptotically minimal regret in online RL, the optimality of these results is only guaranteed in a “large-sample” regime, imposing enormous burn-in cost in order for their algorithms to operate optimally. How to achieve minimax-optimal regret without incurring any burn-in cost has been an open problem in RL theory. We settle this problem for finite-horizon inhomog…

  • Reinforcement Learning for Reasoning in Large Language Models with One Training Example

    ArXiv.org · 2025-04-29 · 1 citations

    preprintOpen access

    We show that reinforcement learning with verifiable reward using one training example (1-shot RLVR) is effective in incentivizing the math reasoning capabilities of large language models (LLMs). Applying RLVR to the base model Qwen2.5-Math-1.5B, we identify a single example that elevates model performance on MATH500 from 36.0% to 73.6% (8.6% improvement beyond format correction), and improves the average performance across six common mathematical reasoning benchmarks from 17.6% to 35.7% (7.0% no…

  • Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation

    2025-06-10 · 1 citations

    article

    The current state-of-the-art video generative models can produce commercial-grade videos with highly realistic details. However, they still struggle to coherently present multiple sequential events in the stories specified by the prompts, which is foreseeable an essential capability for future long video generation scenarios. For example, top T2V generative models still fail to generate a video of the short simple story "how to put an elephant into a refrigerator." While existing detail-oriented…

Frequent coauthors

Education

  • Ph.D., Machine Learning

    Carnegie Mellon University

  • B.S., EECS and EMS

    UC Berkeley

Similar researchers at University of Washington

  • Resume-aware match score
  • Save to shortlist
  • AI-drafted outreach

See your match with Simon Shaolei Du

PhdFit ranks faculty by your research interests, methods, and publications — grounded in their actual work, not templates.

  • Free to start
  • No credit card
  • 30-second signup