
Zhaoran Wang
· Associate Professor of Industrial Engineering and Management Sciences and (by courtesy) Computer ScienceNorthwestern University · Chemical Engineering
Active 2010–2026
Academic metrics are sourced from OpenAlex and public funding records; values may differ from Google Scholar.
About
Zhaoran Wang is an Associate Professor of Industrial Engineering and Management Sciences at Northwestern University, with a courtesy appointment in Computer Science. His research focuses on developing a new generation of data-driven decision-making methods, theories, and systems that leverage artificial intelligence to address pressing societal challenges. His work aims to make autonomous learning agents more efficient both computationally and statistically, enabling their application in critical domains. Additionally, he is dedicated to scaling autonomous learning agents to design and optimize societal-scale multi-agent systems involving cooperation and competition among humans and robots. His research interests span across machine learning, optimization, statistics, game theory, and information theory. Wang has contributed to advancing the understanding and development of reinforcement learning, representation learning, and control systems, with a particular emphasis on provable efficiency and sample complexity. His work has been published in leading conferences such as ICML, NeurIPS, ICLR, and COLT, reflecting his active engagement in cutting-edge research in artificial intelligence and decision-making systems.
Research topics
- Computer Science
- Artificial Intelligence
- Machine Learning
- Information Retrieval
- Mathematics
- Combinatorics
- Mathematical optimization
- Statistics
- Discrete mathematics
- Engineering
Selected publications
Provably Efficient Exploration in Policy Optimization
2020 · 86 citations
Senior authorCorrespondingWhile policy-based reinforcement learning (RL) achieves tremendous successes in practice, it is significantly less understood in theory, especially compared with value-based RL. In particular, it remains elusive how to design a provably efficient policy optimization algorithm that incorporates exploration. To bridge such a gap, this paper proposes an Optimistic variant of the Proximal Policy Optimization algorithm (OPPO), which follows an ``optimistic version'' of the policy gradient direction.…
Offline Reinforcement Learning for Human-Guided Human-Machine Interaction with Private Information
Management Science · 2025-08-06 · 2 citations
articleMotivated by the human-machine interaction such as recommending videos for improving customer engagement, we study human-guided human-machine interaction for decision making with private information. We model this interaction as a two-player turn-based game, where one player (Bob, a human) guides the other player (Alice, a machine) toward a common goal. Specifically, we focus on offline reinforcement learning (RL) in this game, where the goal is to find a policy pair for Alice and Bob that maxim…
Digital Evocation in the Posthuman: The Hauntological Production of Video Games as Cultural Industry
GBP Proceedings Series · 2026-05-02
article1st authorCorrespondingSince the emergence of the posthuman condition, digital technologies and artificial intelligence have profoundly restructured contemporary cultural industries. These developments have not only transformed modes of production but also engendered a novel production logic predicated on the permanence of data storage and the virtualization of material production. Within this context, video games, by virtue of their intrinsic capacity for data persistence, algorithmic modulation, and material virtual…
Learning to Reason as Action Abstractions with Scalable Mid-Training RL
ArXiv.org · 2025-09-30
preprintOpen accessLarge language models excel with reinforcement learning (RL), but fully unlocking this potential requires a mid-training stage. An effective mid-training phase should identify a compact set of useful actions and enable fast selection among them through online RL. We formalize this intuition by presenting the first theoretical result on how mid-training shapes post-training: it characterizes an action subspace that minimizes both the value approximation error from pruning and the RL error during…
Risk-Sensitive Deep RL: Variance-Constrained Actor-Critic Provably Finds Globally Optimal Policy
Journal of the American Statistical Association · 2025-11-12
article
Frequent coauthors
- 125 shared
Zhuoran Yang
- 39 shared
David M. Blei
- 31 shared
Daniel Shu Wei Ting
Duke-NUS Medical School
- 28 shared
Tien Yin Wong
Singapore National Eye Center
- 22 shared
Gilbert Lim
Singapore Eye Research Institute
- 22 shared
Michael I. Jordan
- 21 shared
Lingxiao Wang
Beijing University of Chinese Medicine
- 16 shared
Huiguang He
Institute of Automation
Similar researchers at Northwestern University
- Resume-aware match score
- Save to shortlist
- AI-drafted outreach
See your match with Zhaoran Wang
PhdFit ranks faculty by your research interests, methods, and publications — grounded in their actual work, not templates.
- Free to start
- No credit card
- 30-second signup
