Resume-aware faculty matching

Find professors who actually fit you

Review faculty evidence in public, then use the workspace to turn your background into a shortlist, outreach, and meeting prep.

Profile-awarePaper evidenceSix agents
Guang Cheng

Guang Cheng

· Professor

University of California, Los Angeles · Computer Science

Active 1994–2025

h-index27
Citations2.8k
Papers264108 last 5y
Funding$686k

Academic metrics are sourced from OpenAlex and public funding records; values may differ from Google Scholar.

See your match with Guang Cheng — sign in to PhdFit.Sign in

About

Guang Cheng is a professor in the Department of Computer Science at UCLA Samueli School of Engineering. His research interests include generative data science, machine/deep learning algorithms and theory, and the statistical foundations of data science. He holds a PhD in Statistics from the University of Wisconsin, Madison, obtained in 2006, and a BA in Economics and Management from Tsinghua University, earned in 2002. Cheng has received numerous awards and recognitions, including being named an IMS Fellow in 2020, receiving the Adobe Data Science Faculty Award in 2020, being designated a University Faculty Scholar in 2018, awarded the Simons Fellowship in Mathematics in 2014, the Noether Young Scholar Award in 2012, the NSF CAREER Award in 2012, and the Facebook X Instagram award. His work contributes to advancing the theoretical and practical understanding of data science and machine learning.

Research topics

  • Computer Science
  • Artificial Intelligence
  • Political Science
  • Data Mining
  • Computer Security
  • Machine Learning
  • Data science
  • Algorithm
  • Mathematics

Selected publications

  • Benefit of Interpolation in Nearest Neighbor Algorithms

    SIAM Journal on Mathematics of Data Science · 2022 · 39 citations

    Senior authorCorresponding

    In some studies (e.g., [C. Zhang et al. in Proceedings of the 5th International Conference on Learning Representations, OpenReview.net, 2017]) of deep learning, it is observed that overparametrized deep neural networks achieve a small testing error even when the training error is almost zero. Despite numerous works toward understanding this so-called double-descent phenomenon (e.g., [M. Belkin et al., Proc. Natl. Acad. Sci. USA, 116 (2019), pp. 15849--15854; M. Belkin, D. Hsu, and J. Xu, SIAM J.…

  • A Survey on Statistical Theory of Deep Learning: Approximation, Training Dynamics, and Generative Models

    Annual Review of Statistics and Its Application · 2024-11-21 · 5 citations

    articleOpen accessSenior author

    In this article, we review the literature on statistical theories of neural networks from three perspectives: approximation, training dynamics, and generative models. In the first part, results on excess risks for neural networks are reviewed in the nonparametric framework of regression. These results rely on explicit constructions of neural networks, leading to fast convergence rates of excess risks. Nonetheless, their underlying analysis only applies to the global minimizer in the highly nonco…

  • Data Plagiarism Index: Characterizing the Privacy Risk of Data-Copying in Tabular Generative Models

    arXiv (Cornell University) · 2024-06-18 · 2 citations

    preprintOpen accessSenior author

    The promise of tabular generative models is to produce realistic synthetic data that can be shared and safely used without dangerous leakage of information from the training set. In evaluating these models, a variety of methods have been proposed to measure the tendency to copy data from the training dataset when generating a sample. However, these methods suffer from either not considering data-copying from a privacy threat perspective, not being motivated by recent results in the data-copying…

  • GReaTER: Generate Realistic Tabular data after data Enhancement and Reduction

    2025-05-19 · 1 citations

    articleSenior author

    Tabular data synthesis involves not only multi-table synthesis but also generating multi-modal data (e.g., strings and categories), which enables diverse knowledge synthesis. However, separating numerical and categorical data has limited the effectiveness of tabular data generation. The GReaT (Generate Realistic Tabular Data) framework uses Large Language Models (LLMs) to encode entire rows, eliminating the need to partition data types. Despite this, the framework's performance is constrained by…

  • Rate-Optimal Rank Aggregation with Private Pairwise Rankings

    Journal of the American Statistical Association · 2025-04-03 · 1 citations

    articleSenior author

Recent grants

Frequent coauthors

  • Zuofeng Shang

    49 shared
  • Shilian Kan

    Tianjin Hospital

    25 shared
  • Shuangle Zong

    Second Hospital of Tangshan

    25 shared
  • Weidong Liang

    First Affiliated Hospital of Gannan Medical University

    25 shared
  • Lidong Li

    University of Science and Technology Beijing

    25 shared
  • Ligeng Li

    Second Hospital of Tangshan

    25 shared
  • Aijun Wang

    Qilu Hospital of Shandong University

    25 shared
  • Qiutao Zheng

    25 shared

Awards & honors

  • IMS Fellow (2020)
  • Adobe Data Science Faculty Award (2020)
  • University Faculty Scholar (2018)
  • Simons Fellowship in Mathematics (2014)
  • Noether Young Scholar Award (2012)

Similar researchers at University of California, Los Angeles

  • Resume-aware match score
  • Save to shortlist
  • AI-drafted outreach

See your match with Guang Cheng

PhdFit ranks faculty by your research interests, methods, and publications — grounded in their actual work, not templates.

  • Free to start
  • No credit card
  • 30-second signup