Resume-aware faculty matching

Find professors who actually fit you

Review faculty evidence in public, then use the workspace to turn your background into a shortlist, outreach, and meeting prep.

Profile-awarePaper evidenceSix agents
Lingzhou Xue

Lingzhou Xue

· Professor

Pennsylvania State University · Statistics

Active 2007–2026

h-index23
Citations2.7k
Papers12864 last 5y
Funding$1.7M1 active

Academic metrics are sourced from OpenAlex and public funding records; values may differ from Google Scholar.

See your match with Lingzhou Xue — sign in to PhdFit.Sign in

About

Lingzhou Xue is a Professor of Statistics at Penn State. He received his B.Sc. in Statistics from Peking University in 2008 and his Ph.D. in Statistics from the University of Minnesota in 2012. He was a postdoctoral research associate at Princeton University from 2012-2013. His research interests include high-dimensional statistics, nonparametric statistics, statistical and machine learning, large-scale optimization, and statistical modeling in biomedical, environmental, and social sciences. His recent research focuses on causal inference, federated learning, graphical models, high-dimensional inference, optimal transport, random objects, and reinforcement learning. He is a dedicated mentor to Ph.D. students and postdoctoral researchers, with five of his former advisees becoming tenure-track faculty members in statistics.

Research topics

  • Computer Science
  • Artificial Intelligence
  • Mathematics
  • Statistics
  • Biology
  • Algorithm
  • Geography
  • Waste management
  • Cartography
  • Computational biology

Selected publications

  • Compositional knockoff filter for high‐dimensional regression analysis of microbiome data

    Biometrics · 2020 · 33 citations

    A critical task in microbiome data analysis is to explore the association between a scalar response of interest and a large number of microbial taxa that are summarized as compositional data at different taxonomic levels. Motivated by fine-mapping of the microbiome, we propose a two-step compositional knockoff filter to provide the effective finite-sample false discovery rate (FDR) control in high-dimensional linear log-contrast regression analysis of microbiome compositional data. In the first…

  • An Alternating Manifold Proximal Gradient Method for Sparse Principal Component Analysis and Sparse Canonical Correlation Analysis

    INFORMS Journal on Optimization · 2020 · 28 citations

    Sparse principal component analysis and sparse canonical correlation analysis are two essential techniques from high-dimensional statistics and machine learning for analyzing large-scale data. Both problems can be formulated as an optimization problem with nonsmooth objective and nonconvex constraints. Because nonsmoothness and nonconvexity bring numerical difficulties, most algorithms suggested in the literature either solve some relaxations of them or are heuristic and lack convergence guarant…

  • Assessing Contamination of Stream Networks near Shale Gas Development Using a New Geospatial Tool

    Environmental Science & Technology · 2020 · 17 citations

    Chemical spills in streams can impact ecosystem or human health. Typically, the public learns of spills from reports from industry, media, or government rather than monitoring data. For example, ∼1300 spills (76 ≥ 400 gallons or ∼1500 L) were reported from 2007 to 2014 by the regulator for natural gas wellpads in the Marcellus shale region of Pennsylvania (U.S.), a region of extensive drilling and hydraulic fracturing. Only one such incident of stream contamination in Pennsylvania has been docum…

  • spVelo: RNA velocity inference for multi-batch spatial transcriptomics data

    Genome biology · 2025-08-11 · 4 citations

    articleOpen access

    RNA velocity has emerged as a powerful tool to interpret transcriptional dynamics and infer trajectory from snapshot datasets. However, current methods fail to utilize the spatial information inherent in spatial transcriptomics and lack scalability in multi-batch datasets. Here, we introduce spVelo, a scalable framework for RNA velocity inference of multi-batch spatial transcriptomics data. spVelo supports several downstream applications, including uncertainty quantification, complex trajectory…

  • Composition-on-composition regression analysis for multi-omics integration of metagenomic data

    Bioinformatics · 2025-07-01 · 1 citations

    articleOpen access

    MOTIVATION: Compositional data are frequently encountered in many disciplines, such as in next-generation sequencing experiments widely used in biomedical studies. Regression analysis with compositional data as either responses or predictors has been well studied. However, when both responses and predictors are compositional, the inventory of analysis tools is surprisingly limited, especially in the high-dimensional setting. Among the few existing methods, most of them rely on a log-ratio transf…

Recent grants

Frequent coauthors

Labs

  • Department of StatisticsPI

Education

  • Ph.D., Statistics

    University of Minnesota

    2012

Awards & honors

  • Penn State Schreyer Honors College (SHC) Excellence in Advis…
  • Institute of Mathematical Statistics (IMS) Fellow, 2024
  • Penn State Huck Institutes Leadership Fellow, 2024
  • American Statistical Association (ASA) Fellow, 2023
  • National Institute of Statistical Sciences (NISS) Distinguis…
  • Resume-aware match score
  • Save to shortlist
  • AI-drafted outreach

See your match with Lingzhou Xue

PhdFit ranks faculty by your research interests, methods, and publications — grounded in their actual work, not templates.

  • Free to start
  • No credit card
  • 30-second signup