Resume-aware faculty matching

Find professors who actually fit you

Review faculty evidence in public, then use the workspace to turn your background into a shortlist, outreach, and meeting prep.

Profile-awarePaper evidenceSix agents
Donald B. Rubin

Donald B. Rubin

Harvard University · Biostatistics

Active 1967–2025

h-index151
Citations383.8k
Papers65748 last 5y
Funding$135k

Academic metrics are sourced from OpenAlex and public funding records; values may differ from Google Scholar.

See your match with Donald B. Rubin — sign in to PhdFit.Sign in

About

Donald B. Rubin is an Emeritus Professor of Statistics at Harvard University. His research interests include causal inference in experiments and observational studies, inference in sample surveys with nonresponse and in missing data problems, and the application of Bayesian and empirical Bayesian techniques. Rubin has developed and applied statistical models across various scientific disciplines. He earned his Ph.D. in Statistics from Harvard University in 1970, his M.A. in Computer Science from Harvard in 1966, and his A.B. in Psychology from Princeton University in 1965. Rubin served as a Professor in the Department of Statistics at Harvard from 1984 until his retirement in 2018, and he held the position of Chairman of the department during two separate periods. His professional experience also includes being a Fellow in Theoretical and Applied Statistics at the National Bureau of Economic Research and a Research Associate at NORC in Chicago.

Research topics

  • Computer Science
  • Political Science
  • Data Mining
  • Computer Security
  • Mathematical analysis
  • World Wide Web
  • Mathematics
  • Applied mathematics
  • Data science
  • Medicine

Selected publications

  • Automatic detection of influential actors in disinformation networks

    Proceedings of the National Academy of Sciences · 2021 · 61 citations

    Senior authorCorresponding

    The weaponization of digital communications and social media to conduct disinformation campaigns at immense scale, speed, and reach presents new challenges to identify and counter hostile influence operations (IOs). This paper presents an end-to-end framework to automate detection of disinformation narratives, networks, and influential actors. The framework integrates natural language processing, machine learning, graph analytics, and a network causal inference approach to quantify the impact of…

  • Estimating adjusted risk differences by multiply‐imputing missing control binary potential outcomes following propensity score‐matching

    Statistics in Medicine · 2021-08-10 · 3 citations

    articleOpen access

    We describe a new method to combine propensity-score matching with regression adjustment in treatment-control studies when outcomes are binary by multiply imputing potential outcomes under control for the matched treated subjects. This enables the estimation of clinically meaningful measures of effect such as the risk difference. We used Monte Carlo simulation to explore the effect of the number of imputed potential outcomes under control for the matched treated subjects on inferences about the…

  • PCA Rerandomization

    Canadian Journal of Statistics · 2023-02-16 · 2 citations

    preprintOpen accessSenior author

    Abstract Mahalanobis distance of covariate means between treatment and control groups is often adopted as a balance criterion when implementing a rerandomization strategy. However, this criterion may not work well for high‐dimensional cases because it balances all orthogonalized covariates equally. We propose using principal component analysis (PCA) to identify proper subspaces in which Mahalanobis distance should be calculated. Not only can PCA effectively reduce the dimensionality for high‐dim…

  • Catalytic Priors: Using Synthetic Data to Specify Prior Distributions in Bayesian Analysis

    arXiv (Cornell University) · 2022-08-30 · 2 citations

    preprintOpen access

    Catalytic prior distributions provide general, easy-to-use, and interpretable specifications of prior distributions for Bayesian analysis. They are particularly beneficial when the observed data are inadequate to stably estimate a complex target model. A catalytic prior distribution is constructed by augmenting the observed data with synthetic data that are sampled from the predictive distribution of a simpler model estimated from the observed data. We illustrate the usefulness of the catalytic…

  • Contrast-specific propensity scores

    Biostatistics & Epidemiology · 2021-01-02 · 1 citations

    preprintOpen accessSenior author

    Basic propensity score methodology is designed to balance the distributions of multivariate pre-treatment covariates when comparing one active treatment with one control treatment. However, practical settings often involve comparing more than two treatments, where more complicated contrasts than the basic treatment-control one, (1,−1), are relevant. Here, we propose the use of contrast-specific propensity scores (CSPS), which allows the creation of treatment groups of units that are balanced wit…

Recent grants

Frequent coauthors

Labs

Education

  • B.A., Mathematics

    Harvard University

    1963
  • M.A., Statistics

    Harvard University

    1965
  • Ph.D., Statistics

    Harvard University

    1968

Similar researchers at Harvard University

  • Resume-aware match score
  • Save to shortlist
  • AI-drafted outreach

See your match with Donald B. Rubin

PhdFit ranks faculty by your research interests, methods, and publications — grounded in their actual work, not templates.

  • Free to start
  • No credit card
  • 30-second signup