Resume-aware faculty matching

Find professors who actually fit you

Review faculty evidence in public, then use the workspace to turn your background into a shortlist, outreach, and meeting prep.

Profile-awarePaper evidenceSix agents
Yejin Choi

Yejin Choi

Stanford University · Learning, Design, and Technology

Active 2003–2025

h-index99
Citations42.7k
Papers703479 last 5y
Funding$950k

Academic metrics are sourced from OpenAlex and public funding records; values may differ from Google Scholar.

See your match with Yejin Choi — sign in to PhdFit.Sign in

About

Yejin Choi is a leading figure in AI research, focusing on AI for Science, Pluralistic Alignment & AI for humanity, and Alternative Training and Inference Algorithms. She works on molecular foundation models, protein reasoning, and molecular reasoning and retrosynthesis. She also explores data and algorithms for pluralistic norms and values, deliberate alignment processes, and civic discourse.

Research topics

  • Artificial Intelligence
  • Computer Science
  • Natural Language Processing
  • Data Mining
  • Information Retrieval
  • Programming language
  • Computer vision
  • Chemistry
  • Linguistics
  • Chromatography

Selected publications

  • Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

    arXiv (Cornell University) · 2022 · 548 citations

    Language models demonstrate both quantitative improvement and new qualitative capabilities with increasing scale. Despite their potentially transformative impact, these new capabilities are as yet poorly characterized. In order to inform future research, prepare for disruptive new model capabilities, and ameliorate socially harmful effects, it is vital that we understand the present and near-future capabilities and limitations of language models. To address this challenge, we introduce the Beyon…

  • The Curious Case of Neural Text Degeneration

    arXiv (Cornell University) · 2020 · 528 citations

    Senior authorCorresponding

    Despite considerable advances in neural language modeling, it remains an open question what the best decoding strategy is for text generation from a language model (e.g. to generate a story). The counter-intuitive empirical observation is that even though the use of likelihood as training objective leads to high quality models for a broad range of language understanding tasks, maximization-based decoding methods such as beam search lead to degeneration — output text that is bland, incoherent, or…

  • MAUVE: Measuring the Gap Between Neural Text and Human Text using\n Divergence Frontiers

    arXiv (Cornell University) · 2021 · 91 citations

    As major progress is made in open-ended text generation, measuring how close\nmachine-generated text is to human language remains a critical open problem. We\nintroduce MAUVE, a comparison measure for open-ended text generation, which\ndirectly compares the learnt distribution from a text generation model to the\ndistribution of human-written text using divergence frontiers. MAUVE scales up\nto modern text generation models by computing information divergences in a\nquantized embedding space. Th…

  • HALoGEN: Fantastic LLM Hallucinations and Where to Find Them

    2025-01-01 · 5 citations

    articleOpen accessSenior author

    Despite their impressive ability to generate high-quality and fluent text, generative large language models (LLMs) also produce hallucinations: statements that are misaligned with established world knowledge or provided input context.However, measuring hallucination can be challenging, as having humans verify model generations on-the-fly is both expensive and time-consuming.In this work, we release HALOGEN , a comprehensive hallucination benchmark consisting of: (1) 10,923 prompts for generative…

  • HALoGEN: Fantastic LLM Hallucinations and Where to Find Them

    ArXiv.org · 2025-01-14 · 3 citations

    preprintOpen accessSenior author

    Despite their impressive ability to generate high-quality and fluent text, generative large language models (LLMs) also produce hallucinations: statements that are misaligned with established world knowledge or provided input context. However, measuring hallucination can be challenging, as having humans verify model generations on-the-fly is both expensive and time-consuming. In this work, we release HALoGEN, a comprehensive hallucination benchmark consisting of: (1) 10,923 prompts for generativ…

Recent grants

Frequent coauthors

  • Ronan Le Bras

    186 shared
  • Maarten Sap

    153 shared
  • Noah A. Smith

    137 shared
  • Chandra Bhagavatula

    Allen Institute

    126 shared
  • Ximing Lu

    117 shared
  • Swabha Swayamdipta

    95 shared
  • Antoine Bosselut

    86 shared
  • Jack Hessel

    86 shared

Awards & honors

  • MacArthur Fellow (class of 2022)
  • ACL Fellow (2022)
  • Brett Helsel Career Development Professorship (2020 - 2023)
  • Borg Early Career Award (BECA) (2018)
  • IEEE AI's 10 to Watch (2016)

Similar researchers at Stanford University

  • Resume-aware match score
  • Save to shortlist
  • AI-drafted outreach

See your match with Yejin Choi

PhdFit ranks faculty by your research interests, methods, and publications — grounded in their actual work, not templates.

  • Free to start
  • No credit card
  • 30-second signup