
Yejin Choi
Stanford University · Learning, Design, and Technology
Active 2003–2025
Academic metrics are sourced from OpenAlex and public funding records; values may differ from Google Scholar.
About
Yejin Choi is a leading figure in AI research, focusing on AI for Science, Pluralistic Alignment & AI for humanity, and Alternative Training and Inference Algorithms. She works on molecular foundation models, protein reasoning, and molecular reasoning and retrosynthesis. She also explores data and algorithms for pluralistic norms and values, deliberate alignment processes, and civic discourse.
Research topics
- Artificial Intelligence
- Computer Science
- Natural Language Processing
- Data Mining
- Information Retrieval
- Programming language
- Computer vision
- Chemistry
- Linguistics
- Chromatography
Selected publications
Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
arXiv (Cornell University) · 2022 · 548 citations
Language models demonstrate both quantitative improvement and new qualitative capabilities with increasing scale. Despite their potentially transformative impact, these new capabilities are as yet poorly characterized. In order to inform future research, prepare for disruptive new model capabilities, and ameliorate socially harmful effects, it is vital that we understand the present and near-future capabilities and limitations of language models. To address this challenge, we introduce the Beyon…
The Curious Case of Neural Text Degeneration
arXiv (Cornell University) · 2020 · 528 citations
Senior authorCorrespondingDespite considerable advances in neural language modeling, it remains an open question what the best decoding strategy is for text generation from a language model (e.g. to generate a story). The counter-intuitive empirical observation is that even though the use of likelihood as training objective leads to high quality models for a broad range of language understanding tasks, maximization-based decoding methods such as beam search lead to degeneration — output text that is bland, incoherent, or…
MAUVE: Measuring the Gap Between Neural Text and Human Text using\n Divergence Frontiers
arXiv (Cornell University) · 2021 · 91 citations
As major progress is made in open-ended text generation, measuring how close\nmachine-generated text is to human language remains a critical open problem. We\nintroduce MAUVE, a comparison measure for open-ended text generation, which\ndirectly compares the learnt distribution from a text generation model to the\ndistribution of human-written text using divergence frontiers. MAUVE scales up\nto modern text generation models by computing information divergences in a\nquantized embedding space. Th…
HALoGEN: Fantastic LLM Hallucinations and Where to Find Them
2025-01-01 · 5 citations
articleOpen accessSenior authorDespite their impressive ability to generate high-quality and fluent text, generative large language models (LLMs) also produce hallucinations: statements that are misaligned with established world knowledge or provided input context.However, measuring hallucination can be challenging, as having humans verify model generations on-the-fly is both expensive and time-consuming.In this work, we release HALOGEN , a comprehensive hallucination benchmark consisting of: (1) 10,923 prompts for generative…
HALoGEN: Fantastic LLM Hallucinations and Where to Find Them
ArXiv.org · 2025-01-14 · 3 citations
preprintOpen accessSenior authorDespite their impressive ability to generate high-quality and fluent text, generative large language models (LLMs) also produce hallucinations: statements that are misaligned with established world knowledge or provided input context. However, measuring hallucination can be challenging, as having humans verify model generations on-the-fly is both expensive and time-consuming. In this work, we release HALoGEN, a comprehensive hallucination benchmark consisting of: (1) 10,923 prompts for generativ…
Recent grants
RI: Small: ConnotationNet: Modeling Non-Literal Meaning in Context
NSF · $500k · 2017–2021
RI: Small: A Data-Driven Framework to Sketch-to-Text Generation
NSF · $450k · 2015–2019
Frequent coauthors
- 186 shared
Ronan Le Bras
- 153 shared
Maarten Sap
- 137 shared
Noah A. Smith
- 126 shared
Chandra Bhagavatula
Allen Institute
- 117 shared
Ximing Lu
- 95 shared
Swabha Swayamdipta
- 86 shared
Antoine Bosselut
- 86 shared
Jack Hessel
Awards & honors
- MacArthur Fellow (class of 2022)
- ACL Fellow (2022)
- Brett Helsel Career Development Professorship (2020 - 2023)
- Borg Early Career Award (BECA) (2018)
- IEEE AI's 10 to Watch (2016)
Similar researchers at Stanford University
- Resume-aware match score
- Save to shortlist
- AI-drafted outreach
See your match with Yejin Choi
PhdFit ranks faculty by your research interests, methods, and publications — grounded in their actual work, not templates.
- Free to start
- No credit card
- 30-second signup
