Ali Farhadi
· ProfessorUniversity of Washington · Computer Science & Engineering
Active 2002–2026
Academic metrics are sourced from OpenAlex and public funding records; values may differ from Google Scholar.
About
Ali Farhadi is a Professor in the Department of Computer Science & Engineering at the University of Washington and also serves as CEO of the Allen Institute for Artificial Intelligence (AI2). Before joining the UW faculty, Farhadi spent a year as a postdoctoral fellow at the Robotics Institute at Carnegie Mellon University working with Martial Hebert and Alyosha Efros. He earned his Ph.D. from the University of Illinois at Urbana-Champaign under the supervision of David Forsyth, during which he also worked closely with Derek Hoiem. His main research interests include computer vision, machine learning, the intersection of natural language and vision, analysis of the role of semantics in visual understanding, and visual reasoning.
Research topics
- Computer Science
- Artificial Intelligence
- Natural Language Processing
- Information Retrieval
- Machine Learning
- Computer vision
- Programming language
- Data science
- Engineering
- Algorithm
Selected publications
Objaverse: A Universe of Annotated 3D Objects
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) · 2023 · 562 citations
Senior authorCorrespondingMassive data corpora like WebText, Wikipedia, Conceptual Captions, WebImageText, and LAION have propelled recent dramatic progress in AI. Large neural models trained on such datasets produce impressive results and top many of today's benchmarks. A notable omisslion within this family of large-scale datasets is 3D data. Despite considerable interest and potential applications in 3D vision, datasets of high-fidelity 3D models continue to be mid-sized with limited diversity of object categories. Ad…
Robust fine-tuning of zero-shot models
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) · 2022 · 367 citations
Large pre-trained models such as CLIP or ALIGN offer consistent accuracy across a range of data distributions when performing zero-shot inference (i.e., without fine-tuning on a specific dataset). Although existing fine-tuning methods substantially improve accuracy on a given target distribution, they often reduce robustness to distribution shifts. We address this tension by introducing a simple and effective method for improving robustness while fine-tuning: ensembling the weights of the zero-s…
Fine-Tuning Pretrained Language Models: Weight Initializations, Data\n Orders, and Early Stopping
arXiv (Cornell University) · 2020 · 263 citations
Fine-tuning pretrained contextual word embedding models to supervised\ndownstream tasks has become commonplace in natural language processing. This\nprocess, however, is often brittle: even with the same hyperparameter values,\ndistinct random seeds can lead to substantially different results. To better\nunderstand this phenomenon, we experiment with four datasets from the GLUE\nbenchmark, fine-tuning BERT hundreds of times on each while varying only the\nrandom seeds. We find substantial perfor…
LanguageRefer: Spatial-Language Model for 3D Visual Grounding
arXiv (Cornell University) · 2021 · 22 citations
For robots to understand human instructions and perform meaningful tasks in the near future, it is important to develop learned models that comprehend referential language to identify common objects in real-world 3D scenes. In this paper, we introduce a spatial-language model for a 3D visual grounding problem. Specifically, given a reconstructed 3D scene in the form of point clouds with 3D bounding boxes of potential object candidates, and a language utterance referring to a target object in the…
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
arXiv (Cornell University) · 2024-09-25 · 8 citations
preprintOpen accessToday's most advanced vision-language models (VLMs) remain proprietary. The strongest open-weight models rely heavily on synthetic data from proprietary VLMs to achieve good performance, effectively distilling these closed VLMs into open ones. As a result, the community has been missing foundational knowledge about how to build performant VLMs from scratch. We present Molmo, a new family of VLMs that are state-of-the-art in their class of openness. Our key contribution is a collection of new dat…
Recent grants
CAREER: Computation and Approximation in Structured Learning
NSF · $420k · 2013–2017
RI: Small: Collaborative Research: Detecting Abnormalities in Images
NSF · $120k · 2013–2016
CAREER: Active and Action-Centric Visual Understanding
NSF · $550k · 2017–2023
Frequent coauthors
- 105 shared
Aniruddha Kembhavi
Allen Institute
- 89 shared
Mohammad Rastegari
- 82 shared
Roozbeh Mottaghi
Seattle University
- 68 shared
Hannaneh Hajishirzi
- 49 shared
Yejin Choi
- 43 shared
Abhinav Gupta
- 41 shared
Santosh Divvala
Allen Institute
- 40 shared
Luca Weihs
Allen Institute
Education
Ph.D.
University of Illinois at Urbana-Champaign
Other
Robotics Institute at Carnegie Mellon University
Similar researchers at University of Washington
- Resume-aware match score
- Save to shortlist
- AI-drafted outreach
See your match with Ali Farhadi
PhdFit ranks faculty by your research interests, methods, and publications — grounded in their actual work, not templates.
- Free to start
- No credit card
- 30-second signup
