
Benjamin Van Durme
· Joint Appointment; Associate Professor, Whiting School of Engineering, Computer ScienceJohns Hopkins University · Neuroscience
Active 2003–2026
Academic metrics are sourced from OpenAlex and public funding records; values may differ from Google Scholar.
About
Benjamin Van Durme is an Associate Professor in Computer Science and Cognitive Science at Johns Hopkins University. He is a member of the Center for Language and Speech Processing (CLSP) and leads Natural Language Understanding research at the Human Language Technology Center of Excellence (HLTCOE). His research focuses on helping people work with large amounts of information by understanding the content of documents and images, assisting in information retrieval, and enabling systems to answer questions about that content. Van Durme collaborates on research topics including natural language processing, data mining, social media analysis, machine learning, linguistic semantics, and broader areas within Artificial Intelligence and Cognitive Science. His work in decompositional semantics is organized through Decomp.io. Additionally, he serves as the research lead at Microsoft Semantic Machines.
Research topics
- Computer Science
- Artificial Intelligence
- Natural Language Processing
- Information Retrieval
- Philosophy
- History
- Art history
- Linguistics
- Geology
- Epistemology
Selected publications
Complement Lexical Retrieval Model with Semantic Residual Embeddings
Lecture notes in computer science · 2021 · 87 citations
LOME: Large Ontology Multilingual Extraction
2021 · 24 citations
Senior authorCorrespondingPatrick Xia, Guanghui Qin, Siddharth Vashishtha, Yunmo Chen, Tongfei Chen, Chandler May, Craig Harman, Kyle Rawlins, Aaron Steven White, Benjamin Van Durme. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations. 2021.
Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval
2025-06-10 · 4 citations
articleIn this work, we tackle the problem of text-to-video retrieval (T2VR). Inspired by the success of late interaction techniques in text-document, text-image, and text-video retrieval, our approach, Video-ColBERT, introduces a simple and efficient mechanism for fine-grained similarity assessment between queries and videos. Video-ColBERT is built upon three main components: a fine-grained spatial and temporal token-wise interaction, query and visual expansions, and a dual sigmoid loss during trainin…
mmBERT: A Modern Multilingual Encoder with Annealed Language Learning
ArXiv.org · 2025-09-08 · 2 citations
preprintOpen accessSenior authorEncoder-only languages models are frequently used for a variety of standard machine learning tasks, including classification and retrieval. However, there has been a lack of recent research for encoder models, especially with respect to multilingual models. We introduce mmBERT, an encoder-only language model pretrained on 3T tokens of multilingual text in over 1800 languages. To build mmBERT we introduce several novel elements, including an inverse mask ratio schedule and an inverse temperature…
RE-AdaptIR: Improving Information Retrieval through Reverse Engineered Adaptation
2025-07-13 · 1 citations
articleOpen accessSenior authorLarge language models (LLMs) fine-tuned for text-retrieval have demonstrated state-of-the-art results across several information retrieval (IR) benchmarks. However, supervised training for improving these models requires numerous labeled examples, which are generally unavailable or expensive to acquire. In this work, we explore the effectiveness of extending reverse engineered adaptation to the context of information retrieval (RE-AdaptIR). We use RE-AdaptIR to improve LLM-based IR models using…
Recent grants
Computational Statutory Reasoning
NSF · $597k · 2022–2025
NSF · $124k · 2018–2023
Frequent coauthors
- 72 shared
Adam Poliak
- 67 shared
Patrick Xia
- 66 shared
Catherine Havasi
- 65 shared
Felipe Meneguzzi
- 65 shared
Antoine Raux
Honda (United States)
- 65 shared
Gita Sukthankar
- 65 shared
William F. Lawless
Paine College
- 65 shared
Mirsad Hadžikadić
University of North Carolina at Charlotte
Education
- 2008
Ph.D., Computer Science
University of California, Berkeley
- 2003
M.S., Computer Science
University of California, Berkeley
- 2001
B.S., Computer Science
University of California, Berkeley
Awards & honors
- Celebrating Women in Data Science and AI Symposium recogniti…
- Amazon AI fellowship program for Johns Hopkins students
Similar researchers at Johns Hopkins University
- Resume-aware match score
- Save to shortlist
- AI-drafted outreach
See your match with Benjamin Van Durme
PhdFit ranks faculty by your research interests, methods, and publications — grounded in their actual work, not templates.
- Free to start
- No credit card
- 30-second signup
