Resume-aware faculty matching

Find professors who actually fit you

Review faculty evidence in public, then use the workspace to turn your background into a shortlist, outreach, and meeting prep.

Profile-awarePaper evidenceSix agents
Arjun Guha

Arjun Guha

Northeastern University · Software Engineering

Active 2005–2026

h-index33
Citations3.9k
Papers12852 last 5y
Funding$1.9M

Academic metrics are sourced from OpenAlex and public funding records; values may differ from Google Scholar.

See your match with Arjun Guha — sign in to PhdFit.Sign in

About

Arjun Guha is an associate professor in the Khoury College of Computer Sciences at Northeastern University, based in Boston. His research focuses on programming languages, with particular interest in security and reliability problems in web programming, systems, and robotics. Guha uses tools and techniques from programming languages to address these issues, and one of his recent projects aims to make serverless computing more cost-effective, reliable, and applicable. He is a member of the Programming Research Laboratory. Prior to joining Northeastern, Guha was an associate professor at the University of Massachusetts Amherst and a postdoctoral research associate at Cornell University. His work has received several awards, including an OOPSLA Most Influential Paper Award, a PLDI Distinguished Paper Award, and a PACT Best Paper Award. In his free time, Guha enjoys running, cooking, and reading.

Research topics

  • Artificial Intelligence
  • Computer Science
  • World Wide Web
  • Machine Learning
  • Operating system
  • Programming language
  • Software engineering
  • Theoretical computer science

Selected publications

  • StarCoder: may the source be with you!

    arXiv (Cornell University) · 2023 · 192 citations

    The BigCode community, an open-scientific collaboration working on the responsible development of Large Language Models for Code (Code LLMs), introduces StarCoder and StarCoderBase: 15.5B parameter models with 8K context length, infilling capabilities and fast large-batch inference enabled by multi-query attention. StarCoderBase is trained on 1 trillion tokens sourced from The Stack, a large collection of permissively licensed GitHub repositories with inspection tools and an opt-out process. We…

  • SantaCoder: don't reach for the stars!

    arXiv (Cornell University) · 2023 · 51 citations

    The BigCode project is an open-scientific collaboration working on the responsible development of large language models for code. This tech report describes the progress of the collaboration until December 2022, outlining the current state of the Personally Identifiable Information (PII) redaction pipeline, the experiments conducted to de-risk the model architecture, and the experiments investigating better preprocessing methods for the training data. We train 1.1B parameter models on the Java,…

  • How Beginning Programmers and Code LLMs (Mis)read Each Other

    2024-05-11 · 46 citations

    preprintOpen access

    Generative AI models, specifically large language models (LLMs), have made strides towards the long-standing goal of text-to-code generation. This progress has invited numerous studies of user interaction. However, less is known about the struggles and strategies of non-experts, for whom each step of the text-to-code problem presents challenges: describing their intent in natural language, evaluating the correctness of generated code, and editing prompts when the generated code is incorrect. Thi…

  • Deploying and Evaluating LLMs to Program Service Mobile Robots

    IEEE Robotics and Automation Letters · 2024-01-31 · 30 citations

    article

    Recent advancements in large language models (LLMs) have spurred interest in using them for generating robot programs from natural language, with promising initial results. We investigate the use of LLMs to generate programs for service mobile robots leveraging mobility, perception, and human interaction skills, and where <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">accurate sequencing and ordering</i> of actions is crucial for success. We con…

  • Knowledge Transfer from High-Resource to Low-Resource Programming Languages for Code LLMs

    Proceedings of the ACM on Programming Languages · 2024-10-08 · 29 citations

    articleOpen accessSenior author

    Over the past few years, Large Language Models of Code (Code LLMs) have started to have a significant impact on programming practice. Code LLMs are also emerging as building blocks for research in programming languages and software engineering. However, the quality of code produced by a Code LLM varies significantly by programming language. Code LLMs produce impressive results on high-resource programming languages that are well represented in their training data (e.g., Java, Python, or JavaScri…

Recent grants

Frequent coauthors

Labs

  • Khoury College of Computer SciencesPI

Education

  • Ph.D., Computer Science

    University of California, Los Angeles

    2007
  • M.S., Computer Science

    University of California, Los Angeles

    2003
  • B.S., Computer Science

    University of California, Los Angeles

    2001

Awards & honors

  • OOPSLA Most Influential Paper Award
  • PLDI Distinguished Paper Award
  • PACT Best Paper Award
  • Distinguished Paper Award (2019)

Similar researchers at Northeastern University

  • Resume-aware match score
  • Save to shortlist
  • AI-drafted outreach

See your match with Arjun Guha

PhdFit ranks faculty by your research interests, methods, and publications — grounded in their actual work, not templates.

  • Free to start
  • No credit card
  • 30-second signup