Resume-aware faculty matching

Find professors who actually fit you

Review faculty evidence in public, then use the workspace to turn your background into a shortlist, outreach, and meeting prep.

Profile-awarePaper evidenceSix agents
Marten van Schijndel

Marten van Schijndel

· Assistant Professor

Cornell University · Linguistics

Active 1998–2026

h-index16
Citations949
Papers6529 last 5y
Funding

Academic metrics are sourced from OpenAlex and public funding records; values may differ from Google Scholar.

See your match with Marten van Schijndel — sign in to PhdFit.Sign in

About

Professor Marten van Schijndel is an Assistant Professor in the Department of Linguistics at Cornell University, affiliated with the Cognitive Science Program. His main research interest is in the representations that can be used to process language incrementally. He probes the linguistic processing of humans and neural network language models using methodologies from psycholinguistics, comparing the predictions of computational models with human behavioral and neural responses. He organizes the Computational Psycholinguistic Discussions research group (C.Psyd) and co-organizes the Natural Language Processing research group. His work focuses on computational modelling and psycholinguistics, contributing to understanding linguistic processing through the use of computational tools and experimental methodologies.

Research topics

  • Computer Science
  • Natural Language Processing
  • Artificial Intelligence
  • Linguistics
  • Statistics
  • Mathematics

Selected publications

  • Single‐Stage Prediction Models Do Not Explain the Magnitude of Syntactic Disambiguation Difficulty

    Cognitive Science · 2021 · 64 citations

    1st authorCorresponding

    The disambiguation of a syntactically ambiguous sentence in favor of a less preferred parse can lead to slower reading at the disambiguation point. This phenomenon, referred to as a garden-path effect, has motivated models in which readers initially maintain only a subset of the possible parses of the sentence, and subsequently require time-consuming reanalysis to reconstruct a discarded parse. A more recent proposal argues that the garden-path effect can be reduced to surprisal arising in a ful…

  • All Bark and No Bite: Rogue Dimensions in Transformer Language Models Obscure Representational Quality

    Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing · 2021 · 59 citations

    Senior authorCorresponding

    Similarity measures are a vital tool for understanding how language models represent and process language. Standard representational similarity measures such as cosine similarity and Euclidean distance have been successfully used in static word embedding models to understand how words cluster in semantic space. Recently, these measures have been applied to embeddings from contextualized models such as BERT and GPT-2. In this work, we call into question the informativity of such measures for cont…

  • Semantics or spelling? Probing contextual word embeddings with orthographic noise

    2024-01-01 · 1 citations

    articleOpen accessSenior author

    Pretrained language model (PLM) hidden states are frequently employed as contextual word embeddings (CWE): high-dimensional representations that encode semantic information given linguistic context.Across many areas of computational linguistics research, similarity between CWEs is interpreted as semantic similarity.However, it remains unclear exactly what information is encoded in PLM hidden states.We investigate this practice by probing PLM representations using minimal orthographic noise.We ex…

  • Does Dependency Locality Predict Non-canonical Word Order in Hindi?

    arXiv (Cornell University) · 2024-05-13 · 1 citations

    preprintOpen accessSenior author

    Previous work has shown that isolated non-canonical sentences with Object-before-Subject (OSV) order are initially harder to process than their canonical counterparts with Subject-before-Object (SOV) order. Although this difficulty diminishes with appropriate discourse context, the underlying cognitive factors responsible for alleviating processing challenges in OSV sentences remain a question. In this work, we test the hypothesis that dependency length minimization is a significant predictor of…

  • Linguistic Compression in Single-Sentence Human-Written Summaries

    2023-01-01 · 1 citations

    articleOpen accessSenior author

    Summarizing texts involves significant cognitive efforts to compress information. While advances in automatic summarization systems have drawn attention from the NLP and linguistics communities to this topic, there is a lack of computational studies of linguistic patterns in human-written summaries. This work presents a large-scale corpus study of human-written single-sentence summaries. We analyzed the linguistic compression patterns from source documents to summaries at different granularities…

Frequent coauthors

  • William Schuler

    The Ohio State University

    19 shared
  • Tal Linzen

    18 shared
  • Forrest Davis

    12 shared
  • Evelina Fedorenko

    Massachusetts Institute of Technology

    9 shared
  • Idan Blank

    7 shared
  • William Timkey

    New York University

    6 shared
  • Cory Shain

    Massachusetts Institute of Technology

    6 shared
  • Sidharth Ranjan

    5 shared

Similar researchers at Cornell University

  • Resume-aware match score
  • Save to shortlist
  • AI-drafted outreach

See your match with Marten van Schijndel

PhdFit ranks faculty by your research interests, methods, and publications — grounded in their actual work, not templates.

  • Free to start
  • No credit card
  • 30-second signup