Resume-aware faculty matching

Find professors who actually fit you

Review faculty evidence in public, then use the workspace to turn your background into a shortlist, outreach, and meeting prep.

Profile-awarePaper evidenceSix agents
Bryan Pardo

Bryan Pardo

· Professor of Electrical Engineering and Computer Science

Northwestern University · Radio/Television/Film

Active 1975–2026

h-index30
Citations3.3k
Papers20740 last 5y
Funding$2.2M

Academic metrics are sourced from OpenAlex and public funding records; values may differ from Google Scholar.

See your match with Bryan Pardo — sign in to PhdFit.Sign in

About

Bryan Pardo is a Professor of Computer Science at Northwestern University, where he is also the co-director of the Northwestern University Center for Human Computer Interaction + Design and head of the Interactive Audio Lab. His research focuses on developing new methods in Generative Modeling, Signal Processing, and Human Computer Interaction to create tools for understanding, creating, and manipulating sound. His ongoing research includes applications in music and speech generation, audio scene labeling, audio source separation, inclusive interfaces, new audio production tools, and machine audition models. Pardo's work aims to advance the understanding and manipulation of audio signals, contributing to various societal and technological domains.

Research topics

  • Computer science
  • Speech recognition
  • Artificial intelligence
  • Multimedia
  • Natural language processing

Selected publications

  • Sketch2Sound: Controllable Audio Generation via Time-Varying Signals and Sonic Imitations

    2025-03-12 · 14 citations

    article

    We present Sketch2Sound, a generative audio model capable of creating high-quality sounds from a set of interpretable time-varying control signals: loudness, brightness, and pitch, as well as text prompts. Sketch2Sound can synthesize arbitrary sounds from sonic imitations (i.e., a vocal imitation or a reference sound-shape). Sketch2Sound can be implemented on top of any text-to-audio latent diffusion transformer (DiT), and requires only 40k steps of fine-tuning and a single linear layer per cont…

  • Maskmark: Robust Neuralwatermarking for Real and Synthetic Speech

    2024-03-18 · 10 citations

    articleSenior author

    High-quality speech synthesis models may be used to spread misinformation or impersonate voices. Audio watermarking can combat misuse by embedding a traceable signature in generated audio. However, existing audio watermarks typically demonstrate robustness to only a small set of transformations of the watermarked audio. To address this, we propose MaskMark, a neural network-based digital audio watermarking technique optimized for speech. MaskMark embeds a secret key vector in audio via a multipl…

  • High-Fidelity Neural Phonetic Posteriorgrams

    2024-04-14 · 8 citations

    articleSenior author

    A phonetic posteriorgram (PPG) is a time-varying categorical distribution over acoustic units of speech (e.g., phonemes). PPGs are a popular representation in speech generation due to their ability to disentangle pronunciation features from speaker identity, allowing accurate reconstruction of pronunciation (e.g., voice conversion) and coarse-grained pronunciation editing (e.g., foreign accent conversion). In this paper, we demonstrably improve the quality of PPGs to produce a state-of-the-art i…

  • VampNet: Music Generation via Masked Acoustic Token Modeling

    arXiv (Cornell University) · 2023-07-10 · 7 citations

    preprintOpen accessSenior author

    We introduce VampNet, a masked acoustic token modeling approach to music synthesis, compression, inpainting, and variation. We use a variable masking schedule during training which allows us to sample coherent music from the model by applying a variety of masking approaches (called prompts) during inference. VampNet is non-autoregressive, leveraging a bidirectional transformer architecture that attends to all tokens in a forward pass. With just 36 sampling passes, VampNet can generate coherent h…

  • Fine-Grained and Interpretable Neural Speech Editing

    2024-09-01 · 5 citations

    articleSenior author

Recent grants

Frequent coauthors

  • Prem Seetharaman

    42 shared
  • Max Morrison

    Northwestern University

    22 shared
  • William P. Birmingham

    Saint Vincent College

    20 shared
  • Mark Cartwright

    New York University

    20 shared
  • Zafar Rafii

    19 shared
  • Ethan Manilow

    18 shared
  • Antoine Liutkus

    Institut National Polytechnique de Toulouse

    17 shared
  • Zhiyao Duan

    China United Network Communications Group (China)

    16 shared

Similar researchers at Northwestern University

  • Resume-aware match score
  • Save to shortlist
  • AI-drafted outreach

See your match with Bryan Pardo

PhdFit ranks faculty by your research interests, methods, and publications — grounded in their actual work, not templates.

  • Free to start
  • No credit card
  • 30-second signup