
Bryan Pardo
· Professor of Electrical Engineering and Computer ScienceNorthwestern University · Radio/Television/Film
Active 1975–2026
Academic metrics are sourced from OpenAlex and public funding records; values may differ from Google Scholar.
About
Bryan Pardo is a Professor of Computer Science at Northwestern University, where he is also the co-director of the Northwestern University Center for Human Computer Interaction + Design and head of the Interactive Audio Lab. His research focuses on developing new methods in Generative Modeling, Signal Processing, and Human Computer Interaction to create tools for understanding, creating, and manipulating sound. His ongoing research includes applications in music and speech generation, audio scene labeling, audio source separation, inclusive interfaces, new audio production tools, and machine audition models. Pardo's work aims to advance the understanding and manipulation of audio signals, contributing to various societal and technological domains.
Research topics
- Computer science
- Speech recognition
- Artificial intelligence
- Multimedia
- Natural language processing
Selected publications
Sketch2Sound: Controllable Audio Generation via Time-Varying Signals and Sonic Imitations
2025-03-12 · 14 citations
articleWe present Sketch2Sound, a generative audio model capable of creating high-quality sounds from a set of interpretable time-varying control signals: loudness, brightness, and pitch, as well as text prompts. Sketch2Sound can synthesize arbitrary sounds from sonic imitations (i.e., a vocal imitation or a reference sound-shape). Sketch2Sound can be implemented on top of any text-to-audio latent diffusion transformer (DiT), and requires only 40k steps of fine-tuning and a single linear layer per cont…
Maskmark: Robust Neuralwatermarking for Real and Synthetic Speech
2024-03-18 · 10 citations
articleSenior authorHigh-quality speech synthesis models may be used to spread misinformation or impersonate voices. Audio watermarking can combat misuse by embedding a traceable signature in generated audio. However, existing audio watermarks typically demonstrate robustness to only a small set of transformations of the watermarked audio. To address this, we propose MaskMark, a neural network-based digital audio watermarking technique optimized for speech. MaskMark embeds a secret key vector in audio via a multipl…
High-Fidelity Neural Phonetic Posteriorgrams
2024-04-14 · 8 citations
articleSenior authorA phonetic posteriorgram (PPG) is a time-varying categorical distribution over acoustic units of speech (e.g., phonemes). PPGs are a popular representation in speech generation due to their ability to disentangle pronunciation features from speaker identity, allowing accurate reconstruction of pronunciation (e.g., voice conversion) and coarse-grained pronunciation editing (e.g., foreign accent conversion). In this paper, we demonstrably improve the quality of PPGs to produce a state-of-the-art i…
VampNet: Music Generation via Masked Acoustic Token Modeling
arXiv (Cornell University) · 2023-07-10 · 7 citations
preprintOpen accessSenior authorWe introduce VampNet, a masked acoustic token modeling approach to music synthesis, compression, inpainting, and variation. We use a variable masking schedule during training which allows us to sample coherent music from the model by applying a variety of masking approaches (called prompts) during inference. VampNet is non-autoregressive, leveraging a bidirectional transformer architecture that attends to all tokens in a forward pass. With just 36 sampling passes, VampNet can generate coherent h…
Fine-Grained and Interpretable Neural Speech Editing
2024-09-01 · 5 citations
articleSenior author
Recent grants
NSF · $476k · 2008–2012
CAREER: Making music documents accessible in musical terms
NSF · $507k · 2007–2012
CHS: Small: Robust Interactive Audio Source Separation
NSF · $514k · 2014–2018
Frequent coauthors
- 42 shared
Prem Seetharaman
- 22 shared
Max Morrison
Northwestern University
- 20 shared
William P. Birmingham
Saint Vincent College
- 20 shared
Mark Cartwright
New York University
- 19 shared
Zafar Rafii
- 18 shared
Ethan Manilow
- 17 shared
Antoine Liutkus
Institut National Polytechnique de Toulouse
- 16 shared
Zhiyao Duan
China United Network Communications Group (China)
Similar researchers at Northwestern University
- Resume-aware match score
- Save to shortlist
- AI-drafted outreach
See your match with Bryan Pardo
PhdFit ranks faculty by your research interests, methods, and publications — grounded in their actual work, not templates.
- Free to start
- No credit card
- 30-second signup
