
Nikos Hardavellas
· Professor of Computer ScienceNorthwestern University · Chemical Engineering
Active 1993–2026
Academic metrics are sourced from OpenAlex and public funding records; values may differ from Google Scholar.
About
Nikos Hardavellas is a professor of Computer Science and Computer Engineering at Northwestern University, where he directs the Parallel Architecture Group at Northwestern (PARAG@N). His research focuses on computer architecture, specifically at the intersection of computer architecture with the computer systems stack, including programming languages, compilers, and operating systems. His work also encompasses memory systems, nanophotonics, energy-efficient computing, and quantum computing systems. Hardavellas serves on the Executive Committee of the Northwestern Institute for Quantum Information Research and Engineering (INQUIRE) and the Scientific Advisory Committee of the National Quantum Algorithms Center (NQAC). He has received numerous awards and recognitions, including an NSF CAREER award, being named a Future CRA Leader, and receiving best paper awards at various conferences. Prior to his academic career, he contributed to the design of several generations of Alpha microprocessors and high-end multiprocessor servers at Digital Equipment Corp., Compaq, and Hewlett-Packard.
Research topics
- Computer Science
- Parallel computing
- Programming language
- Artificial Intelligence
- Operating system
- Embedded system
- Database
- Distributed computing
- Algorithm
- Computer engineering
Selected publications
AccelWattch: A Power Modeling Framework for Modern GPUs
2021 · 108 citations
Senior authorCorrespondingGraphics Processing Units (GPUs) are rapidly dominating the accelerator space, as illustrated by their wide-spread adoption in the data analytics and machine learning markets. At the same time, performance per watt has emerged as a crucial evaluation metric together with peak performance. As such, GPU architects require robust tools that will enable them to model both the performance and the power consumption of modern GPUs. However, while GPU performance modeling has progressed in great strides…
CARAT: a case for virtual memory through compiler- and runtime-based address translation
2020 · 14 citations
Virtual memory is a critical abstraction in modern computer systems. Its common model, paging, is currently seeing considerable innovation, yet its implementations continue to be co-designs between power-hungry/latency-adding hardware (e.g., TLBs, pagewalk caches, pagewalkers, etc) and software (the OS kernel). We make a case for a new model for virtual memory, compiler- and runtime-based address translation (CARAT), which instead is a co-design between the compiler and the OS kernel. CARAT can…
SupermarQ: A Scalable Quantum Benchmark Suite
arXiv (Cornell University) · 2022-02-22 · 10 citations
preprintOpen accessThe emergence of quantum computers as a new computational paradigm has been accompanied by speculation concerning the scope and timeline of their anticipated revolutionary changes. While quantum computing is still in its infancy, the variety of different architectures used to implement quantum computations make it difficult to reliably measure and compare performance. This problem motivates our introduction of SupermarQ, a scalable, hardware-agnostic quantum benchmark suite which uses applicatio…
2022-09-01 · 9 citations
articleSenior authorMPI collective communication is an omnipresent communication model for high-performance computing (HPC) systems. The performance of a collective operation depends strongly on the algorithm used to implement it. MPI libraries use inaccurate heuristics to select these algorithms, causing applications to suffer unnecessary slowdowns. Machine learning (ML)-based autotuners are a promising alternative. ML autotuners can intelligently select algorithms for individual jobs, resulting in near-optimal pe…
A FACT-based Approach: Making Machine Learning Collective Autotuning Feasible on Exascale Systems
2021-11-01 · 9 citations
articleAccording to recent performance analyses, MPI collective operations make up a quarter of the execution time on production systems. Machine learning (ML) autotuners use supervised learning to select collective algorithms, significantly improving collective performance. However, we observe two barriers preventing their adoption over the default heuristic-based autotuners. First, a user may find it difficult to compare autotuners because we lack a methodology to quantify their performance. We call…
Recent grants
Frequent coauthors
- 28 shared
Babak Falsafi
- 22 shared
Ippokratis Pandis
Amazon (United States)
- 19 shared
Chris Wilkerson
Intel (United States)
- 19 shared
Tor M. Aamodt
University of British Columbia
- 17 shared
Anastasia Ailamaki
- 17 shared
Lieven Eeckhout
Ghent University
- 17 shared
Jared C. Smolens
Oracle (United States)
- 17 shared
Juanita Hoe
University of West London
Labs
Parallel Architecture Group at Northwestern (PARAG@N)PI
Education
- 2009
PhD, CS
Carnegie Mellon University
- 2006
MSc, CS
Carnegie Mellon University
- 1997
MSc, CS
University of Rochester
- 1995
BSc, CS
University of Crete
Awards & honors
- Future CRA Leader by the Computing Research Association (202…
- NSF CAREER award (2015)
- Best paper awards, nominations and test-of-time awards at HP…
- IEEE Micro Top Picks Award (2010)
- IEEE Micro Top Picks Honorable Mention (2023)
Similar researchers at Northwestern University
- Resume-aware match score
- Save to shortlist
- AI-drafted outreach
See your match with Nikos Hardavellas
PhdFit ranks faculty by your research interests, methods, and publications — grounded in their actual work, not templates.
- Free to start
- No credit card
- 30-second signup
