Anurag Khandelwal
Associate Professor of Computer Science at Yale University
I lead the NOVA Lab in the Department of Computer Science at Yale University. My research interests span computer software and hardware systems, networks, and security. My work addresses challenges in processing, storing, and serving large volumes of data to empower real-world systems: from sprawling internet services like social media to critical tools in health and medicine.
I am always looking for motivated graduate students and postdoctoral researchers!
Recent News
[2026]
- Yanpeng successfully defended his Thesis! Congratulations Dr. Yu!
- CORD selected for inclusion in IEEE Micro’s Top Picks in Computer Architecture in 2025!
- BulletTime accepted to ISCA’26!
- Soul accepted to OSDI’26!
- TimelyLLM accepted to MobiSys’26, wins Best Paper Award Runner-Up and Best Artifact Award Runner-Up!
- CounterPoint accepted to ASPLOS’26, wins Best Paper Award!
[2025]
- NSF Award to fortify cloud confidential computing environments (with Lin, Seung-seob)! Thanks NSF!
- Spirit and Mage accepted to SOSP’25!
- Found In Translation accepted to USENIX Security’25!
- Weave accepted to OSDI’25!
- CORD accepted to ISCA’25, wins Distinguished Artifact Award!
- PULSE accepted to ASPLOS’25!
[Older News]
[2024]
- NetApp Faculty Fellowship! Thanks NetApp!
- Yupeng successfully defended his Thesis! Congrats Dr. Tang!
- Length leakage in oblivious storage accepted to USENIX Security’24!
- Trinity accepted to EuroSys’24, wins Best Student Paper Award! Congratulations Ziming Mao!
- PromptCache accepted to MLSys’24!
- SCALO selected for inclusion in IEEE Micro’s Top Picks in Computer Architecture in 2023!
[2023]
- SCALO accepted to ISCA’23, wins Best Paper Award!
- Work on using brain-inspired prefetching techniques for disaggregated memory accepted to HotOS’23!
- Karma accepted to OSDI’23!
- Received Roberts Innovation Fund Award for work on resource disaggregation!
[2022]
- Congratulations to Ziming Mao for CRA Outstanding Undergraduate Researcher Award (runner up)!
- Shepherd accepted to NSDI’23!
- NSF Award to build disaggregated & serverless infrastructure for virtualizing Massive MIMO PHY layer (with Lin)! Thanks NSF!
- Shortstack accepted to OSDI’22!
- Jiffy accepted to EuroSys’22!
[2021]
- MIND accepted to SOSP’21!
- Member of NSF AI Institute for Edge Computing Leveraging Next-generation Networks!
- NSF Award for work on Pancake (with Rachit and Tom)! Thanks NSF!
- What serverless computing is and should become published in CACM!
- NetApp Faculty Fellowship (with Abhishek Bhattacharjee)! Thanks NetApp!
- NSF CAREER Award! Thanks NSF!
- Caerus accepted to NSDI’21!
[2020]
- Pancake accepted to USENIX Security’20, wins Distinguished Paper Award!
- Le Taureau accepted to SIGMOD’20!
- Started at Yale!
Research
Browse the full list of publications, or read about the main research directions below.
Memory Disaggregation: Scaling resources at server granularity wastes resources when compute and memory needs are mismatched, increasing both costs and carbon emissions. Memory disaggregation separates compute and memory into shared network-connected pools to improve resource efficiency, capacity, and elasticity. We are rethinking the cloud software and hardware stack to realize memory disaggregation:
- Scalable cache coherence for disaggregated shared memory pools: SOSP’21, ISCA’25, OSDI’26
- Fair sharing across disaggregated memory resources: OSDI’23, SOSP’25
- Low-latency, high-throughput access to disaggregated memory: ASPLOS’25, SOSP’25
- Understanding memory performance: ISCA’26, ASPLOS’26
Our work on scalable cache coherence for disaggregated memory has been incorporated into NVIDIA’s Vera Rubin Architecture.
Secure cloud systems: As more privacy-sensitive applications move storage and computation to the cloud, encrypted data and secure enclaves can still leak sensitive information through access patterns. Existing defenses are often too expensive in bandwidth or storage to deploy widely. Our research studies these real-world access-pattern vulnerabilities and designs more efficient protections against them.
- Protecting against access pattern vulnerabilities: USENIX Sec’20, OSDI’22, OSDI’25
- Understanding access pattern vulnerabilities: USENIX Sec’24, USENIX Sec’25
Systems for AI: Today’s AI serving systems waste substantial time and resources because they treat requests as independent and unpredictable, even though real workloads contain rich recurring structure in both arrival patterns and prompt content. We are building cloud AI serving platforms that treat workload structure as a first-class systems primitive:
- Scheduling for low-latency, high-throughput AI inference: NSDI’23, MobiSys’26
- Caching attention state across prompts for low-latency inference: MLSys’24
PromptCache has been adopted in Gemini, OpenAI, and Anthropic for reusing attention states across LLM prompts.
Storage and processing stacks for automated data: Emerging applications that rely on automated data sources — ranging from smart vehicles to brain implants — require processing, storing, and serving massive volumes of semantically rich data. We are developing systems for efficient ingestion of data without compromising query and processing performance by exploiting properties specific to machine-generated data:
- High-throughput compressed storage of high-dimensional data: EuroSys’24
- Distributed system for scalable Brain-Computer Interfacing (BCI): ISCA’23, MICRO Top Picks’23
- Distributed monitoring & diagnosis for high speed networks: NSDI’19
[Past Projects]
Serverless Systems: Serverless analytics workloads increasingly demand fine-grained, rapidly changing compute and memory resources, but existing cloud systems manage them too coarsely, forcing a tradeoff between performance and utilization under bursty, time-varying demand. We built a serverless analytics stack that treats elasticity and workload-aware multiplexing as first-class primitives:
- Position papers: UC Berkeley Tech Report, SIGMOD’20, CACM’21
- Enabling fast and cost-effective analytics over serverless functions: NSDI’21, EuroSys’22
Queries on compressed data: As datasets grow beyond DRAM capacity, maintaining interactive query performance becomes difficult because spilling to slower secondary storage increases latency and lowers throughput. We developed systems that address this challenge using a fundamentally new approach: enabling rich query execution directly on compressed data, reducing the need to fully decompress or rely on large DRAM footprints.
Teaching
Operating Systems:
Computer Networks:
Service
Program Committees:
- 2027: EuroSys
- 2026: NSDI, SOSP
- 2025: EuroSys, OSDI
- 2024: SOSP
- 2023: CoNEXT (Poster Co-Chair), NSDI, EuroSys
- 2022: NSDI, HotNets
- 2021: JSys (Editorial Board (Serverless Area)), NSDI, ASPLOS (EPC)
- 2020: SIGCOMM (Poster/Demo, SRC), NSDI