2026 · Conference paper
TimelyLLM: Time-sensitive LLM Serving System for Physical-I/O Limited Agents
ACM MobiSys, 2026
Best Paper Award Runner-Up, Best Artifact Award Runner-Up
Abstract
Large language models are increasingly integrated into physical-I/O-limited agents such as robots and voice assistants, yet existing serving systems optimize throughput without accounting for the gap between fast token generation and slow physical execution. TimelyLLM coordinates generation with agent behavior through segmented generation and scheduling, using the gap between plan generation and execution to reduce contention and response latency under multi-agent workloads. Built atop a widely used serving framework and evaluated with workloads from drones, robot arms, and quadrupeds, TimelyLLM improves time utility by up to 1.52 times and reduces overall waiting time by 84 percent.
Publication details
- Venue
- ACM MobiSys
- Publication year
- 2026
- Awards
- Best Paper Award Runner-Up, Best Artifact Award Runner-Up
BibTeX
@inproceedings{timelyllm,
author = {Ling, Neiwen and Chen, Guojun and Khandelwal, Anurag and Zhong, Lin},
title = {{TimelyLLM: Time-sensitive LLM Serving System for Physical-I/O Limited Agents}},
year = {2026},
month = jun,
booktitle = {ACM MobiSys},
award = {Best Paper Award Runner-Up, Best Artifact Award Runner-Up}
}