2025/07/08 by Nehal Baganal Krishna, Nehal Baganal-Krishna, Krishna, Nehal Baganal +10
Computer Science · #Acceleration #Age of Information Optimization #Asynchronous communication #Convergence (economics) #IoT and Edge/Fog Computing #Network packet #Queueing theory #Reinforcement learning #Software-Defined Networks and 5G #Transmission (telecommunications) #Traverse
paper · pdf · doi:10.48550/arxiv.2507.05876
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/07/08 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Large-scale training for distributed Machine Learning can cause congestion at bottleneck switch ports, leading to model staleness through update losses. This is particularly detrimental for asynchronous Distributed Reinforcement Learning (DRL) training, as stale updates are known to degrade convergence performance in asynchronous settings. This paper presents Shesha, an in-network DRL accelerator engine, which opportunistically aggregates asynchronously generated model updates on the fly while they traverse the data plane queue. This aggregation operation motivates an alternative queue design, which we prototype and envision for future Top-of-Rack switches. We further present corresponding host-side transmission control in the face of possible congestion, taking advantage of in-network accelerator feedback. A quantification of model staleness, denoted Age-of-Model (AoM), together with a formal verifier allows us to reason on system-wide AoM objectives in multi DRL-cluster scenarios. Shesha shows significant reductions in model staleness and queue congestion, improving overall convergence behavior for asynchronous DRL workloads.