vix.ing · top · new · best · stats · spec

Shesha: Opportunistic In-network Acceleration of Asynchronous Distributed Reinforcement Learning

2025/07/08 by Nehal Baganal Krishna, Nehal Baganal-Krishna, Krishna, Nehal Baganal +10
Computer Science · #Age of Information Optimization #IoT and Edge/Fog Computing #Software-Defined Networks and 5G

paper · pdf · doi:10.48550/arxiv.2507.05876

Abstract

Large-scale training for distributed Machine Learning can cause congestion at bottleneck switch ports, leading to model staleness through update losses. This is particularly detrimental for asynchronous Distributed Reinforcement Learning (DRL) training, as stale updates are known to degrade convergence performance in asynchronous settings. This paper presents Shesha, an in-network DRL accelerator engine, which opportunistically aggregates asynchronously generated model updates on the fly while they traverse the data plane queue. This aggregation operation motivates an alternative queue design, which we prototype and envision for future Top-of-Rack switches. We further present corresponding host-side transmission control in the face of possible congestion, taking advantage of in-network accelerator feedback. A quantification of model staleness, denoted Age-of-Model (AoM), together with a formal verifier allows us to reason on system-wide AoM objectives in multi DRL-cluster scenarios. Shesha shows significant reductions in model staleness and queue congestion, improving overall convergence behavior for asynchronous DRL workloads.

Citations

Related