vix.ing · top · new · best · stats · spec

Task Placement and Resource Allocation for Edge Machine Learning: A GNN-based Multi-Agent Reinforcement Learning Paradigm

2023/02/01 by Yihong Li, Li, Yihong, Xiaoxi Zhang +11 · 2 citations
Computer Science · #Advanced Graph Neural Networks #Age of Information Optimization #FOS: Computer and information sciences #IoT and Edge/Fog Computing #Multiagent Systems (cs.MA)

paper · pdf · doi:10.48550/arxiv.2302.00571

openalex publication_date 2023/02/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Machine learning (ML) tasks are one of the major workloads in today's edge computing networks. Existing edge-cloud schedulers allocate the requested amounts of resources to each task, falling short of best utilizing the limited edge resources for ML tasks. This paper proposes TapFinger, a distributed scheduler for edge clusters that minimizes the total completion time of ML tasks through co-optimizing task placement and fine-grained multi-resource allocation. To learn the tasks' uncertain resource sensitivity and enable distributed scheduling, we adopt multi-agent reinforcement learning (MARL) and propose several techniques to make it efficient, including a heterogeneous graph attention network as the MARL backbone, a tailored task selection phase in the actor network, and the integration of Bayes' theorem and masking schemes. We first implement a single-task scheduling version, which schedules at most one task each time. Then we generalize to the multi-task scheduling case, in which a sequence of tasks is scheduled simultaneously. Our design can mitigate the expanded decision space and yield fast convergence to optimal scheduling solutions. Extensive experiments using synthetic and test-bed ML task traces show that TapFinger can achieve up to 54.9% reduction in the average task completion time and improve resource efficiency as compared to state-of-the-art schedulers.

Cited by

Related