2017/05/30 by Niyazi Sorkunlu, Sorkunlu, Niyazi, Varun Chandola +3 · 1 citation
Computer Science · #Computational Physics and Python Applications #Distributed and Parallel Computing Systems #FOS: Computer and information sciences #Performance (cs.PF) #Software System Performance and Reliability #cs.PF
paper · pdf · doi:10.48550/arxiv.1705.10756
arxiv created 2017/05/30 · openalex publication_date 2017/05/30 · arxiv updated 2017/05/31 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Resource usage data, collected using tools such as TACC Stats, capture the resource utilization by nodes within a high performance computing system. We present methods to analyze the resource usage data to understand the system performance and identify performance anomalies. The core idea is to model the data as a three-way tensor corresponding to the compute nodes, usage metrics, and time. Using the reconstruction error between the original tensor and the tensor reconstructed from a low rank tensor decomposition, as a scalar performance metric, enables us to monitor the performance of the system in an online fashion. This error statistic is then used for anomaly detection that relies on the assumption that the normal/routine behavior of the system can be captured using a low rank approx- imation of the original tensor. We evaluate the performance of the algorithm using information gathered from system logs and show that the performance anomalies identified by the proposed method correlates with critical errors reported in the system logs. Results are shown for data collected for 2013 from the Lonestar4 system at the Texas Advanced Computing Center (TACC)