vix.ing · top · new · best · stats · spec

Exploring Non-Homogeneity and Dynamicity of High Scale Cloud through\n Hive and Pig

2015/03/23 by Kashish Ara Shakil, Shakil, Kashish Ara, Mansaf Alam +3 · 1 voice
Computer Science · Health Professions · #Artificial Intelligence in Healthcare #Cloud Computing and Resource Management #Data Stream Mining Techniques #IoT and Edge/Fog Computing #cs.DC

paper · pdf · doi:10.48550/arxiv.1503.06600

openalex publication_date 2015/03/23 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Cloud computing deals with heterogeneity and dynamicity at all levels and\ntherefore there is a need to manage resources in such an environment and\nproperly allocate them. Resource planning and scheduling requires a proper\nunderstanding of arrival patterns and scheduling of resources. Study of\nworkloads can aid in proper understanding of their associated environment.\nGoogle has released its latest version of cluster trace, trace version 2.1 in\nNovember 2014.The trace consists of cell information of about 29 days spanning\nacross 700k jobs. This paper deals with statistical analysis of this cluster\ntrace. Since the size of trace is very large, Hive which is a Hadoop\ndistributed file system (HDFS) based platform for querying and analysis of Big\ndata, has been used. Hive was accessed through its Beeswax interface. The data\nwas imported into HDFS through HCatalog. Apart from Hive, Pig which is a\nscripting language and provides abstraction on top of Hadoop was used. To the\nbest of our knowledge the analytical method adopted by us is novel and has\nhelped in gaining several useful insights. Clustering of jobs and arrival time\nhas been done in this paper using K-means++ clustering followed by analysis of\ndistribution of arrival time of jobs which revealed weibull distribution while\nresource usage was close to zip-f like distribution and process runtimes\nrevealed heavy tailed distribution.\n

Discussions

Related