2014/03/06 by Shantenu Jha, Jha, Shantenu, Judy Qiu +7 · 1 citation
Business, Management and Accounting · Computer Science · Decision Sciences · #Advanced Data Storage Technologies #Big Data and Business Intelligence #Cloud Computing and Resource Management #Distributed #FOS: Computer and information sciences #Parallel #Scientific Computing and Data Management #and Cluster Computing (cs.DC)
paper · pdf · doi:10.48550/arxiv.1403.1528
openalex publication_date 2014/03/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Scientific problems that depend on processing large amounts of data require\novercoming challenges in multiple areas: managing large-scale data\ndistribution, co-placement and scheduling of data with compute resources, and\nstoring and transferring large volumes of data. We analyze the ecosystems of\nthe two prominent paradigms for data-intensive applications, hereafter referred\nto as the high-performance computing and the Apache-Hadoop paradigm. We propose\na basis, common terminology and functional factors upon which to analyze the\ntwo approaches of both paradigms. We discuss the concept of "Big Data Ogres"\nand their facets as means of understanding and characterizing the most common\napplication workloads found across the two paradigms. We then discuss the\nsalient features of the two paradigms, and compare and contrast the two\napproaches. Specifically, we examine common implementation/approaches of these\nparadigms, shed light upon the reasons for their current "architecture" and\ndiscuss some typical workloads that utilize them. In spite of the significant\nsoftware distinctions, we believe there is architectural similarity. We discuss\nthe potential integration of different implementations, across the different\nlevels and components. Our comparison progresses from a fully qualitative\nexamination of the two paradigms, to a semi-quantitative methodology. We use a\nsimple and broadly used Ogre (K-means clustering), characterize its performance\non a range of representative platforms, covering several implementations from\nboth paradigms. Our experiments provide an insight into the relative strengths\nof the two paradigms. We propose that the set of Ogres will serve as a\nbenchmark to evaluate the two paradigms along different dimensions.\n