STEAMEngine: Driving MapReduce provisioning in the cloud

Michael Cardosa, Piyush Narang, Abhishek Chandra, Himabindu Pucha, Aameek Singh

Research output: Chapter in Book/Report/Conference proceedingConference contribution

16 Scopus citations

Abstract

MapReduce has gained in popularity as a distributed data analysis paradigm, particularly in the cloud, where MapReduce jobs are run on virtual clusters. The provisioning of MapReduce jobs in the cloud is an important problem for optimizing several user as well as provider-side metrics, such as runtime, cost, throughput, energy, and load. In this paper, we present an intelligent provisioning framework called STEAMEngine that consists of provisioning algorithms to optimize these metrics through a set of common building blocks. These building blocks enable spatio-temporal tradeoffs unique to MapReduce provisioning: along with their resource requirements (spatial component), a MapReduce job runtime (temporal component) is a critical element for any provisioning algorithm. We also describe tw o novel provisioning algorithms a user-driven performance optimization and a provider-driven energy optimization that leverage these building blocks. Our experimental results based on an Amazon EC2 cluster and a local Xen/Hadoop cluster show the benefits of STEAMEngine through improvements in performance and energy via the use of these algorithms and building blocks.

Original languageEnglish (US)
Title of host publication18th International Conference on High Performance Computing, HiPC 2011
DOIs
StatePublished - Dec 1 2011
Event18th International Conference on High Performance Computing, HiPC 2011 - Bangalore, India
Duration: Dec 18 2011Dec 21 2011

Publication series

Name18th International Conference on High Performance Computing, HiPC 2011

Other

Other18th International Conference on High Performance Computing, HiPC 2011
CountryIndia
CityBangalore
Period12/18/1112/21/11

Fingerprint Dive into the research topics of 'STEAMEngine: Driving MapReduce provisioning in the cloud'. Together they form a unique fingerprint.

  • Cite this

    Cardosa, M., Narang, P., Chandra, A., Pucha, H., & Singh, A. (2011). STEAMEngine: Driving MapReduce provisioning in the cloud. In 18th International Conference on High Performance Computing, HiPC 2011 [6152649] (18th International Conference on High Performance Computing, HiPC 2011). https://doi.org/10.1109/HiPC.2011.6152649