{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,7,30]],"date-time":"2025-07-30T09:45:21Z","timestamp":1753868721047,"version":"3.41.2"},"reference-count":28,"publisher":"Wiley","issue":"15","license":[{"start":{"date-parts":[[2020,1,8]],"date-time":"2020-01-08T00:00:00Z","timestamp":1578441600000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/onlinelibrary.wiley.com\/termsAndConditions#vor"}],"funder":[{"DOI":"10.13039\/501100004826","name":"Natural Science Foundation of Beijing Municipality","doi-asserted-by":"publisher","award":["4192030"],"award-info":[{"award-number":["4192030"]}],"id":[{"id":"10.13039\/501100004826","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Concurrency and Computation"],"published-print":{"date-parts":[[2020,8,10]]},"abstract":"<jats:title>Summary<\/jats:title><jats:p>With the increase of data processing and Hadoop data center construction requirements, the performance of Hadoop data center is limited by inappropriate resources utilization. This paper introduces a new method to predict utilization for large\u2010scale Hadoop clusters. The new method adopts a two steps model, which includes Hadoop applications' performance simulation and resources utilization prediction. For performance simulation, a new simulator, which integrates baseline test and multilayered network model, is introduced and implemented. A resources utilization predictor is proposed in the second step. By analyzing the pattern of resources utilization, a single task model is proposed. A parallel\u2010batch\u2010task\u2010based (PBT) model, which represents the behavior of real Hadoop applications by integrating the single task model, is introduced. Two test scenarios are configured to verify the performance of our method. For the data center scenario, Terasort, Wordcount, and Hive are selected as benchmarks. In the virtual machines scenario, Terasort is used as benchmark. The experiments show that the error comparing between the simulator results and experimental environment results in most cases is less than 10%. The results confirm that we can locate the resource bottleneck for Hadoop clusters, meanwhile we can agilely configure clusters for applications with massive data.<\/jats:p>","DOI":"10.1002\/cpe.5634","type":"journal-article","created":{"date-parts":[[2020,1,9]],"date-time":"2020-01-09T06:12:50Z","timestamp":1578550370000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["A two steps method of resources utilization predication for large Hadoop data center"],"prefix":"10.1002","volume":"32","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1848-6435","authenticated-orcid":false,"given":"Lei","family":"Yu","sequence":"first","affiliation":[{"name":"Sino\u2010French Engineering School Beihang University  Beijing China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9535-7245","authenticated-orcid":false,"given":"Fei","family":"Teng","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology Southwest Jiaotong University  Chengdu China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shangming","family":"Ning","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology Southwest Jiaotong University  Chengdu China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yunshu","family":"Li","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology Southwest Jiaotong University  Chengdu China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhe","family":"Cui","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology Southwest Jiaotong University  Chengdu China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shengdong","family":"Du","sequence":"additional","affiliation":[{"name":"School of Information Science and Technology Southwest Jiaotong University  Chengdu China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"311","published-online":{"date-parts":[[2020,1,8]]},"reference":[{"key":"e_1_2_7_2_1","unstructured":"Hadoop.Apache Hadoop.http:\/\/hadoop.apache.org\/.2019. Accessed June 13 2019."},{"key":"e_1_2_7_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2017.2700228"},{"key":"e_1_2_7_4_1","doi-asserted-by":"crossref","unstructured":"CaiL QiY LiJ.A recommendation\u2010based parameter tuning approach for Hadoop. Paper presented at: 2017 IEEE 7th International Symposium on Cloud and Service Computing (SC2);2017;Kanazawa Japan.","DOI":"10.1109\/SC2.2017.41"},{"key":"e_1_2_7_5_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.suscom.2017.12.004"},{"key":"e_1_2_7_6_1","unstructured":"HerodotouH LimH LuoG BorisovN DongL CetinFB BabuS.Starfish: a self\u2010tuning system for big data analytics. Paper presented at: Biennial Conference on Innovative Data Systems Research;2011;Monterey CA."},{"key":"e_1_2_7_7_1","doi-asserted-by":"crossref","unstructured":"LinX MengZ XuC WangM.A practical performance model for Hadoop MapReduce. Paper presented at: 2012 IEEE International Conference on Cluster Computing Workshops;2012;Beijing China.","DOI":"10.1109\/ClusterW.2012.24"},{"key":"e_1_2_7_8_1","doi-asserted-by":"crossref","unstructured":"SongG MengZ HuetF MagoulesF YuL LinX.A Hadoop MapReduce performance prediction method. Paper presented at: 2013 IEEE 10th International Conference on High Performance Computing and Communications & 2013 IEEE International Conference on Embedded and Ubiquitous Computing;2013;Zhangjiajie China.","DOI":"10.1109\/HPCC.and.EUC.2013.118"},{"key":"e_1_2_7_9_1","unstructured":"WangG ButtAR PandeyP GuptaK.A simulation approach to evaluating design decisions in MapReduce setups. Paper presented at: 2009 IEEE International Symposium on Modeling Analysis & Simulation of Computer and Telecommunication Systems;2009;London UK."},{"key":"e_1_2_7_10_1","doi-asserted-by":"crossref","unstructured":"HammoudS LiM LiuY AlhamNK LiuZ.MRSim: a discrete event based MapReduce simulator. Paper presented at: 2010 Seventh International Conference on Fuzzy Systems and Knowledge Discovery;2010;Yantai China.","DOI":"10.1109\/FSKD.2010.5569086"},{"key":"e_1_2_7_11_1","doi-asserted-by":"crossref","unstructured":"VermaA CherkasovaL CampbellRH.Play it again simMR!Paper presented at: 2011 IEEE International Conference on Cluster Computing;2011;Austin TX.","DOI":"10.1109\/CLUSTER.2011.36"},{"key":"e_1_2_7_12_1","doi-asserted-by":"crossref","unstructured":"TengF YuL Magoule\u02dcsF.SimMapReduce: a simulator for modeling MapReduce framework. Paper presented at: 2011 Fifth FTRA International Conference on Multimedia and Ubiquitous Engineering;2011;Loutraki Greece.","DOI":"10.1109\/MUE.2011.56"},{"key":"e_1_2_7_13_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.parco.2013.02.001"},{"key":"e_1_2_7_14_1","doi-asserted-by":"crossref","unstructured":"LiuN YangX SunXH JenkinsJ RossR.YARNsim: Simulating Hadoop Yarn. Paper presented at: 2015 15th IEEE\/ACM International Symposium on Cluster Cloud and Grid Computing;2015;Shenzhen China.","DOI":"10.1109\/CCGrid.2015.61"},{"key":"e_1_2_7_15_1","unstructured":"Ganglia.Ganglia monitoring system.http:\/\/ganglia.sourceforge.net\/.2019. Accessed June 13 2019."},{"key":"e_1_2_7_16_1","unstructured":"Nagios.Nagios.https:\/\/www.nagios.org\/.2019. Accessed June 13 2019."},{"key":"e_1_2_7_17_1","unstructured":"Ambari.Apache Ambari.https:\/\/ambari.apache.org\/.2019. Accessed June 13 2019."},{"key":"e_1_2_7_18_1","unstructured":"\u201cDr\u2010elephant\u201d.Linkedin dr\u2010elephant.https:\/\/github.com\/linkedin\/dr-elephant\/.2019. Accessed June 13 2019."},{"key":"e_1_2_7_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSP.2006.889401"},{"key":"e_1_2_7_20_1","doi-asserted-by":"crossref","unstructured":"LiuZ ChoS.Characterizing machines and workloads on a Google cluster. Paper presented at: 2012 41st International Conference on Parallel Processing Workshops;2012;Pittsburgh PA.","DOI":"10.1109\/ICPPW.2012.57"},{"key":"e_1_2_7_21_1","doi-asserted-by":"crossref","unstructured":"MehmoodT LatifS MalikS.Prediction of cloud computing resource utilization. Paper presented at: 2018 15th International Conference on Smart Cities: Improving Quality of Life Using ICT & IoT (HONET\u2010ICT);2018;Islamabad Pakistan.","DOI":"10.1109\/HONET.2018.8551339"},{"key":"e_1_2_7_22_1","doi-asserted-by":"crossref","unstructured":"LeiF YuL ShaoB TengF ZhouB.Large scale data centers simulation based on baseline test model. Paper presented at: 2018 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW);2018;Vancouver Canada.","DOI":"10.1109\/IPDPSW.2018.00018"},{"key":"e_1_2_7_23_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2014.06.008"},{"key":"e_1_2_7_24_1","unstructured":"CarothersCD BauerD PearceS.ROSS: a high\u2010performance low memory modular time warp system. In: Proceedings of the Fourteenth Workshop on Parallel and Distributed Simulation;2000;Bologna Italy."},{"key":"e_1_2_7_25_1","unstructured":"Codes: enabling co\u2010design of multilayer exascale storage architectures.http:\/\/www.mcs.anl.gov\/projects\/codes\/.2019. Accessed June 13 2019."},{"key":"e_1_2_7_26_1","unstructured":"ArmstrongB EigenmannR.Performance forecasting: towards a methodology for characterizing large computational applications. In: Proceedings 1998 International Conference on Parallel Processing;1998;Minneapolis MN."},{"key":"e_1_2_7_27_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.peva.2008.03.005"},{"key":"e_1_2_7_28_1","unstructured":"KarunAK ChitharanjanK.A review on Hadoop \u2010 HDFS infrastructure extensions. Paper presented at: 2013 IEEE Conference on Information & Communication Technologies;2013;Thuckalay India."},{"key":"e_1_2_7_29_1","doi-asserted-by":"crossref","unstructured":"YazdanovL GorbunovM FetzerC.EHadoop: network i\/o aware scheduler for elastic MapReduce cluster. Paper presented at: 2015 IEEE 8th International Conference on Cloud Computing;2015;New York NY.","DOI":"10.1109\/CLOUD.2015.113"}],"container-title":["Concurrency and Computation: Practice and Experience"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/api.wiley.com\/onlinelibrary\/tdm\/v1\/articles\/10.1002%2Fcpe.5634","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1002\/cpe.5634","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/full-xml\/10.1002\/cpe.5634","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1002\/cpe.5634","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,9,6]],"date-time":"2023-09-06T09:23:04Z","timestamp":1693992184000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1002\/cpe.5634"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,1,8]]},"references-count":28,"journal-issue":{"issue":"15","published-print":{"date-parts":[[2020,8,10]]}},"alternative-id":["10.1002\/cpe.5634"],"URL":"https:\/\/doi.org\/10.1002\/cpe.5634","archive":["Portico"],"relation":{},"ISSN":["1532-0626","1532-0634"],"issn-type":[{"type":"print","value":"1532-0626"},{"type":"electronic","value":"1532-0634"}],"subject":[],"published":{"date-parts":[[2020,1,8]]},"assertion":[{"value":"2018-10-24","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-11-18","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-01-08","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e5634"}}