{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,7,30]],"date-time":"2025-07-30T09:45:23Z","timestamp":1753868723171,"version":"3.41.2"},"reference-count":52,"publisher":"Wiley","issue":"15","license":[{"start":{"date-parts":[[2022,7,22]],"date-time":"2022-07-22T00:00:00Z","timestamp":1658448000000},"content-version":"am","delay-in-days":365,"URL":"http:\/\/onlinelibrary.wiley.com\/termsAndConditions#am"},{"start":{"date-parts":[[2021,7,22]],"date-time":"2021-07-22T00:00:00Z","timestamp":1626912000000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/onlinelibrary.wiley.com\/termsAndConditions#vor"}],"funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["CCF\u20101801856"],"award-info":[{"award-number":["CCF\u20101801856"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000015","name":"U.S. Department of Energy","doi-asserted-by":"publisher","award":["DE\u2010AC02\u201006CH11357"],"award-info":[{"award-number":["DE\u2010AC02\u201006CH11357"]}],"id":[{"id":"10.13039\/100000015","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Concurrency and Computation"],"published-print":{"date-parts":[[2023,7,10]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Machine learning (ML) continues to grow in importance across nearly all domains in modeling to learn from data. Often a tradeoff exists between a model's ability to minimize bias and variance. In this article, we utilize ensemble learning to combine linear, nonlinear, and tree\u2010\/rule\u2010based ML methods to cope with the bias\u2010variance tradeoff and result in more accurate models. We use the datasets collected for two parallel cancer deep learning CANDLE benchmarks, NT3 and P1B2, to build performance and power models based on hardware performance counters using single\u2010object and multiple\u2010objects ensemble learning to identify the most important counters for improvement on the Cray XC40 Theta at Argonne National Laboratory. Based on the insights from these models, we improve the performance and energy of P1B2 and NT3 by optimizing the deep learning environments TensorFlow, Keras, Horovod, and Python under the huge page size of 8 MB. Experimental results show that ensemble learning not only produces more accurate models but also provides more robust performance counter ranking. We achieve up to 61.15% performance improvement and up to 62.58% energy saving for P1B2 and up to 55.81% performance improvement and up to 52.60% energy saving for NT3 on up to 24,576 cores.<\/jats:p>","DOI":"10.1002\/cpe.6516","type":"journal-article","created":{"date-parts":[[2021,7,23]],"date-time":"2021-07-23T03:54:35Z","timestamp":1627012475000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Utilizing ensemble learning for performance and power modeling and improvement of parallel cancer deep learning CANDLE benchmarks"],"prefix":"10.1002","volume":"35","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8150-5171","authenticated-orcid":false,"given":"Xingfu","family":"Wu","sequence":"first","affiliation":[{"name":"Mathematics &amp; Computer Science Division, Argonne National Laboratory The University of Chicago  Lemont Illinois USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Valerie","family":"Taylor","sequence":"additional","affiliation":[{"name":"Mathematics &amp; Computer Science Division, Argonne National Laboratory The University of Chicago  Lemont Illinois USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"311","published-online":{"date-parts":[[2021,7,22]]},"reference":[{"key":"e_1_2_10_2_1","doi-asserted-by":"crossref","unstructured":"SinghK BhadhauriaM McKeeS. A. Real time power estimation and thread scheduling via performance counters. ACM SIGARCH Computer Architecture News Volume 37 Issue 2 May 2009 pp 46\u201355.","DOI":"10.1145\/1577129.1577137"},{"key":"e_1_2_10_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2011.47"},{"key":"e_1_2_10_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSII.2013.2285966"},{"key":"e_1_2_10_5_1","unstructured":"MayerZ. A brief Introduction to caretEnsemble;2016."},{"key":"e_1_2_10_6_1","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511973000"},{"key":"e_1_2_10_7_1","unstructured":"IsciC MartonosiM. Runtime power monitoring in high\u2010end processors: methodology and empirical data. Paper presented at: Proceedings of the 36th IEEE\/ACM International Symposium on Microarchitecture;San Diego CA USA 2003."},{"key":"e_1_2_10_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/1077603.1077657"},{"key":"e_1_2_10_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/1183401.1183426"},{"key":"e_1_2_10_10_1","doi-asserted-by":"crossref","unstructured":"LimM PorterfieldA FowlerR. SoftPower: fine\u2010grain power estimations using performance counters. Paper presented at: Proceedings of the 19th ACM International Symposium on High Performance Distributed Computing (HPDC'10);Chicago Illinois 2010.","DOI":"10.1145\/1851476.1851517"},{"key":"e_1_2_10_11_1","doi-asserted-by":"crossref","unstructured":"ChenX XuC DickR MaoZ. Performance and power modeling in a multi\u2010programmed multi\u2010core Environment. Paper presented at: Proceedings of the 47th Design Automation Conference DAC2010;2010.","DOI":"10.1145\/1837274.1837479"},{"key":"e_1_2_10_12_1","doi-asserted-by":"crossref","unstructured":"NagasakaH MaruyamaN NukadaA EndoT MatsuokaS. Statistical power modeling of GPU kernels using performance counters. Paper presented at: Proceedings of the International Green Computing Conference;2010.","DOI":"10.1109\/GREENCOMP.2010.5598315"},{"key":"e_1_2_10_13_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00450-011-0190-0"},{"key":"e_1_2_10_14_1","doi-asserted-by":"crossref","unstructured":"SongS SuC RountreeB CameronK. A simplified and accurate model of power\u2010performance efficiency on emergent GPU architectures. Paper presented at: Proceedings of the 2013 IEEE 27th International Symposium on Parallel and Distributed Processing; Vol. 20 2013:673\u2010686.","DOI":"10.1109\/IPDPS.2013.73"},{"key":"e_1_2_10_15_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00450-013-0239-3"},{"key":"e_1_2_10_16_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.future.2013.07.010"},{"key":"e_1_2_10_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/MC.2016.311"},{"key":"e_1_2_10_18_1","doi-asserted-by":"crossref","unstructured":"WuX TaylorV. Utilizing hardware performance counters to model and optimize the energy and performance of large scaled scientific applications on power\u2010aware supercomputers. Paper presented at: Proceedings of the IPDPS2016 Workshop on High Performance Power\u2010Aware Computing; May 23\u201327 2016; Chicago IL.","DOI":"10.1109\/IPDPSW.2016.78"},{"key":"e_1_2_10_19_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10766-017-0487-0"},{"key":"e_1_2_10_20_1","unstructured":"GreathouseJL LohGH. Machine learning for performance and power modeling of heterogeneous systems (Invited Paper). Paper presented at: Proceedings of the IEEE\/ACM International Conference on Computer\u2010Aided Design (ICCAD'18); November 5\u20138 2018; San Diego CA."},{"key":"e_1_2_10_21_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4614-6849-3"},{"key":"e_1_2_10_22_1","unstructured":"KuhnM. The caret package; March 27 2019.https:\/\/topepo.github.io\/caret\/index.html (https:\/\/cran.r\u2010project.org\/web\/packages\/caret\/)."},{"key":"e_1_2_10_23_1","unstructured":"Scikit\u2010Learn machine learning in python.2019.https:\/\/scikit\u2010learn.org\/stable\/index.html"},{"key":"e_1_2_10_24_1","unstructured":"WuX TaylorV LanZ. Performance and power modeling and prediction using mummi and ten machine learning methods. Paper presented at: Proceedings of the 2020 Cray User Group Conference; October 2020."},{"key":"e_1_2_10_25_1","unstructured":"CANDLE: cancer distributed learning environment.2019.http:\/\/candle.cels.anl.gov"},{"key":"e_1_2_10_26_1","unstructured":"CANDLE benchmarks;2019.https:\/\/github.com\/ECP\u2010CANDLE\/Benchmarks. Accessed September 2019."},{"key":"e_1_2_10_27_1","doi-asserted-by":"crossref","unstructured":"WuX TaylorV WozniakJM StevensR BrettinT XiaF. Performance energy and scalability analysis and improvement of parallel cancer deep learning CANDLE benchmarks. Paper presented at: Proceedings of the 48th International Conference on Parallel Processing (ICPP2019); August 5\u20138 2019; Kyoto Japan.","DOI":"10.1145\/3337821.3337905"},{"key":"e_1_2_10_28_1","unstructured":"Summit.2019.https:\/\/www.olcf.ornl.gov\/olcf\u2010resources\/compute\u2010systems\/summit\/"},{"key":"e_1_2_10_29_1","unstructured":"Theta Cray XC40 system.2019.https:\/\/www.alcf.anl.gov\/theta"},{"key":"e_1_2_10_30_1","unstructured":"MillerP. mvtboost example; December 5 2016.https:\/\/cran.r\u2010project.org\/web\/packages\/mvtboost\/vignettes\/mvtboost_vignette.html."},{"key":"e_1_2_10_31_1","unstructured":"glmnet: lasso and elastic\u2010net regularized generalized linear models.https:\/\/cran.r\u2010project.org\/web\/packages\/glmnet\/vignettes\/glmnet_beta.pdf."},{"key":"e_1_2_10_32_1","unstructured":"ZouH HastieT. Elastic\u2010net for sparse estimation and sparse PCA package \u201celasticnet\u201d; August 31 2018."},{"key":"e_1_2_10_33_1","doi-asserted-by":"publisher","DOI":"10.18637\/jss.v033.i01"},{"key":"e_1_2_10_34_1","doi-asserted-by":"publisher","DOI":"10.18637\/jss.v011.i09"},{"key":"e_1_2_10_35_1","unstructured":"MilborrowS. Multivariate adaptive regression splines package \u201cearth\u201d; November 9 2019."},{"key":"e_1_2_10_36_1","unstructured":"LiawA WienerM. \u201crandomForest\u201d: Breiman and Cutler's random forests for classification and regression. R\u2009package\u2009version 4.6\u201314; March 25 2018."},{"key":"e_1_2_10_37_1","unstructured":"KuhnM WestonS KeeferC CoulterN QuinlanR. \u201cCubist\u201d: rule\u2010 and instance\u2010based regression modeling Package; January 10 2020.https:\/\/topepo.github.io\/Cubist."},{"key":"e_1_2_10_38_1","unstructured":"HothornT HornikK StroblC ZeileisA. A laboratory for recursive partytioning. Package \u201cParty\u201d; March 5 2020."},{"key":"e_1_2_10_39_1","unstructured":"ChenT HeT BenestyM et al. Extreme gradient boosting Package \u201cxgboost\u201d; August 1 2019."},{"key":"e_1_2_10_40_1","unstructured":"PAPI(Performance API).http:\/\/icl.cs.utk.edu\/papi\/"},{"key":"e_1_2_10_41_1","unstructured":"perf_event.https:\/\/perf.wiki.kernel.org\/index.php\/Main_Page."},{"key":"e_1_2_10_42_1","unstructured":"PerfMon2.http:\/\/perfmon2.sourceforge.net."},{"key":"e_1_2_10_43_1","unstructured":"PospiechC. Hardware performance monitor (HPM) toolkit users guide Advanced Computer Technologies\u00a0Center IBM Research; June2008."},{"key":"e_1_2_10_44_1","unstructured":"PyPAPI.https:\/\/flozz.github.io\/pypapi"},{"key":"e_1_2_10_45_1","doi-asserted-by":"publisher","DOI":"10.1214\/aos\/1013203451"},{"key":"e_1_2_10_46_1","unstructured":"Keras: the python deep learning library.https:\/\/keras.io\/#keras\u2010the\u2010python\u2010deep\u2010learning\u2010library."},{"key":"e_1_2_10_47_1","unstructured":"TensorFlow.https:\/\/www.tensorflow.org."},{"key":"e_1_2_10_48_1","unstructured":"Horovod: a distributed training framework for TensorFlow.https:\/\/github.com\/uber\/horovod."},{"key":"e_1_2_10_49_1","unstructured":"XLA architecture.https:\/\/www.tensorflow.org\/xla\/architecture."},{"key":"e_1_2_10_50_1","unstructured":"Pandas.https:\/\/pandas.pydata.org\/pandas\u2010docs\/stable\/."},{"key":"e_1_2_10_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/MCAS.2006.1688199"},{"key":"e_1_2_10_52_1","unstructured":"MuMMI project.2019.http:\/\/www.mummi.org\/info"},{"volume-title":"Tools for High Performance Computing","year":"2014","author":"Wu X","key":"e_1_2_10_53_1"}],"container-title":["Concurrency and Computation: Practice and Experience"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1002\/cpe.6516","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/full-xml\/10.1002\/cpe.6516","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/am-pdf\/10.1002\/cpe.6516","content-type":"application\/pdf","content-version":"am","intended-application":"syndication"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1002\/cpe.6516","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,8,20]],"date-time":"2023-08-20T18:21:00Z","timestamp":1692555660000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1002\/cpe.6516"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,7,22]]},"references-count":52,"journal-issue":{"issue":"15","published-print":{"date-parts":[[2023,7,10]]}},"alternative-id":["10.1002\/cpe.6516"],"URL":"https:\/\/doi.org\/10.1002\/cpe.6516","archive":["Portico"],"relation":{},"ISSN":["1532-0626","1532-0634"],"issn-type":[{"type":"print","value":"1532-0626"},{"type":"electronic","value":"1532-0634"}],"subject":[],"published":{"date-parts":[[2021,7,22]]},"assertion":[{"value":"2020-12-18","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-06-30","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-07-22","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e6516"}}