{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,22]],"date-time":"2026-07-22T09:44:18Z","timestamp":1784713458911,"version":"3.55.0"},"reference-count":366,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2022,3,4]],"date-time":"2022-03-04T00:00:00Z","timestamp":1646352000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Des. Autom. Electron. Syst."],"published-print":{"date-parts":[[2022,5,31]]},"abstract":"<jats:p>Deep neural networks (DNNs) have achieved unprecedented success in the field of artificial intelligence (AI), including computer vision, natural language processing, and speech recognition. However, their superior performance comes at the considerable cost of computational complexity, which greatly hinders their applications in many resource-constrained devices, such as mobile phones and Internet of Things (IoT) devices. Therefore, methods and techniques that are able to lift the efficiency bottleneck while preserving the high accuracy of DNNs are in great demand to enable numerous edge AI applications. This article provides an overview of efficient deep learning methods, systems, and applications. We start from introducing popular model compression methods, including pruning, factorization, quantization, as well as compact model design. To reduce the large design cost of these manual solutions, we discuss the AutoML framework for each of them, such as neural architecture search (NAS) and automated pruning and quantization. We then cover efficient on-device training to enable user customization based on the local data on mobile devices. Apart from general acceleration techniques, we also showcase several task-specific accelerations for point cloud, video, and natural language processing by exploiting their spatial sparsity and temporal\/token redundancy. Finally, to support all these algorithmic advancements, we introduce the efficient deep learning system design from both software and hardware perspectives.<\/jats:p>","DOI":"10.1145\/3486618","type":"journal-article","created":{"date-parts":[[2022,3,4]],"date-time":"2022-03-04T09:54:32Z","timestamp":1646387672000},"page":"1-50","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":128,"title":["Enable Deep Learning on Mobile Devices: Methods, Systems, and Applications"],"prefix":"10.1145","volume":"27","author":[{"given":"Han","family":"Cai","sequence":"first","affiliation":[{"name":"Massachusetts Institute of Technology, Cambridge, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ji","family":"Lin","sequence":"additional","affiliation":[{"name":"Massachusetts Institute of Technology, Cambridge, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yujun","family":"Lin","sequence":"additional","affiliation":[{"name":"Massachusetts Institute of Technology, Cambridge, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhijian","family":"Liu","sequence":"additional","affiliation":[{"name":"Massachusetts Institute of Technology, Cambridge, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Haotian","family":"Tang","sequence":"additional","affiliation":[{"name":"Massachusetts Institute of Technology, Cambridge, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hanrui","family":"Wang","sequence":"additional","affiliation":[{"name":"Massachusetts Institute of Technology, Cambridge, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ligeng","family":"Zhu","sequence":"additional","affiliation":[{"name":"Massachusetts Institute of Technology, Cambridge, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Song","family":"Han","sequence":"additional","affiliation":[{"name":"Massachusetts Institute of Technology, Cambridge, MA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,3,4]]},"reference":[{"key":"e_1_3_1_2_2","volume-title":"USENIX Symposium on Operating Systems Design and Implementation","author":"Abadi Mart\u00edn","year":"2016","unstructured":"Mart\u00edn Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. 2016. TensorFlow: A system for large-scale machine learning. In USENIX Symposium on Operating Systems Design and Implementation."},{"key":"e_1_3_1_3_2","article-title":"A2P-MANN: Adaptive attention inference hops pruned memory-augmented neural networks","author":"Ahmadzadeh Mohsen","year":"2021","unstructured":"Mohsen Ahmadzadeh, Mehdi Kamal, Ali Afzali-Kusha, and Massoud Pedram. 2021. A2P-MANN: Adaptive attention inference hops pruned memory-augmented neural networks. arXiv preprint arXiv:2101.09693 (2021).","journal-title":"arXiv preprint arXiv:2101.09693"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1122"},{"key":"e_1_3_1_5_2","volume-title":"International Symposium on Computer Architecture","author":"Albericio Jorge","year":"2016","unstructured":"Jorge Albericio, Patrick Judd, Tayler H. Hetherington, Tor M. Aamodt, Natalie D. Enright Jerger, and Andreas Moshovos. 2016. Cnvlutin: Ineffectual-neuron-free deep neural network computing. In International Symposium on Computer Architecture."},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2019.8793579"},{"key":"e_1_3_1_7_2","volume-title":"Symposium on VLSI Circuits","author":"Ando Kota","year":"2017","unstructured":"Kota Ando, Kodai Ueyoshi, Kentaro Orimo, Haruyoshi Yonekawa, Shimpei Sato, Hiroki Nakahara, Masayuki Ikebe, Tetsuya Asai, Shinya Takamaeda-Yamazaki, Tadahiro Kuroda, and Masato Motomur. 2017. BRein memory: A 13-layer 4.2 K Neuron\/0.8 M synapse binary\/ternary reconfigurable in-memory deep neural network accelerator in 65 nm CMOS. In Symposium on VLSI Circuits."},{"key":"e_1_3_1_8_2","volume-title":"IEEE Computer Society Annual Symposium on VLSI","author":"Andri Renzo","year":"2016","unstructured":"Renzo Andri, Lukas Cavigelli, Davide Rossi, and Luca Benini. 2016. YodaNN: An ultra-low power convolutional neural network accelerator based on binary weights. In IEEE Computer Society Annual Symposium on VLSI."},{"key":"e_1_3_1_9_2","article-title":"Joint 2D-3D-semantic data for indoor scene understanding","author":"Armeni Iro","year":"2017","unstructured":"Iro Armeni, Alexandar Sax, Amir R. Zamir, and Silvio Savarese. 2017. Joint 2D-3D-semantic data for indoor scene understanding. arXiv preprint arXiv:1702.01105 (2017).","journal-title":"arXiv preprint arXiv:1702.01105"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.170"},{"key":"e_1_3_1_11_2","volume-title":"International Conference on Learning Representations","author":"Bahdanau Dzmitry","year":"2015","unstructured":"Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015. Neural machine translation by jointly learning to align and translate. In International Conference on Learning Representations."},{"key":"e_1_3_1_12_2","volume-title":"Conference of the Association for Computational Linguistics","author":"Bai Haoli","year":"2020","unstructured":"Haoli Bai, Wei Zhang, Lu Hou, Lifeng Shang, Jing Jin, Xin Jiang, Qun Liu, Michael Lyu, and Irwin King. 2020. BinaryBERT: Pushing the limit of BERT quantization. In Conference of the Association for Computational Linguistics."},{"key":"e_1_3_1_13_2","volume-title":"International Conference on Learning Representations","author":"Baker Bowen","year":"2017","unstructured":"Bowen Baker, Otkrist Gupta, Nikhil Naik, and Ramesh Raskar. 2017. Designing neural network architectures using reinforcement learning. In International Conference on Learning Representations."},{"key":"e_1_3_1_14_2","volume-title":"Conference on Machine Learning and Systems","author":"Banbury Colby","year":"2020","unstructured":"Colby Banbury, Chuteng Zhou, Igor Fedorov, Ramon Matas Navarro, Urmish Thakkar, Dibakar Gope, Vijay Janapa Reddi, Matthew Mattina, and Paul N. Whatmough. 2020. MicroNets: Neural network architectures for deploying TinyML applications on commodity microcontrollers. In Conference on Machine Learning and Systems."},{"key":"e_1_3_1_15_2","volume-title":"Conference on Neural Information Processing Systems","author":"Banner Ron","year":"2019","unstructured":"Ron Banner, Yury Nahshan, and Daniel Soudry. 2019. Post training 4-bit quantization of convolutional networks for rapid-deployment. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1338"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00939"},{"key":"e_1_3_1_18_2","article-title":"Longformer: The long-document transformer","author":"Beltagy Iz","year":"2020","unstructured":"Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020. Longformer: The long-document transformer. arXiv preprint arXiv:2004.05150 (2020).","journal-title":"arXiv preprint arXiv:2004.05150"},{"key":"e_1_3_1_19_2","article-title":"Estimating or propagating gradients through stochastic neurons for conditional computation","author":"Bengio Yoshua","year":"2013","unstructured":"Yoshua Bengio, Nicholas L\u00e9onard, and Aaron Courville. 2013. Estimating or propagating gradients through stochastic neurons for conditional computation. arXiv preprint arXiv:1308.3432 (2013).","journal-title":"arXiv preprint arXiv:1308.3432"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.331"},{"key":"e_1_3_1_21_2","article-title":"End to end learning for self-driving cars","author":"Bojarski Mariusz","year":"2016","unstructured":"Mariusz Bojarski, Davide Del Testa, Daniel Dworakowski, Bernhard Firner, Beat Flepp, Prasoon Goyal, Lawrence D. Jackel, Mathew Monfort, Urs Muller, Jiakai Zhang, Xin Zhang, Jake Zhao, and Karol Zieba. 2016. End to end learning for self-driving cars. arXiv preprint arXiv:1604.07316 (2016).","journal-title":"arXiv preprint arXiv:1604.07316"},{"key":"e_1_3_1_22_2","volume-title":"International Conference on Learning Representations","author":"Brock Andrew","year":"2018","unstructured":"Andrew Brock, Theodore Lim, James M. Ritchie, and Nick Weston. 2018. SMASH: One-shot model architecture search through HyperNetworks. In International Conference on Learning Representations."},{"key":"e_1_3_1_23_2","volume-title":"Conference on Neural Information Processing Systems","author":"Brown Tom B.","year":"2020","unstructured":"Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/1150402.1150464"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11709"},{"key":"e_1_3_1_26_2","volume-title":"International Conference on Learning Representations","author":"Cai Han","year":"2020","unstructured":"Han Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang, and Song Han. 2020. Once for all: Train one network and specialize it for efficient deployment. In International Conference on Learning Representations."},{"key":"e_1_3_1_27_2","volume-title":"Conference on Neural Information Processing Systems","author":"Cai Han","year":"2020","unstructured":"Han Cai, Chuang Gan, Ligeng Zhu, and Song Han. 2020. TinyTL: Reduce memory, not parameters for efficient on-device learning. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_28_2","article-title":"AutoML for architecting efficient and specialized neural networks","author":"Cai Han","year":"2019","unstructured":"Han Cai, Ji Lin, Yujun Lin, Zhijian Liu, Kuan Wang, Tianzhe Wang, Ligeng Zhu, and Song Han. 2019. AutoML for architecting efficient and specialized neural networks. IEEE Micro 40, 1 (2019), 75\u201382.","journal-title":"IEEE Micro"},{"key":"e_1_3_1_29_2","volume-title":"International Conference on Machine Learning","author":"Cai Han","year":"2018","unstructured":"Han Cai, Jiacheng Yang, Weinan Zhang, Song Han, and Yong Yu. 2018. Path-level network transformation for efficient architecture search. In International Conference on Machine Learning."},{"key":"e_1_3_1_30_2","volume-title":"International Conference on Learning Representations","author":"Cai Han","year":"2019","unstructured":"Han Cai, Ligeng Zhu, and Song Han. 2019. ProxylessNAS: Direct neural architecture search on target task and hardware. In International Conference on Learning Representations."},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01318"},{"key":"e_1_3_1_32_2","article-title":"Expanding the reach of federated learning by reducing client resource requirements","author":"Caldas Sebastian","year":"2018","unstructured":"Sebastian Caldas, Jakub Kone\u010dny, H. Brendan McMahan, and Ameet Talwalkar. 2018. Expanding the reach of federated learning by reducing client resource requirements. arXiv preprint arXiv:1812.07210 (2018).","journal-title":"arXiv preprint arXiv:1812.07210"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.411"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.502"},{"key":"e_1_3_1_35_2","volume-title":"Great Lakes Symposium on VLSI","author":"Cavigelli Lukas","year":"2015","unstructured":"Lukas Cavigelli, David Gschwend, Christoph Mayer, Samuel Willi, Beat Muheim, and Luca Benini. 2015. Origami: A convolutional network accelerator. In Great Lakes Symposium on VLSI."},{"key":"e_1_3_1_36_2","volume-title":"International Symposium on Computer Architecture","author":"Chakradhar Srimat","year":"2010","unstructured":"Srimat Chakradhar, Murugan Sankaradas, Venkata Jakkula, and Srihari Cadambi. 2010. A dynamically configurable coprocessor for convolutional neural networks. In International Symposium on Computer Architecture."},{"key":"e_1_3_1_37_2","article-title":"ShapeNet: An information-rich 3D model repository","author":"Chang Angel X.","year":"2015","unstructured":"Angel X. Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, Jianxiong Xiao, Li Yi, and Fisher Yu. 2015. ShapeNet: An information-rich 3D model repository. arXiv preprint arXiv:1512.03012 (2015).","journal-title":"arXiv preprint arXiv:1512.03012"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.5244\/C.28.6"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2020\/341"},{"key":"e_1_3_1_40_2","volume-title":"Conference on Neural Information Processing Systems","author":"Chen Guobin","year":"2017","unstructured":"Guobin Chen, Wongun Choi, Xiang Yu, Tony Han, and Manmohan Chandraker. 2017. Learning efficient object detection models with knowledge distillation. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_41_2","volume-title":"Conference on Neural Information Processing Systems","author":"Chen Liang-Chieh","year":"2018","unstructured":"Liang-Chieh Chen, Maxwell Collins, Yukun Zhu, George Papandreou, Barret Zoph, Florian Schroff, Hartwig Adam, and Jon Shlens. 2018. Searching for efficient multi-scale architectures for dense image prediction. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/2541940.2541967"},{"key":"e_1_3_1_43_2","article-title":"MXNet: A flexible and efficient machine learning library for heterogeneous distributed systems","author":"Chen Tianqi","year":"2015","unstructured":"Tianqi Chen, Mu Li, Yutian Li, Min Lin, Naiyan Wang, Minjie Wang, Tianjun Xiao, Bing Xu, Chiyuan Zhang, and Zheng Zhang. 2015. MXNet: A flexible and efficient machine learning library for heterogeneous distributed systems. arXiv preprint arXiv:1512.01274 (2015).","journal-title":"arXiv preprint arXiv:1512.01274"},{"key":"e_1_3_1_44_2","volume-title":"USENIX Symposium on Operating Systems Design and Implementation","author":"Chen Tianqi","year":"2018","unstructured":"Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018. TVM: An automated end-to-end optimizing compiler for deep learning. In USENIX Symposium on Operating Systems Design and Implementation."},{"key":"e_1_3_1_45_2","article-title":"Training deep nets with sublinear memory cost","author":"Chen Tianqi","year":"2016","unstructured":"Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. 2016. Training deep nets with sublinear memory cost. arXiv preprint arXiv:1604.06174 (2016).","journal-title":"arXiv preprint arXiv:1604.06174"},{"key":"e_1_3_1_46_2","volume-title":"International Symposium on Microarchitecture","author":"Chen Yunji","year":"2014","unstructured":"Yunji Chen, Tao Luo, Shaoli Liu, Shijin Zhang, Liqiang He, Jia Wang, Ling Li, Tianshi Chen, Zhiwei Xu, Ninghui Sun, and Olivier Temam. 2014. DaDianNao: A machine-learning supercomputer. In International Symposium on Microarchitecture."},{"key":"e_1_3_1_47_2","volume-title":"Conference on Neural Information Processing Systems","author":"Chen Yukang","year":"2019","unstructured":"Yukang Chen, Tong Yang, Xiangyu Zhang, Gaofeng Meng, Xinyu Xiao, and Jian Sun. 2019. DetNAS: Backbone search for object detection. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/3340531.3412688"},{"key":"e_1_3_1_49_2","article-title":"Eyeriss: An energy-efficient reconfigurable accelerator for deep convolutional neural networks","author":"Chen Yu-Hsin","year":"2017","unstructured":"Yu-Hsin Chen, Tushar Krishna, Joel S. Emer, and Vivienne Sze. 2017. Eyeriss: An energy-efficient reconfigurable accelerator for deep convolutional neural networks. Int. J. Space-based Situat. Comput. 52, 1 (2017), 127\u2013138.","journal-title":"Int. J. Space-based Situat. Comput."},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/JETCAS.2019.2910232"},{"key":"e_1_3_1_51_2","article-title":"A survey of model compression and acceleration for deep neural networks","author":"Cheng Yu","year":"2017","unstructured":"Yu Cheng, Duo Wang, Pan Zhou, and Tao Zhang. 2017. A survey of model compression and acceleration for deep neural networks. IEEE Sig. Process. Mag. 35, 1 (2017), 126\u2013136.","journal-title":"IEEE Sig. Process. Mag."},{"key":"e_1_3_1_52_2","volume-title":"transformers.zip: Compressing Transformers with Pruning and Quantization","author":"Cheong Robin","year":"2019","unstructured":"Robin Cheong and Robel Daniel. 2019. transformers.zip: Compressing Transformers with Pruning and Quantization. Technical Report. Stanford University, Stanford, CA."},{"key":"e_1_3_1_53_2","article-title":"cuDNN: Efficient primitives for deep learning","author":"Chetlur Sharan","year":"2014","unstructured":"Sharan Chetlur, Cliff Woolley, Philippe Vandermersch, Jonathan Cohen, John Tran, Bryan Catanzaro, and Evan Shelhamer. 2014. cuDNN: Efficient primitives for deep learning. arXiv preprint arXiv:1410.0759 (2014).","journal-title":"arXiv preprint arXiv:1410.0759"},{"key":"e_1_3_1_54_2","article-title":"Generating long sequences with sparse transformers","author":"Child Rewon","year":"2019","unstructured":"Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019. Generating long sequences with sparse transformers. arXiv preprint arXiv:1904.10509 (2019).","journal-title":"arXiv preprint arXiv:1904.10509"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10462-020-09816-7"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00259"},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00319"},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00905"},{"key":"e_1_3_1_59_2","volume-title":"International Conference on Medical Image Computing and Computer-assisted Intervention","author":"Cicek Ozgun","year":"2016","unstructured":"Ozgun Cicek, Ahmed Abdulkadir, Soeren S. Lienkamp, Thomas Brox, and Olaf Ronneberger. 2016. 3D U-Net: Learning dense volumetric segmentation from sparse annotation. In International Conference on Medical Image Computing and Computer-assisted Intervention."},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2018.8460487"},{"key":"e_1_3_1_61_2","volume-title":"IEEE Symposium on Field-programmable Custom Computing Machines","author":"Cong Jason","year":"2018","unstructured":"Jason Cong, Zhenman Fang, Michael Lo, Hanrui Wang, Jingxian Xu, and Shaochong Zhang. 2018. Understanding performance differences of FPGAs and GPUs. In IEEE Symposium on Field-programmable Custom Computing Machines."},{"key":"e_1_3_1_62_2","volume-title":"Conference on Neural Information Processing Systems","author":"Courbariaux Matthieu","year":"2015","unstructured":"Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David. 2015. BinaryConnect: Training deep neural networks with binary weights during propagations. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_63_2","article-title":"Binarized neural networks: Training deep neural networks with weights and activations constrained to +1 or \u20131","author":"Courbariaux Matthieu","year":"2016","unstructured":"Matthieu Courbariaux, Itay Hubara, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio. 2016. Binarized neural networks: Training deep neural networks with weights and activations constrained to +1 or \u20131. arXiv preprint arXiv:1602.02830 (2016).","journal-title":"arXiv preprint arXiv:1602.02830"},{"key":"e_1_3_1_64_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00432"},{"key":"e_1_3_1_65_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.261"},{"key":"e_1_3_1_66_2","doi-asserted-by":"publisher","DOI":"10.1109\/DAC18072.2020.9218710"},{"key":"e_1_3_1_67_2","doi-asserted-by":"publisher","DOI":"10.1145\/3361682"},{"key":"e_1_3_1_68_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_1_69_2","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2020.2976475"},{"key":"e_1_3_1_70_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2013.6639345"},{"key":"e_1_3_1_71_2","volume-title":"Conference on Neural Information Processing Systems","author":"Denton Emily L.","year":"2014","unstructured":"Emily L. Denton, Wojciech Zaremba, Joan Bruna, Yann LeCun, and Rob Fergus. 2014. Exploiting linear structure within convolutional networks for efficient evaluation. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_72_2","volume-title":"Conference of the North American Chapter of the Association for Computational Linguistics","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. BERT: Pre-training of deep bidirectional transformers for language understanding. In Conference of the North American Chapter of the Association for Computational Linguistics."},{"key":"e_1_3_1_73_2","article-title":"IOS: Inter-operator scheduler for CNN acceleration","author":"Ding Yaoyao","year":"2020","unstructured":"Yaoyao Ding, Ligeng Zhu, Zhihao Jia, Gennady Pekhimenko, and Song Han. 2020. IOS: Inter-operator scheduler for CNN acceleration. arXiv preprint arXiv:2011.01302 (2020).","journal-title":"arXiv preprint arXiv:2011.01302"},{"key":"e_1_3_1_74_2","volume-title":"International Conference on Machine Learning","author":"Donahue Jeff","year":"2014","unstructured":"Jeff Donahue, Yangqing Jia, Oriol Vinyals, Judy Hoffman, Ning Zhang, Eric Tzeng, and Trevor Darrell. 2014. DeCAF: A deep convolutional activation feature for generic visual recognition. In International Conference on Machine Learning."},{"key":"e_1_3_1_75_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00038"},{"key":"e_1_3_1_76_2","volume-title":"International Symposium on Computer Architecture","author":"Du Zidong","year":"2015","unstructured":"Zidong Du, Robert Fasthuber, Tianshi Chen, Paolo Ienne, Ling Li, Tao Luo, Xiaobing Feng, Yunji Chen, and Olivier Temam. 2015. ShiDianNao: Shifting vision processing closer to the sensor. In International Symposium on Computer Architecture."},{"key":"e_1_3_1_77_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cpc.2007.11.014"},{"key":"e_1_3_1_78_2","volume-title":"International Conference on Learning Representations","author":"Elsken Thomas","year":"2018","unstructured":"Thomas Elsken, Jan Hendrik Metzen, and Frank Hutter. 2018. Efficient multi-objective neural architecture search via Lamarckian evolution. In International Conference on Learning Representations."},{"key":"e_1_3_1_79_2","article-title":"Neural architecture search: A survey","author":"Elsken Thomas","year":"2019","unstructured":"Thomas Elsken, Jan Hendrik Metzen, and Frank Hutter. 2019. Neural architecture search: A survey. J. Mach. Learn. Res. 20, 55 (2019), 1\u201321.","journal-title":"J. Mach. Learn. Res."},{"key":"e_1_3_1_80_2","doi-asserted-by":"publisher","DOI":"10.1109\/72.963775"},{"key":"e_1_3_1_81_2","volume-title":"Conference on Neural Information Processing Systems","author":"Fan Quanfu","year":"2019","unstructured":"Quanfu Fan, Chun-Fu Chen, Hilde Kuehne, Marco Pistoia, and David Cox. 2019. More is less: Learning efficient video representations by big-little network and depthwise temporal aggregation. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_82_2","volume-title":"Conference on Neural Information Processing Systems","author":"Fedorov Igor","year":"2019","unstructured":"Igor Fedorov, Ryan P. Adams, Matthew Mattina, and Paul N. Whatmough. 2019. SpArSe: Sparse architecture search for CNNs on resource-constrained microcontrollers. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_83_2","article-title":"Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity","author":"Fedus William","year":"2021","unstructured":"William Fedus, Barret Zoph, and Noam Shazeer. 2021. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. arXiv preprint arXiv:2101.03961 (2021).","journal-title":"arXiv preprint arXiv:2101.03961"},{"key":"e_1_3_1_84_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00028"},{"key":"e_1_3_1_85_2","volume-title":"Conference on Neural Information Processing Systems","author":"Feichtenhofer Christoph","year":"2016","unstructured":"Christoph Feichtenhofer, Axel Pinz, and Richard Wildes. 2016. Spatiotemporal residual networks for video action recognition. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_86_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.213"},{"key":"e_1_3_1_87_2","volume-title":"International Conference on Learning Representations","author":"Frankle Jonathan","year":"2018","unstructured":"Jonathan Frankle and Michael Carbin. 2018. The lottery ticket hypothesis: Finding sparse, trainable neural networks. In International Conference on Learning Representations."},{"key":"e_1_3_1_88_2","volume-title":"International Conference on Machine Learning","author":"Frankle Jonathan","year":"2020","unstructured":"Jonathan Frankle, Gintare Karolina Dziugaite, Daniel Roy, and Michael Carbin. 2020. Linear mode connectivity and the lottery ticket hypothesis. In International Conference on Machine Learning."},{"key":"e_1_3_1_89_2","article-title":"Training BatchNorm and only BatchNorm: On the expressive power of random features in CNNs","author":"Frankle Jonathan","year":"2020","unstructured":"Jonathan Frankle, David J. Schwab, and Ari S. Morcos. 2020. Training BatchNorm and only BatchNorm: On the expressive power of random features in CNNs. arXiv preprint arXiv:2003.00152 (2020).","journal-title":"arXiv preprint arXiv:2003.00152"},{"key":"e_1_3_1_90_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298872"},{"key":"e_1_3_1_91_2","doi-asserted-by":"publisher","DOI":"10.1145\/3037697.3037702"},{"key":"e_1_3_1_92_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00720"},{"key":"e_1_3_1_93_2","doi-asserted-by":"publisher","DOI":"10.1109\/72.317740"},{"key":"e_1_3_1_94_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.169"},{"key":"e_1_3_1_95_2","unstructured":"Gene H. Golub and Charles F. Van Loan. 1996. Matrix Computations . Stanford University."},{"key":"e_1_3_1_96_2","article-title":"Compressing deep convolutional networks using vector quantization","author":"Gong Yunchao","year":"2014","unstructured":"Yunchao Gong, Liu Liu, Ming Yang, and Lubomir Bourdev. 2014. Compressing deep convolutional networks using vector quantization. arXiv preprint arXiv:1412.6115 (2014).","journal-title":"arXiv preprint arXiv:1412.6115"},{"key":"e_1_3_1_97_2","article-title":"Compressing BERT: Studying the effects of weight pruning on transfer learning","author":"Gordon Mitchell A.","year":"2020","unstructured":"Mitchell A. Gordon, Kevin Duh, and Nicholas Andrews. 2020. Compressing BERT: Studying the effects of weight pruning on transfer learning. arXiv preprint arXiv:2002.08307 (2020).","journal-title":"arXiv preprint arXiv:2002.08307"},{"key":"e_1_3_1_98_2","volume-title":"International Conference on Machine Learning","author":"al Saurabh Goyal et","year":"2020","unstructured":"Saurabh Goyal et al. 2020. PoWER-BERT: Accelerating BERT inference for classification tasks. In International Conference on Machine Learning."},{"key":"e_1_3_1_99_2","doi-asserted-by":"publisher","DOI":"10.5244\/C.29.150"},{"key":"e_1_3_1_100_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00961"},{"key":"e_1_3_1_101_2","doi-asserted-by":"publisher","DOI":"10.1109\/FPL.2016.7577352"},{"key":"e_1_3_1_102_2","volume-title":"Conference on Neural Information Processing Systems","author":"Gruslys Audrunas","year":"2016","unstructured":"Audrunas Gruslys, R\u00e9mi Munos, Ivo Danihelka, Marc Lanctot, and Alex Graves. 2016. Memory-efficient backpropagation through time. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_103_2","volume-title":"International Conference on Learning Representations","author":"Gu Jiatao","year":"2018","unstructured":"Jiatao Gu, James Bradbury, Caiming Xiong, Victor O. K. Li, and Richard Socher. 2018. Non-autoregressive neural machine translation. In International Conference on Learning Representations."},{"key":"e_1_3_1_104_2","volume-title":"Conference on Neural Information Processing Systems","author":"Gu Jiatao","year":"2019","unstructured":"Jiatao Gu, Changhan Wang, and Jake Zhao. 2019. Levenshtein transformer. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_105_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCD50377.2020.00047"},{"key":"e_1_3_1_106_2","volume-title":"Conference on Neural Information Processing Systems","author":"Guo Yiwen","year":"2016","unstructured":"Yiwen Guo, Anbang Yao, and Yurong Chen. 2016. Dynamic network surgery for efficient DNNs. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_107_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58517-4_32"},{"key":"e_1_3_1_108_2","volume-title":"International Conference on Machine Learning","author":"Gupta Suyog","year":"2015","unstructured":"Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan. 2015. Deep learning with limited numerical precision. In International Conference on Machine Learning."},{"key":"e_1_3_1_109_2","volume-title":"IEEE International Symposium on High-performance Computer Architecture","author":"Ham Tae Jun","year":"2020","unstructured":"Tae Jun Ham, Sung Jun Jung, Seonghak Kim, Young H. Oh, Yeonhong Park, Yoonho Song, Jung-Hun Park, Sanghee Lee, Kyoung Park, Jae W. Lee, and Deog-Kyoon Jeong. 2020. A3: Accelerating attention mechanisms in neural networks with approximation. In IEEE International Symposium on High-performance Computer Architecture."},{"key":"e_1_3_1_110_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.668"},{"key":"e_1_3_1_111_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00301"},{"key":"e_1_3_1_112_2","doi-asserted-by":"publisher","DOI":"10.5555\/AAI28115679"},{"key":"e_1_3_1_113_2","volume-title":"ACM\/SIGDA International Symposium on Field-programmable Gate Arrays","author":"Han Song","year":"2017","unstructured":"Song Han, Junlong Kang, Huizi Mao, Yiming Hu, Xin Li, Yubin Li, Dongliang Xie, Hong Luo, Song Yao, Yu Wang, Huazhong Yang, and William J. Dally. 2017. ESE: Efficient speech recognition engine with sparse LSTM on FPGA. In ACM\/SIGDA International Symposium on Field-programmable Gate Arrays."},{"key":"e_1_3_1_114_2","volume-title":"International Symposium on Computer Architecture","author":"Han Song","year":"2016","unstructured":"Song Han, Xingyu Liu, Huizi Mao, Jing Pu, Ardavan Pedram, Mark A. Horowitz, and William J. Dally. 2016. EIE: Efficient inference engine on compressed deep neural network. In International Symposium on Computer Architecture."},{"key":"e_1_3_1_115_2","volume-title":"International Conference on Learning Representations","author":"Han Song","year":"2016","unstructured":"Song Han, Huizi Mao, and William J. Dally. 2016. Deep compression: Compressing deep neural networks with pruning, trained quantization and Huffman coding. In International Conference on Learning Representations."},{"key":"e_1_3_1_116_2","volume-title":"Conference on Neural Information Processing Systems","author":"Han Song","year":"2015","unstructured":"Song Han, Jeff Pool, John Tran, and William J. Dally. 2015. Learning both weights and connections for efficient neural networks. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_117_2","volume-title":"Machine Learning for Healthcare Conference","author":"Haque Albert","year":"2017","unstructured":"Albert Haque, Michelle Guo, Alexandre Alahi, Serena Yeung, Zelun Luo, Alisha Rege, Jeffrey Jopling, Lance Downing, William Beninati, Amit Singh, Terry Platchek, Arnold Milstein, and Li Fei-Fei. 2017. Towards vision-based smart hospitals: A system for tracking and monitoring hand hygiene compliance. In Machine Learning for Healthcare Conference."},{"key":"e_1_3_1_118_2","volume-title":"Conference on Neural Information Processing Systems","author":"Hassibi Babak","year":"1993","unstructured":"Babak Hassibi and David G. Stork. 1993. Second order derivatives for network pruning: Optimal brain surgeon. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_119_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_1_120_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2020.106622"},{"key":"e_1_3_1_121_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01234-2_48"},{"key":"e_1_3_1_122_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.155"},{"key":"e_1_3_1_123_2","doi-asserted-by":"publisher","DOI":"10.1109\/MSP.2012.2205597"},{"key":"e_1_3_1_124_2","article-title":"Distilling the knowledge in a neural network","author":"Hinton Geoffrey","year":"2015","unstructured":"Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531 (2015).","journal-title":"arXiv preprint arXiv:1503.02531"},{"key":"e_1_3_1_125_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00140"},{"key":"e_1_3_1_126_2","article-title":"MobileNets: Efficient convolutional neural networks for mobile vision applications","author":"Howard Andrew G.","year":"2017","unstructured":"Andrew G. Howard, Menglong Zhu, Bo Chen, Dimitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. 2017. MobileNets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861 (2017).","journal-title":"arXiv preprint arXiv:1704.04861"},{"key":"e_1_3_1_127_2","volume-title":"Conference on Neural Information Processing Systems","author":"Hu Baotian","year":"2014","unstructured":"Baotian Hu, Zhengdong Lu, Hang Li, and Qingcai Chen. 2014. Convolutional neural network architectures for matching natural language sentences. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_128_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01112"},{"key":"e_1_3_1_129_2","volume-title":"International Conference on Learning Representations","author":"Huang Gao","year":"2018","unstructured":"Gao Huang, Danlu Chen, Tianhong Li, Felix Wu, Laurens van der Maaten, and Kilian Q. Weinberger. 2018. Multi-scale dense networks for resource efficient image classification. In International Conference on Learning Representations."},{"key":"e_1_3_1_130_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00278"},{"key":"e_1_3_1_131_2","article-title":"SqueezeNet: AlexNet-level accuracy with 50 \\( \\times \\)  fewer parameters and  \\( \\lt \\) 0.5MB model size","author":"Iandola Forrest N.","year":"2016","unstructured":"Forrest N. Iandola, Song Han, Matthew W. Moskewicz, Khalid Ashraf, William J. Dally, and Kurt Keutzer. 2016. SqueezeNet: AlexNet-level accuracy with 50 \\( \\times \\) fewer parameters and \\( \\lt \\) 0.5MB model size. arXiv preprint arXiv:1602.07360 (2016).","journal-title":"arXiv preprint arXiv:1602.07360"},{"key":"e_1_3_1_132_2","article-title":"AI benchmark: Running deep neural networks on Android smartphones","author":"Ignatov Andrey","year":"2018","unstructured":"Andrey Ignatov, Radu Timofte, William Chou, Ke Wang, Max Wu, Tim Hartley, and Luc Van Gool. 2018. AI benchmark: Running deep neural networks on Android smartphones. arXiv preprint arXiv:1810.01109 (2018).","journal-title":"arXiv preprint arXiv:1810.01109"},{"key":"e_1_3_1_133_2","volume-title":"International Conference on Machine Learning","author":"Ioffe Sergey","year":"2015","unstructured":"Sergey Ioffe and Christian Szegedy. 2015. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International Conference on Machine Learning."},{"key":"e_1_3_1_134_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00286"},{"key":"e_1_3_1_135_2","doi-asserted-by":"publisher","DOI":"10.5244\/C.28.88"},{"key":"e_1_3_1_136_2","article-title":"The algorithms for FPGA implementation of sparse matrices multiplication","author":"Jamro Ernest","year":"2015","unstructured":"Ernest Jamro, Tomasz Pabi\u015b, Pawe\u0142 Russek, and Kazimierz Wiatr. 2015. The algorithms for FPGA implementation of sparse matrices multiplication. Comput. Inform. 33, 3 (2015), 667\u2013684.","journal-title":"Comput. Inform."},{"key":"e_1_3_1_137_2","volume-title":"International Symposium on Computer Architecture","author":"Jang Hanhwi","year":"2019","unstructured":"Hanhwi Jang, Joonsung Kim, Jae-Eon Jo, Jaewon Lee, and Jangwoo Kim. 2019. MnnFast: A fast and scalable system architecture for memory-augmented neural networks. In International Symposium on Computer Architecture."},{"key":"e_1_3_1_138_2","volume-title":"ACM Symposium on Operating Systems Principles","author":"Jia Zhihao","year":"2019","unstructured":"Zhihao Jia, Oded Padon, James Thomas, Todd Warszawski, Matei Zaharia, and Alex Aiken. 2019. TASO: Optimizing deep learning computation with automatic generation of graph substitutions. In ACM Symposium on Operating Systems Principles."},{"key":"e_1_3_1_139_2","volume-title":"International Conference on Machine Learning","author":"Jia Zhihao","year":"2018","unstructured":"Zhihao Jia, Matei Zaharia, and Alex Aiken. 2018. Beyond data and model parallelism for deep neural networks. In International Conference on Machine Learning."},{"key":"e_1_3_1_140_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00492"},{"key":"e_1_3_1_141_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2020.2986127"},{"key":"e_1_3_1_142_2","volume-title":"Conference on Machine Learning and Systems","author":"Jiang Xiaotang","year":"2020","unstructured":"Xiaotang Jiang, Huan Wang, Yiliu Chen, Ziqi Wu, Lichuan Wang, Bin Zou, Yafeng Yang, Zongyang Cui, Yu Cai, Tianhang Yu, Chengfei Lv, and Zhihua Wu. 2020. MNN: A universal and efficient inference engine. In Conference on Machine Learning and Systems."},{"key":"e_1_3_1_143_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ijrobp.2017.04.021"},{"key":"e_1_3_1_144_2","volume-title":"IEEE\/ACM International Symposium on Microarchitecture","author":"Judd Patrick","year":"2016","unstructured":"Patrick Judd, Jorge Albericio, Tayler Hetherington, Tor M. Aamodt, and Andreas Moshovos. 2016. Stripes: Bit-serial deep neural network computing. In IEEE\/ACM International Symposium on Microarchitecture."},{"key":"e_1_3_1_145_2","volume-title":"IEEE\/ACM International Symposium on Microarchitecture","author":"Kao Sheng-Chun","year":"2020","unstructured":"Sheng-Chun Kao, Geonhwa Jeong, and Tushar Krishna. 2020. ConfuciuX: Autonomous hardware resource assignment for DNN accelerators using reinforcement learning. In IEEE\/ACM International Symposium on Microarchitecture."},{"key":"e_1_3_1_146_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.223"},{"key":"e_1_3_1_147_2","article-title":"Length-adaptive transformer: Train once with length drop, use anytime with search","author":"Kim Gyuwan","year":"2020","unstructured":"Gyuwan Kim and Kyunghyun Cho. 2020. Length-adaptive transformer: Train once with length drop, use anytime with search. arXiv preprint arXiv:2010.07003 (2020).","journal-title":"arXiv preprint arXiv:2010.07003"},{"key":"e_1_3_1_148_2","article-title":"I-BERT: Integer-only BERT quantization","author":"Kim Sehoon","year":"2021","unstructured":"Sehoon Kim, Amir Gholami, Zhewei Yao, Michael W. Mahoney, and Kurt Keutzer. 2021. I-BERT: Integer-only BERT quantization. arXiv preprint arXiv:2101.01321 (2021).","journal-title":"arXiv preprint arXiv:2101.01321"},{"key":"e_1_3_1_149_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1181"},{"key":"e_1_3_1_150_2","article-title":"Compression of deep convolutional neural networks for fast and low power mobile applications","author":"Kim Yong-Deok","year":"2015","unstructured":"Yong-Deok Kim, Eunhyeok Park, Sungjoo Yoo, Taelim Choi, Lu Yang, and Dongjun Shin. 2015. Compression of deep convolutional neural networks for fast and low power mobile applications. arXiv preprint arXiv:1511.06530 (2015).","journal-title":"arXiv preprint arXiv:1511.06530"},{"key":"e_1_3_1_151_2","volume-title":"International Conference on Learning Representations","author":"Kitaev Nikita","year":"2019","unstructured":"Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya. 2019. Reformer: The efficient transformer. In International Conference on Learning Representations."},{"key":"e_1_3_1_152_2","article-title":"Federated learning: Strategies for improving communication efficiency","author":"Kone\u010dn\u1ef3 Jakub","year":"2016","unstructured":"Jakub Kone\u010dn\u1ef3, H. Brendan McMahan, Felix X. Yu, Peter Richt\u00e1rik, Ananda Theertha Suresh, and Dave Bacon. 2016. Federated learning: Strategies for improving communication efficiency. arXiv preprint arXiv:1610.05492 (2016).","journal-title":"arXiv preprint arXiv:1610.05492"},{"key":"e_1_3_1_153_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00277"},{"key":"e_1_3_1_154_2","article-title":"Kaolin: A Pytorch library for accelerating 3D deep learning research","author":"Smith Jean-Francois Lafleche, Clement Fuji Tsang, Artem Rozantsev, Wenzheng Chen, Tommy Xiang, Rev Lebaredian, Sanja Fidler, Krishna Murthy Jatavallabhula, and Edward","year":"2019","unstructured":"Jean-Francois Lafleche, Clement Fuji Tsang, Artem Rozantsev, Wenzheng Chen, Tommy Xiang, Rev Lebaredian, Sanja Fidler, Krishna Murthy Jatavallabhula, and Edward Smith. 2019. Kaolin: A Pytorch library for accelerating 3D deep learning research. arXiv preprint arXiv:1911.05063 (2019).","journal-title":"arXiv preprint arXiv:1911.05063"},{"key":"e_1_3_1_156_2","volume-title":"Conference on Neural Information Processing Systems","author":"Krizhevsky Alex","year":"2012","unstructured":"Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. 2012. ImageNet classification with deep convolutional neural networks. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_157_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00827"},{"key":"e_1_3_1_158_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00935"},{"key":"e_1_3_1_159_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00109"},{"key":"e_1_3_1_160_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00479"},{"key":"e_1_3_1_161_2","article-title":"Speeding-up convolutional neural networks using fine-tuned CP-decomposition","author":"Lebedev Vadim","year":"2014","unstructured":"Vadim Lebedev, Yaroslav Ganin, Maksim Rakhuba, Ivan Oseledets, and Victor Lempitsky. 2014. Speeding-up convolutional neural networks using fine-tuned CP-decomposition. arXiv preprint arXiv:1412.6553 (2014).","journal-title":"arXiv preprint arXiv:1412.6553"},{"key":"e_1_3_1_162_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.280"},{"key":"e_1_3_1_163_2","article-title":"MNIST handwritten digit database","author":"LeCun Yann","year":"2010","unstructured":"Yann LeCun, Corinna Cortes, and Christopher J. C. Burges. 2010. MNIST handwritten digit database. AT&T Labs. Retrieved from:http:\/\/yann.lecun.com\/exdb\/mnist.","journal-title":"AT&T Labs. Retrieved from:http:\/\/yann.lecun.com\/exdb\/mnist."},{"key":"e_1_3_1_164_2","volume-title":"Conference on Neural Information Processing Systems","author":"LeCun Yann","year":"1989","unstructured":"Yann LeCun, John S. Denker, Sara A. Solla, Richard E. Howard, and Lawrence D. Jackel. 1989. Optimal brain damage. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_165_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISSCC.2018.8310262"},{"key":"e_1_3_1_166_2","volume-title":"International Conference on Machine Learning","author":"Lee Juho","year":"2019","unstructured":"Juho Lee, Yoonho Lee, Jungtaek Kim, Adam Kosiorek, Seungjin Choi, and Yee Whye Teh. 2019. Set transformer: A framework for attention-based permutation-invariant neural networks. In International Conference on Machine Learning."},{"key":"e_1_3_1_167_2","volume-title":"Conference on Empirical Methods in Natural Language Processing","author":"Lei Tao","year":"2017","unstructured":"Tao Lei, Yu Zhang, Sida I. Wang, Hui Dai, and Yoav Artzi. 2017. Simple recurrent units for highly parallelizable recurrence. In Conference on Empirical Methods in Natural Language Processing."},{"key":"e_1_3_1_168_2","article-title":"Ternary weight networks","author":"Li Fengfu","year":"2016","unstructured":"Fengfu Li, Bo Zhang, and Bin Liu. 2016. Ternary weight networks. arXiv preprint arXiv:1605.04711 (2016).","journal-title":"arXiv preprint arXiv:1605.04711"},{"key":"e_1_3_1_169_2","article-title":"DeepGCNs: Making GCNs go as deep as CNNs","author":"Li Guohao","year":"2021","unstructured":"Guohao Li, Matthias M\u00fcller, Guocheng Qian, Itzel C. Delgadillo, Abdulellah Abualshour, Ali Thabet, and Bernard Ghanem. 2021. DeepGCNs: Making GCNs go as deep as CNNs. IEEE Trans. Pattern Anal. Mach. Intell. (2021).","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"e_1_3_1_170_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00169"},{"key":"e_1_3_1_171_2","article-title":"LC-NAS: Latency constrained neural architecture search for point cloud networks","author":"Li Guohao","year":"2020","unstructured":"Guohao Li, Mengmeng Xu, Silvio Giancola, Ali Thabet, and Bernard Ghanem. 2020. LC-NAS: Latency constrained neural architecture search for point cloud networks. arXiv preprint arXiv:2008.10309 (2020).","journal-title":"arXiv preprint arXiv:2008.10309"},{"key":"e_1_3_1_172_2","volume-title":"International Conference on Learning Representations","author":"Li Hao","year":"2017","unstructured":"Hao Li, Asim Kadav, Igor Durdanovic, Hanan Samet, and Hans Peter Graf. 2017. Pruning filters for efficient ConvNets. In International Conference on Learning Representations."},{"key":"e_1_3_1_173_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00533"},{"key":"e_1_3_1_174_2","volume-title":"Conference on Neural Information Processing Systems","author":"Li Yangyan","year":"2018","unstructured":"Yangyan Li, Rui Bu, Mingchao Sun, Wei Wu, Xinhan Di, and Baoquan Chen. 2018. PointCNN: Convolution on \\( \\mathcal {X} \\) -transformed points. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_175_2","volume-title":"International Conference on Machine Learning","author":"Li Zhuohan","year":"2020","unstructured":"Zhuohan Li, Eric Wallace, Sheng Shen, Kevin Lin, Kurt Keutzer, Dan Klein, and Joey Gonzalez. 2020. Train big, then compress: Rethinking model size for efficient training and inference of transformers. In International Conference on Machine Learning."},{"key":"e_1_3_1_176_2","article-title":"MCUNet: Tiny deep learning on IoT devices","author":"Lin Ji","year":"2020","unstructured":"Ji Lin, Wei-Ming Chen, Yujun Lin, John Cohn, Chuang Gan, and Song Han. 2020. MCUNet: Tiny deep learning on IoT devices. arXiv preprint arXiv:2007.10319 (2020).","journal-title":"arXiv preprint arXiv:2007.10319"},{"key":"e_1_3_1_177_2","article-title":"Training kinetics in 15 minutes: Large-scale distributed training on videos","author":"Lin Ji","year":"2019","unstructured":"Ji Lin, Chuang Gan, and Song Han. 2019. Training kinetics in 15 minutes: Large-scale distributed training on videos. arXiv preprint arXiv:1910.00932 (2019).","journal-title":"arXiv preprint arXiv:1910.00932"},{"key":"e_1_3_1_178_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00718"},{"key":"e_1_3_1_179_2","volume-title":"Conference on Neural Information Processing Systems","author":"Lin Ji","year":"2017","unstructured":"Ji Lin, Yongming Rao, and Jiwen Lu. 2017. Runtime neural pruning. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_180_2","volume-title":"Workshop on ML for Systems at NeurIPS","author":"Lin Yujun","year":"2020","unstructured":"Yujun Lin, Driss Hafdi, Kuan Wang, Zhijian Liu, and Song Han. 2020. Neural-hardware architecture search. In Workshop on ML for Systems at NeurIPS."},{"key":"e_1_3_1_181_2","volume-title":"International Conference on Learning Representations","author":"Lin Yujun","year":"2018","unstructured":"Yujun Lin, Song Han, Huizi Mao, Yu Wang, and William J. Dally. 2018. Deep gradient compression: Reducing the communication bandwidth for distributed training. In International Conference on Learning Representations."},{"key":"e_1_3_1_182_2","doi-asserted-by":"publisher","DOI":"10.1109\/DAC18074.2021.9586250"},{"key":"e_1_3_1_183_2","volume-title":"International Conference on Learning Representations","author":"Lin Zhouhan","year":"2016","unstructured":"Zhouhan Lin, Matthieu Courbariaux, Roland Memisevic, and Yoshua Bengio. 2016. Neural networks with few multiplications. In International Conference on Learning Representations."},{"key":"e_1_3_1_184_2","volume-title":"Machine Learning for Healthcare Conference","author":"Liu Bingbin","year":"2018","unstructured":"Bingbin Liu, Michelle Guo, Edward Chou, Rishab Mehra, Serena Yeung, N. Lance Downing, Francesca Salipur, Jeffrey Jopling, Brandi Campbell, Kayla Deru, William Beninati, Arnold Milstein, and Li Fei-Fei. 2018. 3D point cloud-based visual prediction of ICU mobility care activities. In Machine Learning for Healthcare Conference."},{"key":"e_1_3_1_185_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00017"},{"key":"e_1_3_1_186_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01246-5_2"},{"key":"e_1_3_1_187_2","volume-title":"International Conference on Learning Representations","author":"Liu Hanxiao","year":"2018","unstructured":"Hanxiao Liu, Karen Simonyan, Oriol Vinyals, Chrisantha Fernando, and Koray Kavukcuoglu. 2018. Hierarchical representations for efficient architecture search. In International Conference on Learning Representations."},{"key":"e_1_3_1_188_2","volume-title":"International Conference on Learning Representations","author":"Liu Haoxiao","year":"2019","unstructured":"Haoxiao Liu, Karen Simonyan, and Yiming Yang. 2019. DARTS: Differentiable architecture search. In International Conference on Learning Representations."},{"key":"e_1_3_1_189_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11630"},{"key":"e_1_3_1_190_2","volume-title":"International Conference on Learning Representations","author":"Liu Liu","year":"2019","unstructured":"Liu Liu, Lei Deng, Xing Hu, Maohua Zhu, Guoqi Li, Yufei Ding, and Yuan Xie. 2019. Dynamic sparse graph for efficient deep learning. In International Conference on Learning Representations."},{"key":"e_1_3_1_191_2","volume-title":"International Conference on Learning Representations","author":"Liu Peter J.","year":"2018","unstructured":"Peter J. Liu, Mohammad Saleh, Etienne Pot, Ben Goodrich, Ryan Sepassi, Lukasz Kaiser, and Noam Shazeer. 2018. Generating Wikipedia by summarizing long sequences. In International Conference on Learning Representations."},{"key":"e_1_3_1_192_2","volume-title":"Conference of the Association for Computational Linguistics","author":"Liu Shujie","year":"2015","unstructured":"Shujie Liu, Nan Yang, Mu Li, and Ming Zhou. 2015. A recursive recurrent neural network for statistical machine translation. In Conference of the Association for Computational Linguistics."},{"key":"e_1_3_1_193_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00271"},{"key":"e_1_3_1_194_2","volume-title":"Hardware-efficient Deep Learning for 3D Point Cloud","author":"Liu Zhijian","year":"2020","unstructured":"Zhijian Liu. 2020. Hardware-efficient Deep Learning for 3D Point Cloud. Master\u2019s. Thesis Massachusetts Institute of Technology."},{"key":"e_1_3_1_195_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA48506.2021.9561299"},{"key":"e_1_3_1_196_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.298"},{"key":"e_1_3_1_197_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00339"},{"key":"e_1_3_1_198_2","volume-title":"Conference on Neural Information Processing Systems","author":"Liu Zhijian","year":"2019","unstructured":"Zhijian Liu, Haotian Tang, Yujun Lin, and Song Han. 2019. Point-voxel CNN for efficient 3D deep learning. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_199_2","volume-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence","author":"Liu Zhijian","year":"2021","unstructured":"Zhijian Liu, Haotian Tang, Shengyu Zhao, Kevin Shao, and Song Han. 2021. PVNAS: 3D Neural Architecture Search with Point-Voxel Convolution. IEEE Transactions on Pattern Analysis and Machine Intelligence (2021)."},{"key":"e_1_3_1_200_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58621-8_34"},{"key":"e_1_3_1_201_2","doi-asserted-by":"publisher","DOI":"10.1109\/SOCC49529.2020.9524802"},{"key":"e_1_3_1_202_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00561"},{"key":"e_1_3_1_203_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01264-9_8"},{"key":"e_1_3_1_204_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i04.5954"},{"key":"e_1_3_1_205_2","article-title":"Fine-grained visual classification of aircraft","author":"Maji Subhransu","year":"2013","unstructured":"Subhransu Maji, Esa Rahtu, Juho Kannala, Matthew Blaschko, and Andrea Vedaldi. 2013. Fine-grained visual classification of aircraft. arXiv preprint arXiv:1306.5151 (2013).","journal-title":"arXiv preprint arXiv:1306.5151"},{"key":"e_1_3_1_206_2","volume-title":"Conference on Neural Information Processing Systems","author":"Mao Hongzi","year":"2019","unstructured":"Hongzi Mao, Parimarjan Negi, Akshay Narayan, Hanrui Wang, Jiacheng Yang, Haonan Wang, Ryan Marcus, Ravichandra Addanki, Mehrdad Khani, Songtao He, Vikram Nathan, Frank Cangialosi, Shaileshh Venkatakrishnan, Wei-Hung Weng, Song Han, Tim Kraska, and Mohammad Alizadeh. 2019. Park: An open platform for learning-augmented computer systems. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_207_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00166"},{"key":"e_1_3_1_208_2","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2015.7353481"},{"key":"e_1_3_1_209_2","volume-title":"International Conference on Artificial Intelligence and Statistics","author":"McMahan Brendan","year":"2016","unstructured":"Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas. 2016. Communication-efficient learning of deep networks from decentralized data. In International Conference on Artificial Intelligence and Statistics."},{"key":"e_1_3_1_210_2","volume-title":"Conference on Neural Information Processing Systems","author":"Michel Paul","year":"2019","unstructured":"Paul Michel, Omer Levy, and Graham Neubig. 2019. Are sixteen heads really better than one? In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_211_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2010-343"},{"key":"e_1_3_1_212_2","volume-title":"International Conference on Learning Representations","author":"Mirhoseini Azalia","year":"2018","unstructured":"Azalia Mirhoseini, Anna Goldie, Hieu Pham, Benoit Steiner, Quoc V. Le, and Jeff Dean. 2018. A hierarchical model for device placement. In International Conference on Learning Representations."},{"key":"e_1_3_1_213_2","volume-title":"International Conference on Machine Learning","author":"Mirhoseini Azalia","year":"2017","unstructured":"Azalia Mirhoseini, Hieu Pham, Quoc V. Le, Benoit Steiner, Rasmus Larsen, Yuefeng Zhou, Naveen Kumar, Mohammad Norouzi, Samy Bengio, and Jeff Dean. 2017. Device placement optimization with reinforcement learning. In International Conference on Machine Learning."},{"key":"e_1_3_1_214_2","volume-title":"International Conference on Learning Representations","author":"Molchanov Pavlo","year":"2017","unstructured":"Pavlo Molchanov, Stephen Tyree, Tero Karras, Timo Aila, and Jan Kautz. 2017. Pruning convolutional neural networks for resource efficient transfer learning. In International Conference on Learning Representations."},{"key":"e_1_3_1_215_2","volume-title":"International Conference on Learning Representations","author":"Mudrakarta Pramod Kaushik","year":"2019","unstructured":"Pramod Kaushik Mudrakarta, Mark Sandler, Andrey Zhmoginov, and Andrew Howard. 2019. K for the price of 1: Parameter-efficient multi-task and transfer learning. In International Conference on Learning Representations."},{"key":"e_1_3_1_216_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00141"},{"key":"e_1_3_1_217_2","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378534"},{"key":"e_1_3_1_218_2","article-title":"Deep learning for mobile multimedia: A survey","author":"Ota Kaoru","year":"2017","unstructured":"Kaoru Ota, Minh Son Dao, Vasileios Mezaris, and Francesco G. B. De Natale. 2017. Deep learning for mobile multimedia: A survey. ACM Transactions on Multimedia Computing, Communications, and Applications 13, 35 (2017), 1\u201322.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"key":"e_1_3_1_219_2","volume-title":"IEEE International Symposium on High-performance Computer Architecture","author":"Pal Subhankar","year":"2018","unstructured":"Subhankar Pal, Jonathan Beaumont, Dong-Hyeon Park, Aporva Amarnath, Siying Feng, Chaitali Chakrabarti, Hun-Seok Kim, David Blaauw, Trevor Mudge, and Ronald Dreslinski. 2018. OuterSPACE: An outer product based sparse matrix multiplication accelerator. In IEEE International Symposium on High-performance Computer Architecture."},{"key":"e_1_3_1_220_2","volume-title":"International Symposium on Computer Architecture","author":"Parashar Angshuman","year":"2017","unstructured":"Angshuman Parashar, Minsoo Rhu, Anurag Mukkara, Antonio Puglielli, Rangharajan Venkatesan, Brucek Khailany, Joel Emer, Stephen W. Keckler, and William J. Dally. 2017. SCNN: An accelerator for compressed-sparse convolutional neural networks. In International Symposium on Computer Architecture."},{"key":"e_1_3_1_221_2","article-title":"Faster CNNs with direct sparse convolutions and guided pruning","author":"Park Jongsoo","year":"2016","unstructured":"Jongsoo Park, Sheng Li, Wei Wen, Ping Tak Peter Tang, Hai Li, Yiran Chen, and Pradeep Dubey. 2016. Faster CNNs with direct sparse convolutions and guided pruning. arXiv preprint arXiv:1608.01409 (2016).","journal-title":"arXiv preprint arXiv:1608.01409"},{"key":"e_1_3_1_222_2","volume-title":"IEEE International Solid-state Circuits Conference","author":"Park Seongwook","year":"2015","unstructured":"Seongwook Park, Kyeongryeol Bong, Dongjoo Shin, Jinmook Lee, Sungpill Choi, and Hoi-Jun Yoo. 2015. A 1.93TOPS\/W Scalable deep learning\/inference processor with tetra-parallel MIMD architecture for big-data applications. In IEEE International Solid-state Circuits Conference."},{"key":"e_1_3_1_223_2","article-title":"Memory-augmented neural networks on FPGA for real-time and energy-efficient question answering","author":"Park Seongsik","year":"2020","unstructured":"Seongsik Park, Jaehee Jang, Seijoon Kim, Byunggook Na, and Sungroh Yoon. 2020. Memory-augmented neural networks on FPGA for real-time and energy-efficient question answering. IEEE Trans. Very Large Scale Integ. Syst. (2020).","journal-title":"IEEE Trans. Very Large Scale Integ. Syst."},{"key":"e_1_3_1_224_2","volume-title":"International Conference on Machine Learning","author":"Parmar Niki","year":"2018","unstructured":"Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran. 2018. Image transformer. In International Conference on Machine Learning."},{"key":"e_1_3_1_225_2","volume-title":"Conference on Neural Information Processing Systems","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. PyTorch: An imperative style, high-performance deep learning library. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_226_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCD.2013.6657019"},{"key":"e_1_3_1_227_2","volume-title":"International Conference on Machine Learning","author":"Pham Hieu","year":"2018","unstructured":"Hieu Pham, Melody Y. Guan, Barret Zoph, Quoc V. Le, and Jeff Dean. 2018. Efficient neural architecture search via parameter sharing. In International Conference on Machine Learning."},{"key":"e_1_3_1_228_2","article-title":"Tiny video networks","author":"Piergiovanni A. J.","year":"2019","unstructured":"A. J. Piergiovanni, Anelia Angelova, and Michael S. Ryoo. 2019. Tiny video networks. arXiv preprint arXiv:1910.06961 (2019).","journal-title":"arXiv preprint arXiv:1910.06961"},{"key":"e_1_3_1_229_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00188"},{"key":"e_1_3_1_230_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00937"},{"key":"e_1_3_1_231_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00102"},{"key":"e_1_3_1_232_2","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Qi Charles R.","year":"2017","unstructured":"Charles R. Qi, Hao Su, Kaichun Mo, and Leonidas J. Guibas. 2017. PointNet: Deep learning on point sets for 3D classification and segmentation. In IEEE Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_3_1_233_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.609"},{"key":"e_1_3_1_234_2","volume-title":"Conference on Neural Information Processing Systems","author":"Qi Charles R.","year":"2017","unstructured":"Charles R. Qi, Li Yi, Hao Su, and Leonidas J. Guibas. 2017. PointNet++: Deep hierarchical feature learning on point sets in a metric space. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_235_2","volume-title":"IEEE International Symposium on High-performance Computer Architecture","author":"Qin Eric","year":"2020","unstructured":"Eric Qin, Ananda Samajdar, Hyoukjun Kwon, Vineet Nadella, Sudarshan Srinivasan, Dipankar Das, Bharat Kaul, and Tushar Krishna. 2020. SIGMA: A sparse and irregular GEMM accelerator with flexible interconnects for DNN training. In IEEE International Symposium on High-performance Computer Architecture."},{"key":"e_1_3_1_236_2","article-title":"Blockwise self-attention for long document understanding","author":"Qiu Jiezhong","year":"2019","unstructured":"Jiezhong Qiu, Hao Ma, Omer Levy, Scott Wen-tau Yih, Sinong Wang, and Jie Tang. 2019. Blockwise self-attention for long document understanding. arXiv preprint arXiv:1911.02972 (2019).","journal-title":"arXiv preprint arXiv:1911.02972"},{"key":"e_1_3_1_237_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.590"},{"key":"e_1_3_1_238_2","unstructured":"Alec Radford Jeff Wu Rewon Child David Luan Dario Amodei and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. (2019)."},{"key":"e_1_3_1_239_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01044"},{"key":"e_1_3_1_240_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D16-1264"},{"key":"e_1_3_1_241_2","article-title":"Searching for activation functions","author":"Ramachandran Prajit","year":"2017","unstructured":"Prajit Ramachandran, Barret Zoph, and Quoc V. Le. 2017. Searching for activation functions. arXiv preprint arXiv:1710.05941 (2017).","journal-title":"arXiv preprint arXiv:1710.05941"},{"key":"e_1_3_1_242_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46493-0_32"},{"key":"e_1_3_1_243_2","article-title":"Accelerating 3D deep learning with PyTorch3D","author":"Ravi Nikhila","year":"2020","unstructured":"Nikhila Ravi, Jeremy Reizenstein, David Novotny, Taylor Gordon, Wan-Yen Lo, Justin Johnson, and Georgia Gkioxari. 2020. Accelerating 3D deep learning with PyTorch3D. arXiv preprint arXiv:2007.08501 (2020).","journal-title":"arXiv preprint arXiv:2007.08501"},{"key":"e_1_3_1_244_2","doi-asserted-by":"crossref","unstructured":"Esteban Real Alok Aggarwal Yanping Huang and Quoc V. Le. 2019. Regularized evolution for image classifier architecture search. In AAAI Conference on Artificial Intelligence .","DOI":"10.1609\/aaai.v33i01.33014780"},{"key":"e_1_3_1_245_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.701"},{"key":"e_1_3_1_246_2","article-title":"FitNets: Hints for thin deep nets","author":"Romero Adriana","year":"2014","unstructured":"Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio. 2014. FitNets: Hints for thin deep nets. arXiv preprint arXiv:1412.6550 (2014).","journal-title":"arXiv preprint arXiv:1412.6550"},{"key":"e_1_3_1_247_2","article-title":"Efficient content-based sparse attention with routing transformers","author":"Roy Aurko","year":"2020","unstructured":"Aurko Roy, Mohammad Saffar, Ashish Vaswani, and David Grangier. 2020. Efficient content-based sparse attention with routing transformers. Trans. Assoc. Comput. Ling. 9, 3 (2020), 53\u201368.","journal-title":"Trans. Assoc. Comput. Ling."},{"key":"e_1_3_1_248_2","doi-asserted-by":"crossref","unstructured":"Manuele Rusci Marco Fariselli Alessandro Capotondi and Luca Benini. 2020. Leveraging automated mixed-low-precision quantization for tiny edge microcontrollers. arXiv preprint arXiv:2008.05124 (2020).","DOI":"10.1007\/978-3-030-66770-2_22"},{"key":"e_1_3_1_249_2","volume-title":"International Conference on Learning Representations","author":"Ryoo Michael S.","year":"2020","unstructured":"Michael S. Ryoo, A. J. Piergiovanni, Mingxing Tan, and Anelia Angelova. 2020. AssembleNet: Searching for multi-stream neural connectivity in video architectures. In International Conference on Learning Representations."},{"key":"e_1_3_1_250_2","volume-title":"International Joint Conference on Neural Networks","author":"Rzayev Tayyar","year":"2017","unstructured":"Tayyar Rzayev, Saber Moradi, David H. Albonesi, and Rajit Manchar. 2017. DeepRecon: Dynamically reconfigurable architecture for accelerating deep neural networks. In International Joint Conference on Neural Networks."},{"key":"e_1_3_1_251_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00474"},{"key":"e_1_3_1_252_2","volume-title":"Conference on Neural Information Processing Systems","author":"Sanh Victor","year":"2019","unstructured":"Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019. DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_253_2","doi-asserted-by":"publisher","DOI":"10.1109\/ASAP.2009.25"},{"key":"e_1_3_1_254_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2014-274"},{"key":"e_1_3_1_255_2","article-title":"CNN features off-the-shelf: An astounding baseline for recognition","author":"Razavian Ali Sharif","year":"2014","unstructured":"Ali Sharif Razavian, Hossein Azizpour, Josephine Sullivan, and Stefan Carlsson. 2014. CNN features off-the-shelf: An astounding baseline for recognition. arXiv preprint arXiv:1403.6382 (2014).","journal-title":"arXiv preprint arXiv:1403.6382"},{"key":"e_1_3_1_256_2","doi-asserted-by":"publisher","DOI":"10.1145\/3195970.3196072"},{"key":"e_1_3_1_257_2","volume-title":"International Symposium on Computer Architecture","author":"Sharma Hardik","year":"2018","unstructured":"Hardik Sharma, Jongse Park, Naveen Suda, Liangzhen Lai, Benson Chau, Vikas Chandra, and Hadi Esmaeilzadeh. 2018. Bit fusion: Bit-level dynamically composable architecture for accelerating deep neural networks. In International Symposium on Computer Architecture."},{"key":"e_1_3_1_258_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.331"},{"key":"e_1_3_1_259_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i05.6409"},{"key":"e_1_3_1_260_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01054"},{"key":"e_1_3_1_261_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00086"},{"key":"e_1_3_1_262_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2020.2977026"},{"key":"e_1_3_1_263_2","volume-title":"Conference on Neural Information Processing Systems","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Two-stream convolutional networks for action recognition in videos. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_264_2","volume-title":"International Conference on Learning Representations","author":"Simonyan Karen","year":"2015","unstructured":"Karen Simonyan and Andrew Zisserman. 2015. Very deep convolutional networks for large-scale image recognition. In International Conference on Learning Representations."},{"key":"e_1_3_1_265_2","volume-title":"International Conference on Machine Learning","author":"So David R.","year":"2019","unstructured":"David R. So, Chen Liang, and Quoc V. Le. 2019. The evolved transformer. In International Conference on Machine Learning."},{"key":"e_1_3_1_266_2","doi-asserted-by":"publisher","DOI":"10.5244\/C.29.31"},{"key":"e_1_3_1_267_2","doi-asserted-by":"publisher","DOI":"10.1109\/FPT.2010.5681487"},{"key":"e_1_3_1_268_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV45572.2020.9093274"},{"key":"e_1_3_1_269_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1355"},{"key":"e_1_3_1_270_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00118"},{"key":"e_1_3_1_271_2","article-title":"Optimizing network performance for distributed DNN training on GPU clusters: ImageNet\/AlexNet training in 1.5 minutes","author":"Sun Peng","year":"2019","unstructured":"Peng Sun, Wansen Feng, Ruobing Han, Shengen Yan, and Yonggang Wen. 2019. Optimizing network performance for distributed DNN training on GPU clusters: ImageNet\/AlexNet training in 1.5 minutes. arXiv preprint arXiv:1902.06855 (2019).","journal-title":"arXiv preprint arXiv:1902.06855"},{"key":"e_1_3_1_272_2","unstructured":"Ilya Sutskever Oriol Vinyals and Quoc V. Le. 2014. Sequence to sequence learning with neural networks. In Conference on Neural Information Processing Systems ."},{"key":"e_1_3_1_273_2","doi-asserted-by":"crossref","unstructured":"Vivienne Sze Yu-Hsin Chen Tien-Ju Yang and Joel S. Emer. 2017. Efficient processing of deep neural networks: A tutorial and survey. Proceedings of the IEEE 105 12 (2017) 2295\u20132329.","DOI":"10.1109\/JPROC.2017.2761740"},{"key":"e_1_3_1_274_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"e_1_3_1_275_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.308"},{"key":"e_1_3_1_276_2","article-title":"EdgeBERT: Optimizing on-chip inference for multi-task NLP","author":"Tambe Thierry","year":"2020","unstructured":"Thierry Tambe, Coleman Hooper, Lillian Pentecost, En-Yu Yang, Marco Donato, Victor Sanh, Alexander M. Rush, David Brooks, and Gu-Yeon Wei. 2020. EdgeBERT: Optimizing on-chip inference for multi-task NLP. arXiv preprint arXiv:2011.14203 (2020).","journal-title":"arXiv preprint arXiv:2011.14203"},{"key":"e_1_3_1_277_2","doi-asserted-by":"publisher","DOI":"10.1109\/DAC18072.2020.9218516"},{"key":"e_1_3_1_278_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00293"},{"key":"e_1_3_1_279_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01079"},{"key":"e_1_3_1_280_2","article-title":"PCNN: Pattern-based fine-grained regular pruning towards optimizing CNN accelerators","author":"Tan Zhanhong","year":"2020","unstructured":"Zhanhong Tan, Jiebo Song, Xiaolong Ma, Sia-Huat Tan, Hongyang Chen, Yuanqing Miao, Yifu Wu, Shaokai Ye, Yanzhi Wang, Dehui Li, and Kaisheng Ma. 2020. PCNN: Pattern-based fine-grained regular pruning towards optimizing CNN accelerators. arXiv preprint arXiv:2002.04997 (2020).","journal-title":"arXiv preprint arXiv:2002.04997"},{"key":"e_1_3_1_281_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58604-1_41"},{"key":"e_1_3_1_282_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00409"},{"key":"e_1_3_1_283_2","article-title":"Efficient transformers: A survey","author":"Tay Yi","year":"2020","unstructured":"Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. 2020. Efficient transformers: A survey. arXiv preprint arXiv:2009.06732 (2020).","journal-title":"arXiv preprint arXiv:2009.06732"},{"key":"e_1_3_1_284_2","doi-asserted-by":"publisher","DOI":"10.1109\/3DV.2017.00067"},{"key":"e_1_3_1_285_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00651"},{"key":"e_1_3_1_286_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.510"},{"key":"e_1_3_1_287_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00565"},{"key":"e_1_3_1_288_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00675"},{"key":"e_1_3_1_289_2","volume-title":"International Symposium on Field-programmable Gate Arrays","author":"Umuroglu Yaman","year":"2017","unstructured":"Yaman Umuroglu, Nicholas J. Fraser, Giulio Gambardella, Michaela Blott, Philip Leong, Magnus Jahre, and Kees Vissers. 2017. FINN: A framework for fast, scalable binarized neural network inference. In International Symposium on Field-programmable Gate Arrays."},{"key":"e_1_3_1_290_2","doi-asserted-by":"publisher","DOI":"10.1109\/FPL.2018.00059"},{"key":"e_1_3_1_291_2","volume-title":"Conference on Neural Information Processing Systems","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_292_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1580"},{"key":"e_1_3_1_293_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W18-5446"},{"key":"e_1_3_1_294_2","volume-title":"Efficient Algorithms and Hardware for Natural Language Processing","author":"Wang Hanrui","year":"2020","unstructured":"Hanrui Wang. 2020. Efficient Algorithms and Hardware for Natural Language Processing. Master\u2019s. Thesis Massachusetts Institute of Technology."},{"key":"e_1_3_1_295_2","doi-asserted-by":"publisher","DOI":"10.1109\/DAC18072.2020.9218757"},{"key":"e_1_3_1_296_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.686"},{"key":"e_1_3_1_297_2","volume-title":"Workshop on ML for Systems at NeurIPS","author":"Wang Hanrui","year":"2018","unstructured":"Hanrui Wang, Jiacheng Yang, Hae-Seung Lee, and Song Han. 2018. Learning to design circuits. In Workshop on ML for Systems at NeurIPS."},{"key":"e_1_3_1_298_2","volume-title":"IEEE International Symposium on High-performance Computer Architecture","author":"Wang Hanrui","year":"2021","unstructured":"Hanrui Wang, Zhekai Zhang, and Song Han. 2021. SpAtten: Efficient sparse attention architecture with cascade token and head pruning. In IEEE International Symposium on High-performance Computer Architecture."},{"key":"e_1_3_1_299_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00881"},{"key":"e_1_3_1_300_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-020-01339-6"},{"key":"e_1_3_1_301_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46484-8_2"},{"key":"e_1_3_1_302_2","doi-asserted-by":"publisher","DOI":"10.1145\/3072959.3073608"},{"key":"e_1_3_1_303_2","doi-asserted-by":"publisher","DOI":"10.1145\/3272127.3275050"},{"key":"e_1_3_1_304_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1176"},{"key":"e_1_3_1_305_2","article-title":"Linformer: Self-attention with linear complexity","author":"Wang Sinong","year":"2020","unstructured":"Sinong Wang, Belinda Li, Madian Khabsa, Han Fang, and Hao Ma. 2020. Linformer: Self-attention with linear complexity. arXiv preprint arXiv:2006.04768 (2020).","journal-title":"arXiv preprint arXiv:2006.04768"},{"key":"e_1_3_1_306_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00215"},{"key":"e_1_3_1_307_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00272"},{"key":"e_1_3_1_308_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01261-8_25"},{"key":"e_1_3_1_309_2","doi-asserted-by":"publisher","DOI":"10.1145\/3326362"},{"key":"e_1_3_1_310_2","article-title":"SparseDNN: Fast sparse deep learning inference on CPUs","author":"Wang Ziheng","year":"2021","unstructured":"Ziheng Wang. 2021. SparseDNN: Fast sparse deep learning inference on CPUs. arXiv preprint arXiv:2101.07948 (2021).","journal-title":"arXiv preprint arXiv:2101.07948"},{"key":"e_1_3_1_311_2","article-title":"PV-NAS: Practical neural architecture search for video recognition","author":"Wang Zihao","year":"2020","unstructured":"Zihao Wang, Chen Lin, Lu Sheng, Junjie Yan, and Jing Shao. 2020. PV-NAS: Practical neural architecture search for video recognition. arXiv preprint arXiv:2011.00826 (2020).","journal-title":"arXiv preprint arXiv:2011.00826"},{"key":"e_1_3_1_312_2","doi-asserted-by":"crossref","unstructured":"Zongji Wang and Feng Lu. 2019. VoxSegNet: Volumetric CNNs for semantic part segmentation of 3D shapes. IEEE Transactions on Visualization and Computer Graphics 26 9 (2019) 2919\u20132930.","DOI":"10.1109\/TVCG.2019.2896310"},{"key":"e_1_3_1_313_2","volume-title":"Conference on Neural Information Processing Systems","author":"Wangni Jianqiao","year":"2018","unstructured":"Jianqiao Wangni, Jialei Wang, Ji Liu, and Tong Zhang. 2018. Gradient sparsification for communication-efficient distributed optimization. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_314_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1125"},{"key":"e_1_3_1_315_2","volume-title":"Conference on Neural Information Processing Systems","author":"Wen Wei","year":"2016","unstructured":"Wei Wen, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li. 2016. Learning structured sparsity in deep neural networks. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_316_2","volume-title":"Conference on Neural Information Processing Systems","author":"Wen Wei","year":"2017","unstructured":"Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li. 2017. TernGrad: Ternary gradients to reduce communication in distributed deep learning. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_317_2","article-title":"A survey on neural architecture search","author":"Wistuba Martin","year":"2019","unstructured":"Martin Wistuba, Ambrish Rawat, and Tejaswini Pedapati. 2019. A survey on neural architecture search. arXiv preprint arXiv:1905.01392 (2019).","journal-title":"arXiv preprint arXiv:1905.01392"},{"key":"e_1_3_1_318_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01099"},{"key":"e_1_3_1_319_2","volume-title":"International Conference on Learning Representations","author":"Wu Felix","year":"2019","unstructured":"Felix Wu, Angela Fan, Alexei Baevski, Yann N. Dauphin, and Michael Auli. 2019. Pay less attention with lightweight and dynamic convolutions. In International Conference on Learning Representations."},{"key":"e_1_3_1_320_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.521"},{"key":"e_1_3_1_321_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00985"},{"key":"e_1_3_1_322_2","volume-title":"International Conference on Learning Representations","author":"Wu Zhanghao","year":"2020","unstructured":"Zhanghao Wu, Zhijian Liu, Ji Lin, Yujun Lin, and Song Han. 2020. Lite transformer with long-short range attention. In International Conference on Learning Representations."},{"key":"e_1_3_1_323_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00919"},{"key":"e_1_3_1_324_2","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Wu Zhirong","year":"2015","unstructured":"Zhirong Wu, Shuran Song, Aditya Khosla, Fisher Yu, Linguang Zhang, Xiaoou Tang, and Jianxiong Xiao. 2015. 3D ShapeNets: A deep representation for volumetric shapes. In IEEE Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_3_1_325_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58580-8_34"},{"key":"e_1_3_1_326_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01267-0_19"},{"key":"e_1_3_1_327_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.204"},{"key":"e_1_3_1_328_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2018.8461870"},{"key":"e_1_3_1_329_2","doi-asserted-by":"publisher","DOI":"10.1145\/3308558.3313591"},{"key":"e_1_3_1_330_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00570"},{"key":"e_1_3_1_331_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01237-3_6"},{"key":"e_1_3_1_332_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2013-552"},{"key":"e_1_3_1_333_2","doi-asserted-by":"publisher","DOI":"10.3390\/s18103337"},{"key":"e_1_3_1_334_2","article-title":"MicroNet for efficient language modeling","author":"Yan Zhongxia","year":"2020","unstructured":"Zhongxia Yan, Hanrui Wang, Demi Guo, and Song Han. 2020. MicroNet for efficient language modeling. J. Mach. Learn. Res. 123, 20 (2020), 215\u2013231.","journal-title":"J. Mach. Learn. Res."},{"key":"e_1_3_1_335_2","doi-asserted-by":"publisher","DOI":"10.1109\/DAC18072.2020.9218676"},{"key":"e_1_3_1_336_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01249-6_18"},{"key":"e_1_3_1_337_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00204"},{"key":"e_1_3_1_338_2","doi-asserted-by":"publisher","DOI":"10.1056\/NEJMp1716891"},{"key":"e_1_3_1_339_2","article-title":"AutoSlim: Towards one-shot architecture search for channel numbers","author":"Yu Jiahui","year":"2019","unstructured":"Jiahui Yu and Thomas Huang. 2019. AutoSlim: Towards one-shot architecture search for channel numbers. arXiv preprint arXiv:1903.11728 (2019).","journal-title":"arXiv preprint arXiv:1903.11728"},{"key":"e_1_3_1_340_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00189"},{"key":"e_1_3_1_341_2","volume-title":"International Symposium on Computer Architecture","author":"Yu Jiecao","year":"2017","unstructured":"Jiecao Yu, Andrew Lukefahr, David Palframan, Ganesh Dasika, Reetuparna Das, and Scott Mahlke. 2017. Scalpel: Customizing DNN pruning to the underlying hardware parallelism. In International Symposium on Computer Architecture."},{"key":"e_1_3_1_342_2","volume-title":"International Conference on Learning Representations","author":"Yu Jiahui","year":"2019","unstructured":"Jiahui Yu, Linjie Yang, Ning Xu, Jianchao Yang, and Thomas Huang. 2019. Slimmable neural networks. In International Conference on Learning Representations."},{"key":"e_1_3_1_343_2","volume-title":"Joint Pattern Recognition Symposium","author":"Zach Christopher","year":"2007","unstructured":"Christopher Zach, Thomas Pock, and Horst Bischof. 2007. A duality based approach for realtime TV-L1 optical flow. In Joint Pattern Recognition Symposium."},{"key":"e_1_3_1_344_2","volume-title":"IEEE\/ACM International Symposium on Microarchitecture","author":"Zadeh Ali Hadi","year":"2020","unstructured":"Ali Hadi Zadeh, Isak Edo, Omar Mohamed Awad, and Andreas Moshovos. 2020. GOBO: Quantizing attention-based NLP models for low latency and energy efficient inference. In IEEE\/ACM International Symposium on Microarchitecture."},{"key":"e_1_3_1_345_2","volume-title":"International Conference on Learning Representations","author":"Zagoruyko Sergey","year":"2017","unstructured":"Sergey Zagoruyko and Nikos Komodakis. 2017. Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer. In International Conference on Learning Representations."},{"key":"e_1_3_1_346_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1166"},{"key":"e_1_3_1_347_2","volume-title":"International Symposium on Field-programmable Gate Arrays","author":"Zhang Chen","year":"2015","unstructured":"Chen Zhang, Peng Li, Guangyu Sun, Yijin Guan, Bingjun Xiao, and Jason Cong. 2015. Optimizing FPGA-based accelerator design for deep convolutional neural networks. In International Symposium on Field-programmable Gate Arrays."},{"key":"e_1_3_1_348_2","volume-title":"IEEE\/ACM International Symposium on Microarchitecture","author":"Zhang Shijin","year":"2016","unstructured":"Shijin Zhang, Zidong Du, Lei Zhang, Huiying Lan, Shaoli Liu, Ling Li, Qi Guo, Tianshi Chen, and Yunji Chen. 2016. Cambricon-X: An accelerator for sparse neural networks. In IEEE\/ACM International Symposium on Microarchitecture."},{"key":"e_1_3_1_349_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.37"},{"key":"e_1_3_1_350_2","volume-title":"Conference on Neural Information Processing Systems","author":"Zhang Xiang","year":"2015","unstructured":"Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_351_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00716"},{"key":"e_1_3_1_352_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2015.2502579"},{"key":"e_1_3_1_353_2","volume-title":"IEEE International Symposium on High-performance Computer Architecture","author":"Zhang Zhekai","year":"2020","unstructured":"Zhekai Zhang, Hanrui Wang, Song Han, and William J. Dally. 2020. SpArch: Efficient architecture for sparse matrix multiplication. In IEEE International Symposium on High-performance Computer Architecture."},{"key":"e_1_3_1_354_2","article-title":"AutoEmb: Automated embedding dimensionality search in streaming recommendations","author":"Zhao Xiangyu","year":"2020","unstructured":"Xiangyu Zhao, Chong Wang, Ming Chen, Xudong Zheng, Xiaobing Liu, and Jiliang Tang. 2020. AutoEmb: Automated embedding dimensionality search in streaming recommendations. arXiv preprint arXiv:2002.11252 (2020).","journal-title":"arXiv preprint arXiv:2002.11252"},{"key":"e_1_3_1_355_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00257"},{"key":"e_1_3_1_356_2","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Zhou Shuchang","year":"2018","unstructured":"Shuchang Zhou, Yuxin Wu, Zekun Ni, Xinyu Zhou, He Wen, and Yuheng Zou. 2018. DoReFa-Net: Training low bitwidth convolutional neural networks with low bitwidth gradients. In IEEE Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_3_1_357_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00472"},{"key":"e_1_3_1_358_2","volume-title":"International Conference on Learning Representations","author":"Zhu Chenzhuo","year":"2017","unstructured":"Chenzhuo Zhu, Song Han, Huizi Mao, and William Dally. 2017. Trained ternary quantization. In International Conference on Learning Representations."},{"key":"e_1_3_1_359_2","volume-title":"Conference on Neural Information Processing Systems","author":"Zhu Ligeng","year":"2021","unstructured":"Ligeng Zhu, Hongzhou Lin, Yao Lu, Yujun Lin, and Song Han. 2021. Delayed gradient averaging: Tolerate the communication latency in federated learning. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_360_2","volume-title":"Conference on Neural Information Processing Systems","author":"Zhu Ligeng","year":"2019","unstructured":"Ligeng Zhu, Zhijian Liu, and Han Song. 2019. Deep leakage for gradient. In Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_361_2","unstructured":"Ligeng Zhu Yao Lu Yujun Lin and Song Han. 2019. Distributed training across the World. In NeurIPS Workshop on Systems for ML ."},{"key":"e_1_3_1_362_2","article-title":"A3D: Adaptive 3D networks for video action recognition","author":"Zhu Sijie","year":"2020","unstructured":"Sijie Zhu, Taojiannan Yang, Matias Mendieta, and Chen Chen. 2020. A3D: Adaptive 3D networks for video action recognition. arXiv preprint arXiv:2011.12384 (2020).","journal-title":"arXiv preprint arXiv:2011.12384"},{"key":"e_1_3_1_363_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.52"},{"key":"e_1_3_1_364_2","doi-asserted-by":"publisher","DOI":"10.1109\/3DV.2019.00035"},{"key":"e_1_3_1_365_2","volume-title":"International Symposium on Field-programmable Gate Arrays","author":"Zhuo Ling","year":"2005","unstructured":"Ling Zhuo and Viktor K. Prasanna. 2005. Sparse matrix-vector multiplication on FPGAs. In International Symposium on Field-programmable Gate Arrays."},{"key":"e_1_3_1_366_2","volume-title":"International Conference on Learning Representations","author":"Zoph Barret","year":"2017","unstructured":"Barret Zoph and Quoc V. Le. 2017. Neural architecture search with reinforcement learning. In International Conference on Learning Representations."},{"key":"e_1_3_1_367_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00907"},{"key":"e_1_3_1_368_2","doi-asserted-by":"publisher","DOI":"10.1587\/elex.10.20130529"}],"container-title":["ACM Transactions on Design Automation of Electronic Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3486618","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3486618","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:12:06Z","timestamp":1750191126000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3486618"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,3,4]]},"references-count":366,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2022,5,31]]}},"alternative-id":["10.1145\/3486618"],"URL":"https:\/\/doi.org\/10.1145\/3486618","relation":{},"ISSN":["1084-4309","1557-7309"],"issn-type":[{"value":"1084-4309","type":"print"},{"value":"1557-7309","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,3,4]]},"assertion":[{"value":"2021-04-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-09-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-03-04","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}