Efficient Hardware Architectures for Accelerating Deep Neural Networks: Survey

Dhilleswararao, Pudi; Boppu, Srinivas; Manikandan, M. Sabarimalai; Cenkeramaddi, Linga Reddy

dc.contributor.author	Dhilleswararao, Pudi
dc.contributor.author	Boppu, Srinivas
dc.contributor.author	Manikandan, M. Sabarimalai
dc.contributor.author	Cenkeramaddi, Linga Reddy
dc.date.accessioned	2023-01-03T13:45:51Z
dc.date.available	2023-01-03T13:45:51Z
dc.date.created	2022-12-16T09:22:00Z
dc.date.issued	2022
dc.identifier.citation	Dhilleswararao, P, Boppu, S, Manikandan, M. S. & Cenkeramaddi, L. R. (2022). Efficient Hardware Architectures for Accelerating Deep Neural Networks: Survey. IEEE Access, 10, 131788 - 131828.	en_US
dc.identifier.issn	2169-3536
dc.identifier.uri	https://hdl.handle.net/11250/3040701
dc.description.abstract	In the modern-day era of technology, a paradigm shift has been witnessed in the areas involving applications of Artificial Intelligence (AI), Machine Learning (ML), and Deep Learning (DL). Specifically, Deep Neural Networks (DNNs) have emerged as a popular field of interest in most AI applications such as computer vision, image and video processing, robotics, etc. In the context of developed digital technologies and the availability of authentic data and data handling infrastructure, DNNs have been a credible choice for solving more complex real-life problems. The performance and accuracy of a DNN is a way better than human intelligence in certain situations. However, it is noteworthy that the DNN is computationally too cumbersome in terms of the resources and time to handle these computations. Furthermore, general-purpose architectures like CPUs have issues in handling such computationally intensive algorithms. Therefore, a lot of interest and efforts have been invested by the research fraternity in specialized hardware architectures such as Graphics Processing Unit (GPU), Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), and Coarse Grained Reconfigurable Array (CGRA) in the context of effective implementation of computationally intensive algorithms. This paper brings forward the various research works carried out on the development and deployment of DNNs using the aforementioned specialized hardware architectures and embedded AI accelerators. The review discusses the detailed description of the specialized hardware-based accelerators used in the training and/or inference of DNN. A comparative study based on factors like power, area, and throughput, is also made on the various accelerators discussed. Finally, future research and development directions are discussed, such as future trends in DNN implementation on specialized hardware accelerators. This review article is intended to serve as a guide for hardware architectures for accelerating and improving the effectiveness of deep learning research.	en_US
dc.language.iso	eng	en_US
dc.publisher	IEEE	en_US
dc.rights	Navngivelse 4.0 Internasjonal	*
dc.rights.uri	http://creativecommons.org/licenses/by/4.0/deed.no	*
dc.title	Efficient Hardware Architectures for Accelerating Deep Neural Networks: Survey	en_US
dc.type	Peer reviewed	en_US
dc.type	Journal article	en_US
dc.description.version	publishedVersion	en_US
dc.rights.holder	© 2022 The Author(s)	en_US
dc.subject.nsi	VDP::Teknologi: 500	en_US
dc.subject.nsi	VDP::Teknologi: 500::Informasjons- og kommunikasjonsteknologi: 550	en_US
dc.source.pagenumber	131788 - 131828	en_US
dc.source.volume	10	en_US
dc.source.journal	IEEE Access	en_US
dc.identifier.doi	10.1109/ACCESS.2022.3229767
dc.identifier.cristin	2094131
dc.relation.project	Norges forskningsråd: 287918	en_US
dc.relation.project	The Seed Grant of IIT Bhubaneswar (TAML: Timing Analysis with Machine Learning): SP088.	en_US
cristin.qualitycode	1

Files in this item

Name:: Article.pdf
Size:: 6.121Mb
Format:: PDF

View/Open

This item appears in the following Collection(s)

Show simple item record

Except where otherwise noted, this item's license is described as Navngivelse 4.0 Internasjonal