BitBlade: Energy-Efficient Variable Bit-Precision Hardware Accelerator for Quantized Neural Networks

  • Ryu, Sungju
  • Kim, Hyungjun
  • Yi, Wooseok
  • Kim, Eunhwan
  • Kim, Yulhwa
  • 외 2명
Citations

WEB OF SCIENCE

50
Citations

SCOPUS

51

초록

We introduce an area/energy-efficient precisionscalable neural network accelerator architecture. Previous precision-scalable hardware accelerators have limitations such as the under-utilization of multipliers for low bit-width operations and the large area overhead to support various bit precisions. To mitigate the problems, we first propose a bitwise summation, which reduces the area overhead for the bit-width scaling. In addition, we present a channel-wise aligning scheme (CAS) to efficiently fetch inputs and weights from on-chip SRAM buffers and a channel-first and pixel-last tiling (CFPL) scheme to maximize the utilization of multipliers on various kernel sizes. A test chip was implemented in 28-nm CMOS technology, and the experimental results show that the throughput and energy efficiency of our chip are up to 7.7x and 1.64x higher than those of the state-of-the-art designs, respectively. Moreover, additional 1.5-3.4x throughput gains can be achieved using the CFPL method compared to the CAS.

키워드

Computer architectureNeural networksHardware accelerationAddersArraysRandom access memoryThroughputBit-precision scalingbitwise summationchannel-first and pixel-last tiling (CFPL)channel-wise aligningdeep neural networkhardware acceleratormultiply-accumulate unit
제목
BitBlade: Energy-Efficient Variable Bit-Precision Hardware Accelerator for Quantized Neural Networks
저자
Ryu, SungjuKim, HyungjunYi, WooseokKim, EunhwanKim, YulhwaKim, TaesuKim, Jae-Joon
DOI
10.1109/JSSC.2022.3141050
발행일
2022-06
유형
Article
저널명
IEEE Journal of Solid-State Circuits
57
6
페이지
1924 ~ 1935