SQUAD: A Scalable Quantization Accelerator Toward Energy-Efficient On-Device Quantization-Aware Training

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Quantization-aware training (QAT) is a crucial technique for maintaining model accuracy at extremely low-bit precision (i.e., 4-bit or lower). However, the fake quantization operation of QAT introduces significant training time overhead compared to post-training quantization (PTQ). To address this challenge, we propose SQUAD, a scalable hardware accelerator designed to perform fake quantization operations in any precision at high-performance and energy efficiency. Our proposed SQUAD consists of multiple parallel and pipelined quantization accelerator cores (Agents), with peripherals such as buffer, control units, etc. According to our analysis on QAT of ResNet, MobileNetV2, and DeiT, SQUAD reduces the fake quantization latency by 93.1% and energy consumption by 97.6%, compared to a commercial edge GPU. When SQUAD is applied to QAT on an edge GPU, it significantly accelerates the forward pass, resulting in 48.8% lower latency and 51.2% lower energy consumption. Consequently, these improvements contribute to an overall training latency reduction and energy reduction of 14.6% and 15.3%, with only 0.43% area overhead.

키워드

Quantization (signal)TrainingAccuracyComputational modelingGraphics processing unitsAdaptation modelsHardware accelerationArtificial neural networksEnergy efficiencyTransformersEdge devicehardware acceleratorneural network accelerationquantization-aware training
제목
SQUAD: A Scalable Quantization Accelerator Toward Energy-Efficient On-Device Quantization-Aware Training
저자
Kwon, Ye-BinLee, Young SeoGong, Young-Ho
DOI
10.1109/ACCESS.2025.3608445
발행일
2025-09
유형
Article
저널명
IEEE Access
13
페이지
159487 ~ 159498