상세 보기
SQUAD: A Scalable Quantization Accelerator Toward Energy-Efficient On-Device Quantization-Aware Training
- Kwon, Ye-Bin;
- Lee, Young Seo;
- Gong, Young-Ho
WEB OF SCIENCE
0SCOPUS
0초록
Quantization-aware training (QAT) is a crucial technique for maintaining model accuracy at extremely low-bit precision (i.e., 4-bit or lower). However, the fake quantization operation of QAT introduces significant training time overhead compared to post-training quantization (PTQ). To address this challenge, we propose SQUAD, a scalable hardware accelerator designed to perform fake quantization operations in any precision at high-performance and energy efficiency. Our proposed SQUAD consists of multiple parallel and pipelined quantization accelerator cores (Agents), with peripherals such as buffer, control units, etc. According to our analysis on QAT of ResNet, MobileNetV2, and DeiT, SQUAD reduces the fake quantization latency by 93.1% and energy consumption by 97.6%, compared to a commercial edge GPU. When SQUAD is applied to QAT on an edge GPU, it significantly accelerates the forward pass, resulting in 48.8% lower latency and 51.2% lower energy consumption. Consequently, these improvements contribute to an overall training latency reduction and energy reduction of 14.6% and 15.3%, with only 0.43% area overhead.
키워드
- 제목
- SQUAD: A Scalable Quantization Accelerator Toward Energy-Efficient On-Device Quantization-Aware Training
- 저자
- Kwon, Ye-Bin; Lee, Young Seo; Gong, Young-Ho
- 발행일
- 2025-09
- 유형
- Article
- 저널명
- IEEE Access
- 권
- 13
- 페이지
- 159487 ~ 159498