상세 보기
Performance and Efficiency Gains of NPU-Based Servers over GPUs for AI Model Inference
- Hong, Youngpyo;
- Kim, Dongsoo
WEB OF SCIENCE
0SCOPUS
1초록
The exponential growth of AI applications has intensified the demand for efficient inference hardware capable of delivering low-latency, high-throughput, and energy-efficient performance. This study presents a systematic, empirical comparison of GPU- and NPU-based server platforms across key AI inference domains: text-to-text, text-to-image, multimodal understanding, and object detection. We configure representative models-LLama-family for text generation, Stable Diffusion variants for image synthesis, LLaVA-NeXT for multimodal tasks, and YOLO11 series for object detection-on a dual NVIDIA A100 GPU server and an eight-chip RBLN-CA12 NPU server. Performance metrics including latency, throughput, power consumption, and energy efficiency are measured under realistic workloads. Results demonstrate that NPUs match or exceed GPU throughput in many inference scenarios while consuming 35-70% less power. Moreover, optimization with the vLLM library on NPUs nearly doubles the tokens-per-second and yields a 92% increase in power efficiency. Our findings validate the potential of NPU-based inference architectures to reduce operational costs and energy footprints, offering a viable alternative to the prevailing GPU-dominated paradigm.
키워드
- 제목
- Performance and Efficiency Gains of NPU-Based Servers over GPUs for AI Model Inference
- 저자
- Hong, Youngpyo; Kim, Dongsoo
- 발행일
- 2025-09
- 유형
- Article
- 저널명
- SYSTEMS
- 권
- 13
- 호
- 9