Performance and Efficiency Gains of NPU-Based Servers over GPUs for AI Model Inference

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

1

초록

The exponential growth of AI applications has intensified the demand for efficient inference hardware capable of delivering low-latency, high-throughput, and energy-efficient performance. This study presents a systematic, empirical comparison of GPU- and NPU-based server platforms across key AI inference domains: text-to-text, text-to-image, multimodal understanding, and object detection. We configure representative models-LLama-family for text generation, Stable Diffusion variants for image synthesis, LLaVA-NeXT for multimodal tasks, and YOLO11 series for object detection-on a dual NVIDIA A100 GPU server and an eight-chip RBLN-CA12 NPU server. Performance metrics including latency, throughput, power consumption, and energy efficiency are measured under realistic workloads. Results demonstrate that NPUs match or exceed GPU throughput in many inference scenarios while consuming 35-70% less power. Moreover, optimization with the vLLM library on NPUs nearly doubles the tokens-per-second and yields a 92% increase in power efficiency. Our findings validate the potential of NPU-based inference architectures to reduce operational costs and energy footprints, offering a viable alternative to the prevailing GPU-dominated paradigm.

키워드

AI inferenceNeural Processing Unit (NPU)Graphics Processing Unit (GPU)performance benchmarkingenergy efficiencyheterogeneous computingvLLM optimization
제목
Performance and Efficiency Gains of NPU-Based Servers over GPUs for AI Model Inference
저자
Hong, YoungpyoKim, Dongsoo
DOI
10.3390/systems13090797
발행일
2025-09
유형
Article
저널명
SYSTEMS
13
9