상세 보기
TransCrowd-AVL: Token-Level Audio and Vision-Language Fusion for Robust Crowd Counting Under Degraded Conditions
- Lee, Geunwo;
- Cho, Won-Yang;
- Lee, Sangjun
WEB OF SCIENCE
0초록
Crowd counting underpins urban-safety and smart-city applications, yet vision-only counters degrade severely under low light, sensor noise, or partial occlusion, common in real-world deployments. Prior audio-visual approaches inject ambient sound through feature-wise modulation on a convolutional backbone, but this reduces to a per-channel affine transform and cannot encode high-level scene semantics. We propose TransCrowd-AVL, a trimodal crowd counter that adds two token-level side channels to a TransCrowd Vision Transformer (ViT) backbone: an audio token from a frozen audio encoder and a scene-context token from a frozen vision-language model conditioned on a short scene description. The work extends the FiLM and channel-concat fusion of prior convolutional audio-visual counters to the transformer-token level, fusing the two tokens either by per-block FiLM, where their sum scales and shifts each patch token at every block, or by prepending both as extra tokens after the CLS. Experiments on the DISCO (AC) benchmark across four image perturbations, three density bins, and eight baselines show that TransCrowd-AVL attains the lowest mean absolute error (MAE) among the compared models under the two most degraded conditions, low light and high noise, and is the only one to improve the mean MAE over the baseline on every degraded condition, reducing dense-region MAE from 38.52 to 33.92 under low light and from 106.46 to 25.88 under high noise, a 75.7% reduction over the convolutional baseline. A subvariant analysis further shows that the language channel helps only when the audio encoder is representationally complementary to the scene context, offering practical guidance for multimodal counter design.
키워드
- 제목
- TransCrowd-AVL: Token-Level Audio and Vision-Language Fusion for Robust Crowd Counting Under Degraded Conditions
- 저자
- Lee, Geunwo; Cho, Won-Yang; Lee, Sangjun
- 발행일
- 2026-08
- 유형
- Article
- 저널명
- IEEE Access
- 권
- 14
- 페이지
- 124297 ~ 124310