Dual-Channel Deepfake Audio Detection: Leveraging Direct and Reverberant Waveforms

Citations

WEB OF SCIENCE

1
Citations

SCOPUS

6

초록

Deepfake content-including audio, video, images, and text-synthesized or modified using artificial intelligence is designed to convincingly mimic real content. As deepfake generation technology advances, detecting deepfake content presents significant challenges. While recent progress has been made in detection techniques, identifying deepfake audio remains particularly challenging. Previous approaches have attempted to capture deepfake features by combining video and audio content; however, these methods are ineffective when video and audio are mismatched due to occlusion. To address this, we propose a novel dual-channel deepfake audio detection model that leverages the direct and reverberant components extracted from raw audio signals, focusing exclusively on audio-based detection without reliance on video content. Across various datasets, including ASVspoof2019, FakeAVCeleb, and sport press conference datasets collected by our group, the proposed dual-channel model demonstrates significant improvements in quantitative metrics such as equal error rate and area under the curve. The implementation is available at https://github.com/gunwoo5034/Dual-Channel-Audio-Deepfake-Detection.

키워드

DeepfakesSpectrogramReflectionPipelinesSportsSolid modelingReceiversPressesFeature extractionTransformersDeepfake audio detectiondual-channel datadirect waveformreverberant waveform
제목
Dual-Channel Deepfake Audio Detection: Leveraging Direct and Reverberant Waveforms
저자
Lee, GunwooLee, JungminJung, MinkyoLee, JosephHong, KihunJung, SouhwanHan, Yoseob
DOI
10.1109/ACCESS.2025.3532775
발행일
2025-01
유형
Article
저널명
IEEE Access
13
페이지
18040 ~ 18052