Prosody Disruptor for Voice Protection Against Unauthorized Speech Synthesis

  • Park, Seoyoung
  • Nguyen, An Thien
  • Doan, Thien-Phuc
  • Jung, Souhwan
Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Recent advances in neural speech synthesis enable highly realistic speech generation, raising concerns about unauthorized voice impersonation. To prevent this, proactive voice protection introduces imperceptible perturbations into speech signals that hinder synthesis models from learning a target speaker's identity. However, many existing approaches rely on complex optimization procedures and multiple auxiliary models, resulting in high computational cost. This paper proposes a lightweight prosody-driven voice protection method that disrupts prosodic cues used for speaker adaptation during neural speech synthesis training. The method simply perturbs pitch confidence and energy contours of the original waveform using a fixed pitch tracking model, enabling efficient and lightweight protection. Since prosodic features play a role in representing speaker identity, corrupting these cues reduces speaker similarity in synthesized speech while preserving perceptual quality. Experiments across multiple neural speech synthesis architectures demonstrate that the proposed approach achieves comparable protection performance while enabling a fourfold reduction in processing time.

키워드

BroadcastingDigital audio broadcastingBroadcast technologyFilteringFiltersCircuits and systemsVideosDeepfakesCommunications technologyInformation and communication technologyAdversarial machine learningbiometric authenticationdeep learninggenerative AIinformation securitymachine learningspeech processingspeech synthesisspeech spoofingvoice biometrics
제목
Prosody Disruptor for Voice Protection Against Unauthorized Speech Synthesis
저자
Park, SeoyoungNguyen, An ThienDoan, Thien-PhucJung, Souhwan
DOI
10.1109/ACCESS.2026.3688217
발행일
2026-04
유형
Article
저널명
IEEE Access
14
페이지
65537 ~ 65551