EchoFake: A Replay-Aware Dataset for Practical Speech Deepfake Detection
Researchers have introduced EchoFake, a comprehensive dataset designed to improve the detection of speech deepfakes in real-world scenarios. Current anti-spoofing systems often fail against physical replay attacks, a common and low-cost method used in telephone fraud and identity theft. Experiments revealed that models trained on existing datasets suffer severe performance degradation, with accuracy dropping to 59.6% when evaluating replayed audio. To address this gap, EchoFake comprises over 120 hours of audio from more than 13,000 speakers. It features both advanced zero-shot text-to-speech synthetic audio and physical replay recordings collected across varied devices and environmental settings. The study evaluates three baseline detection models, demonstrating that those trained on EchoFake achieve lower average Equal Error Rates (EERs) across different datasets, indicating superior generalization capabilities. By incorporating practical challenges relevant to actual deployment, EchoFake provides a more realistic foundation for advancing spoofing detection methods, aiming to enhance security against increasingly sophisticated voice synthesis threats in practical applications.
Wire timeline
EchoFake: A Replay-Aware Dataset for Practical Speech Deepfake Detection
Researchers have introduced EchoFake, a comprehensive dataset designed to improve the detection of speech deepfakes in real-world scenarios. Current anti-spoofing systems often fail against physical replay attacks, a common and low-cost method used in telephone fraud and identity theft. Experiments revealed that models trained on existing datasets suffer severe performance degradation, with accuracy dropping to 59.6% when evaluating replayed audio. To address this gap, EchoFake comprises over 120 hours of audio from more than 13,000 speakers. It features both advanced zero-shot text-to-speech synthetic audio and physical replay recordings collected across varied devices and environmental settings. The study evaluates three baseline detection models, demonstrating that those trained on EchoFake achieve lower average Equal Error Rates (EERs) across different datasets, indicating superior generalization capabilities. By incorporating practical challenges relevant to actual deployment, EchoFake provides a more realistic foundation for advancing spoofing detection methods, aiming to enhance security against increasingly sophisticated voice synthesis threats in practical applications.
cs.AI updates on arXiv.org