English
Related papers

Related papers: Multi-Granularity Adaptive Time-Frequency Attentio…

200 papers

We introduce a novel, general-purpose audio generation framework specifically designed for anomaly detection and localization. Unlike existing datasets that predominantly focus on industrial and machine-related sounds, our framework focuses…

There are increasing concerns about malicious attacks on autonomous vehicles. In particular, inaudible voice command attacks pose a significant threat as voice commands become available in autonomous driving systems. How to empirically…

Cryptography and Security · Computer Science 2023-06-09 Jiwei Guan , Lei Pan , Chen Wang , Shui Yu , Longxiang Gao , Xi Zheng

Denoising diffusion models have shown remarkable potential in various generation tasks. The open-source large-scale text-to-image model, Stable Diffusion, becomes prevalent as it can generate realistic artistic or facial images with…

Computer Vision and Pattern Recognition · Computer Science 2023-05-09 Ruijia Wu , Yuhang Wang , Huafeng Shi , Zhipeng Yu , Yichao Wu , Ding Liang

Target speaker extraction, which aims at extracting a target speaker's voice from a mixture of voices using audio, visual or locational clues, has received much interest. Recently an audio-visual target speaker extraction has been proposed…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-03 Hiroshi Sato , Tsubasa Ochiai , Keisuke Kinoshita , Marc Delcroix , Tomohiro Nakatani , Shoko Araki

Test-time domain adaption (TTDA) for semantic segmentation aims to adapt a segmentation model trained on a source domain to a target domain for inference on-the-fly, where both efficiency and effectiveness are critical. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Taorong Liu , Zhen Zhang , Liang Liao , Jing Xiao , Chia-Wen Lin

Decoding the attended speaker in a multi-speaker environment from electroencephalography (EEG) has attracted growing interest in recent years, with neuro-steered hearing devices as a driver application. Current approaches typically rely on…

Signal Processing · Electrical Eng. & Systems 2026-02-05 Yuanyuan Yao , Simon Geirnaert , Tinne Tuytelaars , Alexander Bertrand

Several speech processing systems have demonstrated considerable performance improvements when deep complex neural networks (DCNN) are coupled with self-attention (SA) networks. However, the majority of DCNN-based studies on speech…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-24 Vinay Kothapally , John H. L. Hansen

In this paper, we propose Localized Artifact Attention X (LAA-X), a novel deepfake detection framework that is both robust to high-quality forgeries and capable of generalizing to unseen manipulations. Existing approaches typically rely on…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Dat Nguyen , Enjie Ghorbel , Anis Kacem , Marcella Astrid , Djamila Aouada

Current deepfake detection models achieve state-of-the-art performance on pristine academic datasets but suffer severe spatial attention drift under real-world compound degradations, such as blurring and severe lossy compression. To address…

Computer Vision and Pattern Recognition · Computer Science 2026-04-29 Minh-Khoa Le-Phan , Minh-Hoang Le , Trong-Le Do , Minh-Triet Tran

The rapid advancements in computer vision have stimulated remarkable progress in face forgery techniques, capturing the dedicated attention of researchers committed to detecting forgeries and precisely localizing manipulated areas.…

Computer Vision and Pattern Recognition · Computer Science 2023-06-30 Yingxin Lai , Zhiming Luo , Zitong Yu

Recent synthetic speech detection models typically adapt a pre-trained SSL model via finetuning, which is computationally demanding. Parameter-Efficient Fine-Tuning (PEFT) offers an alternative. However, existing methods lack the specific…

Sound · Computer Science 2025-10-30 Yassine El Kheir , Fabian Ritter-Guttierez , Arnab Das , Tim Polzehl , Sebastian Möller

This paper presents a comprehensive analysis of an enhanced asynchronous AdaBoost framework for federated learning (FL), focusing on its application across five distinct domains: computer vision on edge devices, blockchain-based model…

Machine Learning · Computer Science 2025-06-12 Arthur Oghlukyan , Nuria Gomez Blas

Text-to-audio (TTA) generation can significantly benefit the media industry by reducing production costs and enhancing work efficiency. However, most current TTA models (primarily diffusion-based) suffer from slow inference speeds and high…

Sound · Computer Science 2025-12-30 HaeChun Chung

The rapid development of audio-driven talking head generators and advanced Text-To-Speech (TTS) models has led to more sophisticated temporal deepfakes. These advances highlight the need for robust methods capable of detecting and…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-12 Ivan Kukanov , Jun Wah Ng

The continuous development of new adaptive filters (AFs) based on novel cost functions (CFs) is driven by the demands of various application scenarios and noise environments. However, these algorithms typically demonstrate optimal…

Signal Processing · Electrical Eng. & Systems 2025-06-03 Yi Peng , Haiquan Zhao , Jinhui Hu

The malicious misuse and widespread dissemination of AI-generated images pose a significant threat to the authenticity of online information. Current detection methods often struggle to generalize to unseen generative models, and the rapid…

Computer Vision and Pattern Recognition · Computer Science 2026-01-12 Hanyi Wang , Jun Lan , Yaoyu Kang , Huijia Zhu , Weiqiang Wang , Zhuosheng Zhang , Shilin Wang

With the development of deep neural networks, the performance of crowd counting and pixel-wise density estimation are continually being refreshed. Despite this, there are still two challenging problems in this field: 1) current supervised…

Computer Vision and Pattern Recognition · Computer Science 2020-10-28 Junyu Gao , Yuan Yuan , Qi Wang

This paper describes the deepfake audio detection system submitted to the Audio Deep Synthesis Detection (ADD) Challenge Track 3.2 and gives an analysis of score fusion. The proposed system is a score-level fusion of several light…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-14 Yuxiang Zhang , Jingze Lu , Xingming Wang , Zhuo Li , Runqiu Xiao , Wenchao Wang , Ming Li , Pengyuan Zhang

Acoustic Echo Cancellation (AEC) plays a key role in speech interaction by suppressing the echo received at microphone introduced by acoustic reverberations from loudspeakers. Since the performance of linear adaptive filter (AF) would…

Sound · Computer Science 2021-06-02 Lu Ma , Song Yang , Yaguang Gong , Zhongqin Wu

Graph anomaly detection (GAD) has garnered increasing attention in recent years, yet remains challenging due to two key factors: (1) label scarcity stemming from the high cost of annotations and (2) homophily disparity at node and class…

Machine Learning · Computer Science 2026-01-30 Yunhui Liu , Jiashun Cheng , Yiqing Lin , Qizhuo Xie , Jia Li , Fugee Tsung , Hongzhi Yin , Tao Zheng , Jianhua Zhao , Tieke He