English
Related papers

Related papers: SONAR: Spectral-Contrastive Audio Residuals for Ge…

200 papers

Recent advances in Text-to-Speech (TTS) and Voice-Conversion (VC) using generative Artificial Intelligence (AI) technology have made it possible to generate high-quality and realistic human-like audio. This poses growing challenges in…

Sound · Computer Science 2025-03-25 Xiang Li , Pin-Yu Chen , Wenqi Wei

Self-supervised learning (SSL) on large-scale datasets like AudioSet has become the dominant paradigm for audio representation learning. While the continuous influx of new, unlabeled audio presents an opportunity to enrich these static…

Sound · Computer Science 2026-01-26 Yizhou Zhang , Yuan Gao , Wangjin Zhou , Zicheng Yuan , Keisuke Imoto , Tatsuya Kawahara

Audio deepfake detection has become increasingly challenging due to rapid advances in speech synthesis and voice conversion technologies, particularly under channel distortions, replay attacks, and real-world recording conditions. This…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-13 K. A. Shahriar

Recently, pioneer research works have proposed a large number of acoustic features (log power spectrogram, linear frequency cepstral coefficients, constant Q cepstral coefficients, etc.) for audio deepfake detection, obtaining good…

Synthetic Aperture Radar (SAR) images are inherently corrupted by speckle noise, limiting their utility in high-precision applications. While deep learning methods have shown promise in SAR despeckling, most methods employ a single unified…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Ziqing Ma , Chang Yang , Zhichang Guo , Yao Li

Recent advances in synthetic speech have made audio deepfakes increasingly realistic, posing significant security risks. Existing detection methods that rely on a single modality, either raw waveform embeddings or spectral based features,…

Anomalous sound detection (ASD) in the wild requires robustness to distribution shifts such as unseen low-SNR input mixtures of machine and noise types. State-of-the-art systems extract embeddings from an adapted audio encoder and detect…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-30 Phurich Saengthong , Tomoya Nishida , Kota Dohi , Natsuo Yamashita , Yohei Kawaguchi

Acoustic sonar image analysis plays a critical role in object detection and classification, with applications in both civilian and defense domains. Despite the availability of real and synthetic datasets, existing AI models that achieve…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Kamal Basha S , Athira Nambiar

Most of existing audio fingerprinting systems have limitations to be used for high-specific audio retrieval at scale. In this work, we generate a low-dimensional representation from a short unit segment of audio, and couple this fingerprint…

Sound · Computer Science 2021-02-11 Sungkyun Chang , Donmoon Lee , Jeongsoo Park , Hyungui Lim , Kyogu Lee , Karam Ko , Yoonchang Han

State-of-the-art methods for audio generation suffer from fingerprint artifacts and repeated inconsistencies across temporal and spectral domains. Such artifacts could be well captured by the frequency domain analysis over the spectrogram.…

Sound · Computer Science 2021-06-29 Yang Gao , Tyler Vuong , Mahsa Elyasi , Gaurav Bharaj , Rita Singh

This letter introduces a physics-informed self-supervised framework for sonar image despeckling that reformulates despeckling as residual consistency in the homomorphic log domain. By constraining the log-ratio residual to obey…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Swapna Pillai , Siddharth Singh Savner , Sujit Kumar Sahoo

Although recent works on neural vocoder have improved the quality of synthesized audio, there still exists a gap between generated and ground-truth audio in frequency space. This difference leads to spectral artifacts such as hissing noise…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-15 Ji-Hoon Kim , Sang-Hoon Lee , Ji-Hyun Lee , Seong-Whan Lee

An ideal audio retrieval system efficiently and robustly recognizes a short query snippet from an extensive database. However, the performance of well-known audio fingerprinting systems falls short at high signal distortion levels. This…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-22 Anup Singh , Kris Demuynck , Vipul Arora

Overlapping sound events are ubiquitous in real-world environments, but existing end-to-end sound event detection (SED) methods still struggle to detect them effectively. A critical reason is that these methods represent overlapping events…

Sound · Computer Science 2024-01-12 Yadong Guan , Jiqing Han , Hongwei Song , Wenjie Song , Guibin Zheng , Tieran Zheng , Yongjun He

Recent studies have demonstrated that vision models can effectively learn multimodal audio-image representations when paired. However, the challenge of enabling deep models to learn representations from unpaired modalities remains…

Sound · Computer Science 2025-04-15 Yasar Abbas Ur Rehman , Kin Wai Lau , Yuyang Xie , Ma Lan , JiaJun Shen

In this paper, we propose a deep learning based system for the task of deepfake audio detection. In particular, the draw input audio is first transformed into various spectrograms using three transformation methods of Short-time Fourier…

Sound · Computer Science 2024-07-03 Lam Pham , Phat Lam , Truong Nguyen , Huyen Nguyen , Alexander Schindler

Collaborative filtering (CF) recommendation has been significantly advanced by integrating Graph Neural Networks (GNNs) and Graph Contrastive Learning (GCL). However, (i) random edge perturbations often distort critical structural signals…

Machine Learning · Computer Science 2026-03-18 Yixuan Huang , Jiawei Chen , Shengfan Zhang , Zongsheng Cao

The rapid advancement of generative models has enabled highly realistic audio deepfakes, yet current detectors suffer from a critical bias problem, leading to poor generalization across unseen datasets. This paper proposes Artifact-Focused…

Deep learning has enabled highly realistic synthetic speech, raising concerns about fraud, impersonation, and disinformation. Despite rapid progress in neural detectors, transparent baselines are needed to reveal which acoustic cues…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-16 Faheem Ahmad , Ajan Ahmed , Masudul Imtiaz

Audio deepfakes generated by modern TTS and voice conversion systems are increasingly difficult to distinguish from real speech, raising serious risks for security and online trust. While state-of-the-art self-supervised models provide rich…

Sound · Computer Science 2026-02-03 Zhili Nicholas Liang , Soyeon Caren Han , Qizhou Wang , Christopher Leckie
‹ Prev 1 2 3 10 Next ›