English
Related papers

Related papers: A Noval Feature via Color Quantisation for Fake Au…

200 papers

The process of reconstructing missing parts of speech audio from context is called speech in-painting. Human perception of speech is inherently multi-modal, involving both audio and visual (AV) cues. In this paper, we introduce and study a…

Multimedia · Computer Science 2024-06-04 Mahsa Kadkhodaei Elyaderani , Shahram Shirani

Current fake audio detection relies on hand-crafted features, which lose information during extraction. To overcome this, recent studies use direct feature extraction from raw audio signals. For example, RawNet is one of the representative…

Sound · Computer Science 2023-05-24 Chenglong Wang , Jiangyan Yi , Jianhua Tao , Chuyuan Zhang , Shuai Zhang , Ruibo Fu , Xun Chen

We propose DiffQ a differentiable method for model compression for quantizing model parameters without gradient approximations (e.g., Straight Through Estimator). We suggest adding independent pseudo quantization noise to model parameters…

Machine Learning · Statistics 2022-10-18 Alexandre Défossez , Yossi Adi , Gabriel Synnaeve

This paper proposes to use low-level spatial features extracted from multichannel audio for sound event detection. We extend the convolutional recurrent neural network to handle more than one type of these multichannel features by learning…

Sound · Computer Science 2017-06-09 Sharath Adavanne , Pasi Pertilä , Tuomas Virtanen

The existing fake audio detection systems often rely on expert experience to design the acoustic features or manually design the hyperparameters of the network structure. However, artificial adjustment of the parameters can have a…

Recent state-of-the-art neural audio compression models have progressively adopted residual vector quantization (RVQ). Despite this success, these models employ a fixed number of codebooks per frame, which can be suboptimal in terms of…

Objective: Voice disorders significantly compromise individuals' ability to speak in their daily lives. Without early diagnosis and treatment, these disorders may deteriorate drastically. Thus, automatic classification systems at home are…

Audio and Speech Processing · Electrical Eng. & Systems 2023-04-27 Heng-Cheng Kuo , Yu-Peng Hsieh , Huan-Hsin Tseng , Chi-Te Wang , Shih-Hau Fang , Yu Tsao

We present a novel approach for the detection of deepfake videos using a pair of vision transformers pre-trained by a self-supervised masked autoencoding setup. Our method consists of two distinct components, one of which focuses on…

Computer Vision and Pattern Recognition · Computer Science 2024-02-12 Sayantan Das , Mojtaba Kolahdouzi , Levent Özparlak , Will Hickie , Ali Etemad

Sound Event Detection and Audio Classification tasks are traditionally addressed through time-frequency representations of audio signals such as spectrograms. However, the emergence of deep neural networks as efficient feature extractors…

Denoising diffusion models have found applications in image segmentation by generating segmented masks conditioned on images. Existing studies predominantly focus on adjusting model architecture or improving inference, such as test-time…

Image and Video Processing · Electrical Eng. & Systems 2023-12-11 Yunguan Fu , Yiwen Li , Shaheer U Saeed , Matthew J Clarkson , Yipeng Hu

Density reconstruction from X-ray projections is an important problem in radiography with key applications in scientific and industrial X-ray computed tomography (CT). Often, such projections are corrupted by unknown sources of noise and…

Image and Video Processing · Electrical Eng. & Systems 2026-02-26 Siddhant Gautam , Marc L. Klasky , Balasubramanya T. Nadiga , Trevor Wilcox , Gary Salazar , Saiprasad Ravishankar

Audio is one of the most used ways of human communication, but at the same time it can be easily misused to trick people. With the revolution of AI, the related technologies are now accessible to almost everyone, thus making it simple for…

Ultrafast electron beam X-ray computed tomography produces noisy data due to short measurement times, causing reconstruction artifacts and limiting overall image quality. To counteract these issues, two self-supervised deep learning methods…

Machine Learning · Computer Science 2025-11-24 Israt Jahan Tulin , Sebastian Starke , Dominic Windisch , André Bieberle , Peter Steinbach

Super-resolution and denoising are ill-posed yet fundamental image restoration tasks. In blind settings, the degradation kernel or the noise level are unknown. This makes restoration even more challenging, notably for learning-based…

Image and Video Processing · Electrical Eng. & Systems 2020-07-24 Majed El Helou , Ruofan Zhou , Sabine Süsstrunk

Deepfake detection remains a challenging task due to the difficulty of generalizing to new types of forgeries. This problem primarily stems from the overfitting of existing detection methods to forgery-irrelevant features and…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Zhiyuan Yan , Yong Zhang , Yanbo Fan , Baoyuan Wu

Automatic speaker recognition algorithms typically characterize speech audio using short-term spectral features that encode the physiological and anatomical aspects of speech production. Such algorithms do not fully capitalize on…

Sound · Computer Science 2021-02-16 Anurag Chowdhury , Arun Ross , Prabu David

Robust audio anti-spoofing has been increasingly challenging due to the recent advancements on deepfake techniques. While spectrograms have demonstrated their capability for anti-spoofing, complementary information presented in multi-order…

Sound · Computer Science 2024-10-04 Penghui Wen , Kun Hu , Wenxi Yue , Sen Zhang , Wanlei Zhou , Zhiyong Wang

The deepfake generation of singing vocals is a concerning issue for artists in the music industry. In this work, we propose a singing voice deepfake detection (SVDD) system, which uses noise-variant encodings of open-AI's Whisper model. As…

Sound · Computer Science 2025-02-03 Falguni Sharma , Priyanka Gupta

Recent success in speech representation learning enables a new way to leverage unlabeled data to train speech recognition model. In speech representation learning, a large amount of unlabeled data is used in a self-supervised manner to…

Audio and Speech Processing · Electrical Eng. & Systems 2020-12-15 Shaoshi Ling , Yuzong Liu

This paper presents a video encoding method in which noise is encoded using a novel parametric model representing spectral envelope and spatial distribution of energy. The proposed method has been experimentally assessed using video test…

Image and Video Processing · Electrical Eng. & Systems 2019-09-04 Olgierd Stankiewicz
‹ Prev 1 3 4 5 6 7 10 Next ›