English
Related papers

Related papers: Upsampling artifacts in neural audio synthesis

200 papers

Speech deepfake detection has recently gained significant attention within the multimedia forensics community. Related issues have also been explored, such as the identification of partially fake signals, i.e., tracks that include both real…

Sound · Computer Science 2024-08-27 Viola Negroni , Davide Salvi , Paolo Bestagini , Stefano Tubaro

Deep convolutional neural networks (CNNs) have been actively adopted in the field of music information retrieval, e.g. genre classification, mood detection, and chord recognition. However, the process of learning and prediction is little…

Machine Learning · Computer Science 2016-07-11 Keunwoo Choi , George Fazekas , Mark Sandler

Deep learning is currently the most widespread and successful technology in artificial intelligence. It promises to push the frontier of scientific discovery beyond current limits. However, skeptics have worried that deep neural networks…

Machine Learning · Computer Science 2020-03-27 Cameron Buckner

Several recent polyphonic music transcription systems have utilized deep neural networks to achieve state of the art results on various benchmark datasets, pushing the envelope on framewise and note-level performance measures. Unfortunately…

Sound · Computer Science 2017-02-02 Rainer Kelz , Gerhard Widmer

Current methods for performing 3D reconstruction and novel view synthesis (NVS) in ultrasound imaging data often face severe artifacts when training NeRF-based approaches. The artifacts produced by current approaches differ from NeRF…

Computer Vision and Pattern Recognition · Computer Science 2024-08-22 Rishit Dagli , Atsuhiro Hibi , Rahul G. Krishnan , Pascal N. Tyrrell

We propose the Neuralogram -- a deep neural network based representation for understanding audio signals which, as the name suggests, transforms an audio signal to a dense, compact representation based upon embeddings learned via a neural…

Sound · Computer Science 2019-04-11 Prateek Verma , Chris Chafe , Jonathan Berger

Advances in techniques for thermal sampling in classical and quantum systems would deepen understanding of the underlying physics. Unfortunately, one often has to rely solely on inexact numerical simulation, due to the intractability of…

Quantum Physics · Physics 2020-04-10 Jeffrey Marshall , Andrea Di Gioacchino , Eleanor G. Rieffel

Speech enhancement using artificial neural networks aims to remove noise from noisy speech signals while preserving the speech content. However, speech enhancement networks often introduce distortions to the speech signal, referred to as…

Sound · Computer Science 2025-08-15 Iksoon Jeong , Kyung-Joong Kim , Kang-Hun Ahn

An direction of development in the extraction of features from audio signals is based on processing raw samples in the time domain. Such an approach appears to be effective, especially in the era of neural networks. An example is SincNet.…

Sound · Computer Science 2026-04-22 Waldek Maciejko

We consider the problem of audio voice separation for binaural applications, such as earphones and hearing aids. While today's neural networks perform remarkably well (separating $4+$ sources with 2 microphones) they assume a known or fixed…

Sound · Computer Science 2022-07-18 Zhongweiyang Xu , Romit Roy Choudhury

Ptychography is a well-studied phase imaging method that makes non-invasive imaging possible at a nanometer scale. It has developed into a mainstream technique with various applications across a range of areas such as material science or…

Image and Video Processing · Electrical Eng. & Systems 2022-08-01 Semih Barutcu , Aggelos K. Katsaggelos , Doğa Gürsoy

We propose a novel universal detector for detecting images generated by using CNNs. In this paper, properties of checkerboard artifacts in CNN-generated images are considered, and the spectrum of images is enhanced in accordance with the…

Computer Vision and Pattern Recognition · Computer Science 2021-08-05 Miki Tanaka , Sayaka Shiota , Hitoshi Kiya

Deepfake speech attribution remains challenging for existing solutions. Classifier-based solutions often fail to generalize to domain-shifted samples, and watermarking-based solutions are easily compromised by distortions like codec…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-16 Wanying Ge , Xin Wang , Junichi Yamagishi

Self-supervised representations excel at many vision and speech tasks, but their potential for audio-visual deepfake detection remains underexplored. Unlike prior work that uses these features in isolation or buried within complex…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Dragos-Alexandru Boldisor , Stefan Smeu , Dan Oneata , Elisabeta Oneata

Given an input sound signal and a target virtual sound source, sound spatialisation algorithms manipulate the signal so that a listener perceives it as though it were emitted from the target source. There exist several established…

Sound · Computer Science 2017-11-28 Ali Tarzan , Marco Alunno , Paolo Bientinesi

Time-of-Flight (ToF) depth sensing camera is able to obtain depth maps at a high frame rate. However, its low resolution and sensitivity to the noise are always a concern. A popular solution is upsampling the obtained noisy low resolution…

Computer Vision and Pattern Recognition · Computer Science 2015-06-18 Wei Liu , Yijun Li , Xiaogang Chen , Jie Yang , Qiang Wu , Jingyi Yu

Deepfake is content or material that is synthetically generated or manipulated using artificial intelligence (AI) methods, to be passed off as real and can include audio, video, image, and text synthesis. This survey has been conducted with…

Sound · Computer Science 2021-11-30 Zahra Khanjani , Gabrielle Watson , Vandana P. Janeja

Music segmentation refers to the dual problem of identifying boundaries between, and labeling, distinct music segments, e.g., the chorus, verse, bridge etc. in popular music. The performance of a range of music segmentation algorithms has…

Sound · Computer Science 2021-08-31 Matthew C. McCallum

This paper introduces two ongoing projects where audio augmented reality is implemented as a means of engaging museum and gallery visitors with audio archive material and associated objects, artworks and artefacts. It outlines some of the…

Human-Computer Interaction · Computer Science 2024-12-13 Laurence Cliffe , James Mansell , Joanne Cormac , Chris Greenhalgh , Adrian Hazzard

The applications of Electroencephalogram (EEG) have been extended to out of laboratory and clinics recently due to the advancements in the technical capabilities. There are various advantageous of EEG, making it a preferable method for a…

Signal Processing · Electrical Eng. & Systems 2021-09-14 Ibrahim Kaya
‹ Prev 1 3 4 5 6 7 10 Next ›