English
Related papers

Related papers: Upsampling artifacts in neural audio synthesis

200 papers

Banding artifacts, which manifest as staircase-like color bands on pictures or video frames, is a common distortion caused by compression of low-textured smooth regions. These false contours can be very noticeable even on high-quality…

Image and Video Processing · Electrical Eng. & Systems 2020-10-28 Zhengzhong Tu , Jessie Lin , Yilin Wang , Balu Adsumilli , Alan C. Bovik

Augmented listening devices, such as hearing aids and augmented reality headsets, enhance human perception by changing the sounds that we hear. Microphone arrays can improve the performance of listening systems in noisy environments, but…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-28 Ryan M. Corey , Andrew C. Singer

Audio processing methods operating on a time-frequency representation of the signal can introduce unpleasant sounding artifacts known as musical noise. These artifacts are observed in the context of audio coding, speech enhancement, and…

Audio and Speech Processing · Electrical Eng. & Systems 2021-05-28 Matteo Torcoli

Music recordings often suffer from audio quality issues such as excessive reverberation, distortion, clipping, tonal imbalances, and a narrowed stereo image, especially when created in non-professional settings without specialized equipment…

Sound · Computer Science 2026-05-06 Jan Melechovsky , Ambuj Mehrish , Abhinaba Roy , Dorien Herremans

The convolutional neural network (CNN) remains an essential tool in solving computer vision problems. Standard convolutional architectures consist of stacked layers of operations that progressively downscale the image. Aliasing is a…

Image and Video Processing · Electrical Eng. & Systems 2021-02-16 Antônio H. Ribeiro , Thomas B. Schön

It is well-known that a number of excellent super-resolution (SR) methods using convolutional neural networks (CNNs) generate checkerboard artifacts. A condition to avoid the checkerboard artifacts is proposed in this paper. So far,…

Computer Vision and Pattern Recognition · Computer Science 2018-06-08 Yusuke Sugawara , Sayaka Shiota , Hitoshi Kiya

Convolutional layers are an integral part of many deep neural network solutions in computer vision. Recent work shows that replacing the standard convolution operation with mechanisms based on self-attention leads to improved performance on…

Computer Vision and Pattern Recognition · Computer Science 2020-12-21 Souvik Kundu , Hesham Mostafa , Sharath Nittur Sridhar , Sairam Sundaresan

Neural Radiance Field (NeRF), capable of synthesizing high-quality novel viewpoint images, suffers from issues like artifact occurrence due to its fixed sampling points during rendering. This study proposes a method that optimizes sampling…

Computer Vision and Pattern Recognition · Computer Science 2024-10-22 Kazuhiro Ohta , Satoshi Ono

Human auditory perception is compositional in nature -- we identify auditory streams from auditory scenes with multiple sound events. However, such auditory scenes are typically represented using clip-level representations that do not…

Sound · Computer Science 2025-03-04 Sripathi Sridhar , Mark Cartwright

A typical 2D-to-3D pipeline takes multi-view images as input, where a Vision Foundation Model (VFM) extracts features that are spatially upsampled to dense representations for 3D reconstruction. If dense features across views preserve…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Ling Xiao , Yuliang Xiu , Yue Chen , Guoming Wang , Toshihiko Yamasaki

Deep learning appears as an appealing solution for Automatic Synthesizer Programming (ASP), which aims to assist musicians and sound designers in programming sound synthesizers. However, integrating software synthesizers into training…

Sound · Computer Science 2025-09-10 Paolo Combes , Stefan Weinzierl , Klaus Obermayer

Cochlear implant users struggle to understand speech in reverberant environments. To restore speech perception, artifacts dominated by reverberant reflections can be removed from the cochlear implant stimulus. Artifacts can be identified…

Sound · Computer Science 2021-08-16 Lidea K. Shahidi , Leslie M. Collins , Boyla O. Mainsah

Visual artifacts remain a persistent challenge in diffusion models, even with training on massive datasets. Current solutions primarily rely on supervised detectors, yet lack understanding of why these artifacts occur in the first place. In…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Yu Cao , Zengqun Zhao , Ioannis Patras , Shaogang Gong

Convolutional Neural Networks (CNNs) are the current de-facto models used for many imaging tasks due to their high learning capacity as well as their architectural qualities. The ubiquitous UNet architecture provides an efficient and…

Image and Video Processing · Electrical Eng. & Systems 2020-09-30 Demetris Marnerides , Thomas Bashford-Rogers , Kurt Debattista

In NMR spectroscopy, undersampling in the indirect dimensions causes reconstruction artifacts whose size can be bounded using the so-called {\it coherence}. In experiments with multiple indirect dimensions, new undersampling approaches were…

Applications · Statistics 2017-02-08 Hatef Monajemi , David L. Donoho , Jeffrey C. Hoch , Adam D. Schuyler

Image denoising or artefact removal using deep learning is possible in the availability of supervised training dataset acquired in real experiments or synthesized using known noise models. Neither of the conditions can be fulfilled for…

Image and Video Processing · Electrical Eng. & Systems 2020-11-23 Suyog Jadhav , Sebastian Acuña , Krishna Agarwal , Dilip K. prasad

We study how visual artifacts introduced by diffusion-based inpainting affect language generation in vision-language models. We use a two-stage diagnostic setup in which masked image regions are reconstructed and then provided to captioning…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Pratham Yashwante , Davit Abrahamyan , Shresth Grover , Sukruth Rao

Existing convolutional neural networks widely adopt spatial down-/up-sampling for multi-scale modeling. However, spatial up-sampling operators (\emph{e.g.}, interpolation, transposed convolution, and un-pooling) heavily depend on local…

Computer Vision and Pattern Recognition · Computer Science 2022-10-12 Man Zhou , Hu Yu , Jie Huang , Feng Zhao , Jinwei Gu , Chen Change Loy , Deyu Meng , Chongyi Li

Real-world sound scenes consist of time-varying collections of sound sources, each generating characteristic sound events that are mixed together in audio recordings. The association of these constituent sound events with their mixture and…

Jointly training a speech enhancement (SE) front-end and an automatic speech recognition (ASR) back-end has been investigated as a way to mitigate the influence of \emph{processing distortion} generated by single-channel SE on ASR. In this…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-21 Kazuma Iwamoto , Tsubasa Ochiai , Marc Delcroix , Rintaro Ikeshita , Hiroshi Sato , Shoko Araki , Shigeru Katagiri