中文
相关论文

相关论文: Upsampling artifacts in neural audio synthesis

200 篇论文

Upsampling artifacts are caused by problematic upsampling layers and due to spectral replicas that emerge while upsampling. Also, depending on the used upsampling layer, such artifacts can either be tonal artifacts (additive high-frequency…

声音 · 计算机科学 2021-11-24 Jordi Pons , Joan Serrà , Santiago Pascual , Giulio Cengarle , Daniel Arteaga , Davide Scaini

Pixel-wise predictions are required in a wide variety of tasks such as image restoration, image segmentation, or disparity estimation. Common models involve several stages of data resampling, in which the resolution of feature maps is first…

计算机视觉与模式识别 · 计算机科学 2024-07-16 Shashank Agnihotri , Julia Grabinski , Margret Keuper

Spoofed utterances always contain artifacts introduced by generative models. While several countermeasures have been proposed to detect spoofed utterances, most primarily focus on architectural improvements. In this work, we investigate how…

声音 · 计算机科学 2025-06-16 Thanapat Trachu , Thanathai Lertpetchpun , Ekapol Chuangsuwanich

We introduce a new audio processing technique that increases the sampling rate of signals such as speech or music using deep convolutional neural networks. Our model is trained on pairs of low and high-quality audio examples; at test-time,…

声音 · 计算机科学 2017-08-03 Volodymyr Kuleshov , S. Zayd Enam , Stefano Ermon

Photoacoustic imaging (PAI) is rapidly moving from the laboratory to the clinic, increasing the need to understand confounders which might adversely affect patient care. Over the past five years, landmark studies have shown the clinical…

In this paper, we propose a novel convolutional neural network (CNN) that never causes checkerboard artifacts, for image enhancement. In research fields of image-to-image translation problems, it is well-known that images generated by usual…

图像与视频处理 · 电气工程与系统科学 2020-10-26 Yuma Kinoshita , Hitoshi Kiya

The most prominent problem associated with the deconvolution layer is the presence of checkerboard artifacts in output images and dense labels. To combat this problem, smoothness constraints, post processing and different architecture…

计算机视觉与模式识别 · 计算机科学 2017-07-11 Andrew Aitken , Christian Ledig , Lucas Theis , Jose Caballero , Zehan Wang , Wenzhe Shi

There have been several successful deep learning models that perform audio super-resolution. Many of these approaches involve using preprocessed feature extraction which requires a lot of domain-specific signal processing knowledge to…

音频与语音处理 · 电气工程与系统科学 2021-10-01 James King , Ramon Viñas Torné , Alexander Campbell , Pietro Liò

Neural image compression, based on auto-encoders and overfitted representations, relies on a latent representation of the coded signal. This representation needs to be compact and uses low resolution feature maps. In the decoding process,…

图像与视频处理 · 电气工程与系统科学 2024-12-02 Pierrick Philippe , Théo Ladune , Gordon Clare , Félix Henry , Théophile Blard , Thomas Leguay

The advancements of AI-synthesized human voices have introduced a growing threat of impersonation and disinformation. It is therefore of practical importance to developdetection methods for synthetic human voices. This work proposes a new…

声音 · 计算机科学 2023-04-28 Chengzhe Sun , Shan Jia , Shuwei Hou , Ehab AlBadawy , Siwei Lyu

We propose a novel method of efficient upsampling of a single natural image. Current methods for image upsampling tend to produce high-resolution images with either blurry salient edges, or loss of fine textural detail, or spurious noise…

计算机视觉与模式识别 · 计算机科学 2015-03-03 Chinmay Hegde , Oncel Tuzel , Fatih Porikli

Recently, the proliferation of highly realistic synthetic images, facilitated through a variety of GANs and Diffusions, has significantly heightened the susceptibility to misuse. While the primary focus of deepfake detection has…

计算机视觉与模式识别 · 计算机科学 2023-12-21 Chuangchuang Tan , Huan Liu , Yao Zhao , Shikui Wei , Guanghua Gu , Ping Liu , Yunchao Wei

Although recent advances in deep learning technology improved automatic speech recognition (ASR), it remains difficult to recognize speech when it overlaps other people's voices. Speech separation or extraction is often used as a front-end…

音频与语音处理 · 电气工程与系统科学 2022-06-17 Hiroshi Sato , Tsubasa Ochiai , Marc Delcroix , Keisuke Kinoshita , Takafumi Moriya , Naoyuki Kamo

Speech enhancement attenuates interfering sounds in speech signals but may introduce artifacts that perceivably deteriorate the output signal. We propose a method for controlling the trade-off between the attenuation of the interfering…

音频与语音处理 · 电气工程与系统科学 2021-07-23 Christian Uhle , Matteo Torcoli , Jouni Paulus

The rapid rise of generative AI has transformed music creation, with millions of users engaging in AI-generated music. Despite its popularity, concerns regarding copyright infringement, job displacement, and ethical implications have led to…

声音 · 计算机科学 2025-06-25 Darius Afchar , Gabriel Meseguer-Brocal , Kamil Akesbi , Romain Hennequin

Novel Magnetic Resonance (MR) imaging modalities can quantify hemodynamics but require long acquisition times, precluding its widespread use for early diagnosis of cardiovascular disease. To reduce the acquisition times, reconstruction…

图像与视频处理 · 电气工程与系统科学 2022-01-12 Lauren Partin , Daniele E. Schiavazzi , Carlos A. Sing Long

Distance transforms are a central tool in shape analysis, morphometry, and curve evolution problems. This work describes and investigates an artifact present in distance maps computed from sampled signals. Namely, sampling reflects through…

图像与视频处理 · 电气工程与系统科学 2020-11-19 Bryce A. Besler , Tannis D. Kemp , Nils D. Forkert , Steven K. Boyd

Neural fields have rapidly been adopted for representing 3D signals, but their application to more classical 2D image-processing has been relatively limited. In this paper, we consider one of the most important operations in image…

It is challenging to improve automatic speech recognition (ASR) performance in noisy conditions with single-channel speech enhancement (SE). In this paper, we investigate the causes of ASR performance degradation by decomposing the SE…

音频与语音处理 · 电气工程与系统科学 2022-03-31 Kazuma Iwamoto , Tsubasa Ochiai , Marc Delcroix , Rintaro Ikeshita , Hiroshi Sato , Shoko Araki , Shigeru Katagiri

Deep generative modeling has the potential to cause significant harm to society. Recognizing this threat, a magnitude of research into detecting so-called "Deepfakes" has emerged. This research most often focuses on the image domain, while…

机器学习 · 计算机科学 2021-11-05 Joel Frank , Lea Schönherr
‹ 上一页 1 2 3 10 下一页 ›