English
Related papers

Related papers: Audio Inpainting in Time-Frequency Domain with Pha…

200 papers

Spatial sound field interpolation relies on suitable models to both conform to available measurements and predict the sound field in the domain of interest. A suitable model can be difficult to determine when the spatial domain of interest…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-30 Manuel Hahmann , Efren Fernandez-Grande

Ear occlusions (arising from the presence of ear accessories such as earrings and earphones) can negatively impact performance in ear-based biometric recognition systems, especially in unconstrained imaging circumstances. In this study, we…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 Deeksha Arun , Kevin W. Bowyer , Patrick Flynn

Contemporary deep learning based inpainting algorithms are mainly based on a hybrid dual stage training policy of supervised reconstruction loss followed by an unsupervised adversarial critic loss. However, there is a dearth of literature…

Computer Vision and Pattern Recognition · Computer Science 2019-08-19 Avisek Lahiri , Arnav Kumar Jain , Prabir Kumar Biswas

Diffusion-based inpainting can reconstruct missing image areas with high quality from sparse data, provided that their location and their values are well optimised. This is particularly useful for applications such as image compression,…

Image and Video Processing · Electrical Eng. & Systems 2023-03-24 Pascal Peter , Karl Schrader , Tobias Alt , Joachim Weickert

Recently, a novel form of audio partial forgery has posed challenges to its forensics, requiring advanced countermeasures to detect subtle forgery manipulations within long-duration audio. However, existing countermeasures still serve a…

Multimedia · Computer Science 2024-07-24 Junyan Wu , Wei Lu , Xiangyang Luo , Rui Yang , Qian Wang , Xiaochun Cao

The paper presents a unified, flexible framework for the tasks of audio inpainting, declipping, and dequantization. The concept is further extended to cover analogous degradation models in a transformed domain, e.g. quantization of the…

Audio and Speech Processing · Electrical Eng. & Systems 2021-01-05 Ondřej Mokrý , Pavel Rajmic , Pavel Záviška

Many studies combine text and audio to capture multi-modal information but they overlook the model's generalization ability on new datasets. Introducing new datasets may affect the feature space of the original dataset, leading to…

Sound · Computer Science 2025-07-29 Yingfei Sun , Xu Gu , Wei Ji , Hanbin Zhao , Yifang Yin , Roger Zimmermann

Speech in-painting is the task of regenerating missing audio contents using reliable context information. Despite various recent studies in multi-modal perception of audio in-painting, there is still a need for an effective infusion of…

Sound · Computer Science 2024-06-04 Mahsa Kadkhodaei Elyaderani , Shahram Shirani

In audio processing applications, phase retrieval (PR) is often performed from the magnitude of short-time Fourier transform (STFT) coefficients. Although PR performance has been observed to depend on the considered STFT parameters and…

Signal Processing · Electrical Eng. & Systems 2021-06-10 Andrés Marafioti , Nicki Holighaus , Piotr Majdak

We introduce a model-agnostic forward diffusion process for time-series forecasting that decomposes signals into spectral components, preserving structured temporal patterns such as seasonality more effectively than standard diffusion.…

Machine Learning · Statistics 2026-02-17 Francisco Caldas , Sahil Kumar , Cláudia Soares

Anomalous audio in speech recordings is often caused by speaker voice distortion, external noise, or even electric interferences. These obstacles have become a serious problem in some fields, such as high-quality music mixing and speech…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-11 Qiang Huang , Thomas Hain

This letter introduces a novel speech enhancement method in the Hilbert-Huang Transform domain to mitigate the effects of acoustic impulsive noises. The estimation and selection of noise components is based on the impulsiveness index of…

Audio and Speech Processing · Electrical Eng. & Systems 2019-10-08 C. Medina , R. Coelho

Modern approaches to sound synthesis using deep neural networks are hard to control, especially when fine-grained conditioning information is not available, hindering their adoption by musicians. In this paper, we cast the generation of…

Sound · Computer Science 2021-04-16 Théis Bazin , Gaëtan Hadjeres , Philippe Esling , Mikhail Malt

Although recent inpainting approaches have demonstrated significant improvements with deep neural networks, they still suffer from artifacts such as blunt structures and abrupt colors when filling in the missing regions. To address these…

Computer Vision and Pattern Recognition · Computer Science 2021-04-20 Tengfei Wang , Hao Ouyang , Qifeng Chen

We study timbre transfer as an inference-time editing problem for music audio. Starting from a strong pre-trained latent diffusion model, we introduce a lightweight procedure that requires no additional training: (i) a dimension-wise noise…

Sound · Computer Science 2026-01-29 Ching Ho Lee , Javier Nistal , Stefan Lattner , Marco Pasini , George Fazekas

One of the challenges in computational acoustics is the identification of models that can simulate and predict the physical behavior of a system generating an acoustic signal. Whenever such models are used for commercial applications an…

Photoacoustic tomography is a hybrid biomedical technology, which combines the advantages of acoustic and optical imaging. However, for the conventional image reconstruction method, the image quality is affected obviously by artifacts under…

Image and Video Processing · Electrical Eng. & Systems 2024-06-26 Bowei Yao , Yi Zeng , Haizhao Dai , Qing Wu , Youshen Xiao , Fei Gao , Yuyao Zhang , Jingyi Yu , Xiran Cai

Photoacoustic tomography (PAT) is a hybrid medical imaging technique that offer high contrast and a high spatial resolution. One challenging mathematical problem associated with PAT is reconstructing the initial pressure of the wave…

Numerical Analysis · Mathematics 2024-07-16 Gyeongha Hwang , Gihyeon Jeon , Sunghwan Moon , Dabin Park

The high-intensity, repetitive noise associated with functional magnetic resonance imaging hinders on-line monitoring of subjects' speech and/or recording speech signals suitable for off-line analysis. The proposed algorithm enhances the…

Sound · Computer Science 2012-07-26 Satrajit S. Ghosh

Packet loss is a major cause of voice quality degradation in VoIP transmissions with serious impact on intelligibility and user experience. This paper describes a system based on a generative adversarial approach, which aims to repair the…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-31 Carlo Aironi , Samuele Cornell , Luca Serafini , Stefano Squartini
‹ Prev 1 4 5 6 7 8 10 Next ›