English
Related papers

Related papers: GAP-URGENet: A Generative-Predictive Fusion Framew…

200 papers

Active Speaker Detection (ASD) aims to identify who is currently speaking in each frame of a video. Most state-of-the-art approaches rely on late fusion to combine visual and audio features, but late fusion often fails to capture…

Computer Vision and Pattern Recognition · Computer Science 2025-12-18 Yu Wang , Juhyung Ha , Frangil M. Ramirez , Yuchen Wang , David J. Crandall

Single-channel speech enhancement algorithms are often used in resource-constrained embedded devices, where low latency and low complexity designs gain more importance. In recent years, researchers have proposed a wide variety of novel…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-29 Nicolás Arrieta Larraza , Niels de Koeijer

We propose Parallel WaveGAN, a distillation-free, fast, and small-footprint waveform generation method using a generative adversarial network. In the proposed method, a non-autoregressive WaveNet is trained by jointly optimizing…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-07 Ryuichi Yamamoto , Eunwoo Song , Jae-Min Kim

Creating high-quality sound effects from videos and text prompts requires precise alignment between visual and audio domains, both semantically and temporally, along with step-by-step guidance for professional audio generation. However,…

Sound · Computer Science 2025-03-31 Haomin Zhang , Sizhe Shan , Haoyu Wang , Zihao Chen , Xiulong Liu , Chaofan Ding , Xinhan Di

Most recently, there has been significant interest in learning contextual representations for various NLP tasks, by leveraging large scale text corpora to train large neural language models with self-supervised learning objectives, such as…

Computation and Language · Computer Science 2020-12-21 Peng Shi , Patrick Ng , Zhiguo Wang , Henghui Zhu , Alexander Hanbo Li , Jun Wang , Cicero Nogueira dos Santos , Bing Xiang

Phase information has a significant impact on speech perceptual quality and intelligibility. However, existing speech enhancement methods encounter limitations in explicit phase estimation due to the non-structural nature and wrapping…

Audio and Speech Processing · Electrical Eng. & Systems 2024-04-02 Ye-Xin Lu , Yang Ai , Zhen-Hua Ling

This paper proposes a novel way of doing audio synthesis at the waveform level using Transformer architectures. We propose a deep neural network for generating waveforms, similar to wavenet. This is fully probabilistic, auto-regressive, and…

Sound · Computer Science 2021-07-09 Prateek Verma , Chris Chafe

In Generalized Linear Estimation (GLE) problems, we seek to estimate a signal that is observed through a linear transform followed by a component-wise, possibly nonlinear and noisy, channel. In the Bayesian optimal setting, Generalized…

Disordered Systems and Neural Networks · Physics 2021-02-03 Luca Saglietti , Yue M. Lu , Carlo Lucibello

The advent of hyper-scale and general-purpose pre-trained models is shifting the paradigm of building task-specific models for target tasks. In the field of audio research, task-agnostic pre-trained models with high transferability and…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-03 Ju-ho Kim , Jungwoo Heo , Hyun-seo Shin , Chan-yeong Lim , Ha-Jin Yu

Recent advances in the design of neural network architectures, in particular those specialized in modeling sequences, have provided significant improvements in speech separation performance. In this work, we propose to use a bio-inspired…

Sound · Computer Science 2021-12-07 Xiaolin Hu , Kai Li , Weiyi Zhang , Yi Luo , Jean-Marie Lemercier , Timo Gerkmann

FullSubNet is our recently proposed real-time single-channel speech enhancement network that achieves outstanding performance on the Deep Noise Suppression (DNS) Challenge dataset. A number of variants of FullSubNet have been proposed, but…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-08 Xiang Hao , Xiaofei Li

Accurate assessment of gastric content from ultrasound is critical for stratifying aspiration risk at induction of general anesthesia. However, traditional methods rely on manual tracing of gastric antra and empirical formulas, which face…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Nu-Fnag Xiao , De-Xing Huang , Le-Tian Wang , Mei-Jiang Gui , Qi Fu , Xiao-Liang Xie , Shi-Qi Liu , Shuangyi Wang , Zeng-Guang Hou , Ying-Wei Wang , Xiao-Hu Zhou

There have been several successful deep learning models that perform audio super-resolution. Many of these approaches involve using preprocessed feature extraction which requires a lot of domain-specific signal processing knowledge to…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-01 James King , Ramon Viñas Torné , Alexander Campbell , Pietro Liò

Recent image restoration methods can be broadly categorized into two classes: (1) regression methods that recover the rough structure of the original image without synthesizing high-frequency details and (2) generative methods that…

Computer Vision and Pattern Recognition · Computer Science 2024-01-02 Hwayoon Lee , Kyoungkook Kang , Hyeongmin Lee , Seung-Hwan Baek , Sunghyun Cho

We explore diverse representations of speech audio, and their effect on a performance of late fusion ensemble of E-Branchformer models, applied to Automatic Speech Recognition (ASR) task. Although it is generally known that ensemble methods…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-04 Marin Jezidžić , Matej Mihelčić

This paper presents the speech restoration and enhancement system created by the 1024K team for the ICASSP 2024 Speech Signal Improvement (SSI) Challenge. Our system consists of a generative adversarial network (GAN) in complex-domain for…

Sound · Computer Science 2024-02-06 Guochen Yu , Runqiang Han , Chenglin Xu , Haoran Zhao , Nan Li , Chen Zhang , Xiguang Zheng , Chao Zhou , Qi Huang , Bing Yu

Grasping is fundamental to robotic manipulation, and recent advances in large-scale grasping datasets have provided essential training data and evaluation benchmarks, accelerating the development of learning-based methods for robust object…

Robotics · Computer Science 2025-07-04 Siyu Ma , Wenxin Du , Chang Yu , Ying Jiang , Zeshun Zong , Tianyi Xie , Yunuo Chen , Yin Yang , Xuchen Han , Chenfanfu Jiang

Efficient audio synthesis is an inherently difficult machine learning task, as human perception is sensitive to both global structure and fine-scale waveform coherence. Autoregressive models, such as WaveNet, model local structure at the…

Speech bandwidth extension (BWE) refers to widening the frequency bandwidth range of speech signals, enhancing the speech quality towards brighter and fuller. This paper proposes a generative adversarial network (GAN) based BWE model with…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-17 Ye-Xin Lu , Yang Ai , Hui-Peng Du , Zhen-Hua Ling

This paper proposes a novel multimodal self-supervised architecture for energy-efficient audio-visual (AV) speech enhancement that integrates Graph Neural Networks with canonical correlation analysis (CCA-GNN). The proposed approach lays…