English
Related papers

Related papers: RaDur: A Reference-aware and Duration-robust Netwo…

200 papers

Simultaneous speech translation (SST) produces target text incrementally from partial speech input. Recent speech large language models (Speech LLMs) have substantially improved SST quality, yet they still struggle to correctly translate…

Computation and Language · Computer Science 2026-02-02 Jiaxuan Luo , Siqi Ouyang , Lei Li

Facing the complex marine environment, it is extremely challenging to conduct underwater acoustic target recognition (UATR) using ship-radiated noise. Inspired by neural mechanism of auditory perception, this paper provides a new deep…

Sound · Computer Science 2020-12-01 Gang Hu , Kejun Wang , Liangliang Liu

In multi-speaker applications is common to have pre-computed models from enrolled speakers. Using these models to identify the instances in which these speakers intervene in a recording is the task of speaker tracking. In this paper, we…

Sound event detection (SED) is the task of tagging the absence or presence of audio events and their corresponding interval within a given audio clip. While SED can be done using supervised machine learning, where training data is fully…

Sound · Computer Science 2021-02-08 Heinrich Dinkel , Mengyue Wu , Kai Yu

Audio-visual speech recognition (AVSR) combines audio-visual modalities to improve speech recognition, especially in noisy environments. However, most existing methods deploy the unidirectional enhancement or symmetric fusion manner, which…

Multimedia · Computer Science 2025-08-12 Junxiao Xue , Xiaozhen Liu , Xuecheng Wu , Xinyi Yin , Danlei Huang , Fei Yu

Deep learning models usually require a large amount of labeled data to achieve satisfactory performance. In multimedia analysis, domain adaptation studies the problem of cross-domain knowledge transfer from a label rich source domain to a…

Computer Vision and Pattern Recognition · Computer Science 2021-09-10 Lei Zhu , Zhaojing Luo , Wei Wang , Meihui Zhang , Gang Chen , Kaiping Zheng

Despite the remarkable progresses made in deep-learning based depth map super-resolution (DSR), how to tackle real-world degradation in low-resolution (LR) depth maps remains a major challenge. Existing DSR model is generally trained and…

Computer Vision and Pattern Recognition · Computer Science 2020-06-03 Xibin Song , Yuchao Dai , Dingfu Zhou , Liu Liu , Wei Li , Hongdng Li , Ruigang Yang

Localizing sounds and detecting events in different room environments is a difficult task, mainly due to the wide range of reflections and reverberations. When training neural network models with sounds recorded in only a few room…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-06 Yusun Shul , Byeong-Yun Ko , Jung-Woo Choi

SAR image classification naturally has to deal with huge noise and a high dynamic range particularly requiring robust classification models. Additionally, the deployment of these models on edge devices, such as drones and military aircraft,…

At a cocktail party, humans exhibit an impressive ability to direct their attention. The auditory attention detection (AAD) approach seeks to identify the attended speaker by analyzing brain signals, such as EEG signals. However, current…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-19 Sheng Yan , Cunhang fan , Hongyu Zhang , Xiaoke Yang , Jianhua Tao , Zhao Lv

Audio-visual speech recognition (AVSR) aims to transcribe human speech using both audio and video modalities. In practical environments with noise-corrupted audio, the role of video information becomes crucial. However, prior works have…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-15 Sungnyun Kim , Kangwook Jang , Sangmin Bae , Hoirin Kim , Se-Young Yun

Streaming end-to-end multi-talker speech recognition aims at transcribing the overlapped speech from conversations or meetings with an all-neural model in a streaming fashion, which is fundamentally different from a modular-based approach…

Audio and Speech Processing · Electrical Eng. & Systems 2022-01-26 Liang Lu , Jinyu Li , Yifan Gong

We propose a new two-pass E2E speech recognition model that improves ASR performance by training on a combination of paired data and unpaired text data. Previously, the joint acoustic and text decoder (JATD) has shown promising results…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-28 Sepand Mavandadi , Tara N. Sainath , Ke Hu , Zelin Wu

Over the past few decades, extensive research has been devoted to the design of artificial reverberation algorithms aimed at emulating the room acoustics of physical environments. Despite significant advancements, automatic parameter tuning…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-10 Alessandro Ilic Mezza , Riccardo Giampiccolo , Enzo De Sena , Alberto Bernardini

Natural language understanding (NLU) using neural network pipelines often requires additional context that is not solely present in the input data. Through Prior research, it has been evident that NLU benchmarks are susceptible to…

Computation and Language · Computer Science 2024-03-06 Yuxin Zi , Hariram Veeramani , Kaushik Roy , Amit Sheth

The rise of advanced large language models such as GPT-4, GPT-4o, and the Claude family has made fake audio detection increasingly challenging. Traditional fine-tuning methods struggle to keep pace with the evolving landscape of synthetic…

Sound · Computer Science 2024-08-14 Xiaohui Zhang , Jiangyan Yi , Jianhua Tao

Infrared small target detection (ISTD) is highly sensitive to sensor type, observation conditions, and the intrinsic properties of the target. These factors can introduce substantial variations in the distribution of acquired infrared image…

Computer Vision and Pattern Recognition · Computer Science 2025-04-24 Yahao Lu , Yuehui Li , Xingyuan Guo , Shuai Yuan , Yukai Shi , Liang Lin

Target speaker extraction (TSE) relies on a reference cue of the target to extract the target speech from a speech mixture. While a speaker embedding is commonly used as the reference cue, such embedding pre-trained with a large number of…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-12 Ke Zhang , Junjie Li , Shuai Wang , Yangjie Wei , Yi Wang , Yannan Wang , Haizhou Li

The recent popularity of deep neural networks (DNNs) has generated a lot of research interest in performing DNN-related computation efficiently. However, the primary focus is usually very narrow and limited to (i) inference -- i.e. how to…

Machine Learning · Computer Science 2018-04-17 Hongyu Zhu , Mohamed Akrout , Bojian Zheng , Andrew Pelegris , Amar Phanishayee , Bianca Schroeder , Gennady Pekhimenko

In radar systems, tracking targets in low signal-to-noise ratio (SNR) environments is a very important task. There are some algorithms designed for multitarget tracking. Their performances, however, are not satisfactory in low SNR…

Applications · Statistics 2015-05-30 Huisi Tong , Hao Zhang , Huadong Meng , Xiqin Wang