English
Related papers

Related papers: RaDur: A Reference-aware and Duration-robust Netwo…

200 papers

Recently, it can be noticed that most models based on spiking neural networks (SNNs) only use a same level temporal resolution to deal with speech classification problems, which makes these models cannot learn the information of input data…

Sound · Computer Science 2025-01-03 Qi Zhang , Huamin Wang , Hangchi Shen , Shukai Duan , Shiping Wen , Tingwen Huang

Infrared small target detection (ISTD) remains a long-standing challenge due to weak signal contrast, limited spatial extent, and cluttered backgrounds. Despite performance improvements from convolutional neural networks (CNNs) and Vision…

Computer Vision and Pattern Recognition · Computer Science 2026-01-12 Hongyang Xie , Hongyang He , Victor Sanchez

This paper presents a speech BERT model to extract embedded prosody information in speech segments for improving the prosody of synthesized speech in neural text-to-speech (TTS). As a pre-trained model, it can learn prosody attributes from…

Audio and Speech Processing · Electrical Eng. & Systems 2021-09-15 Liping Chen , Yan Deng , Xi Wang , Frank K. Soong , Lei He

Listening to the audio of TV broadcast signals can be challenging for hearing-impaired as well as normal-hearing listeners, especially when background sounds are prominent or too loud compared to the speech signal. This can result in a…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-04 Nils L. Westhausen , Rainer Huber , Hannah Baumgartner , Ragini Sinha , Jan Rennies , Bernd T. Meyer

Developing a robust speech emotion recognition (SER) system in noisy conditions faces challenges posed by different noise properties. Most previous studies have not considered the impact of human speech noise, thus limiting the application…

Sound · Computer Science 2024-12-18 Jinyi Mi , Xiaohan Shi , Ding Ma , Jiajun He , Takuya Fujimura , Tomoki Toda

Neural transducer (RNNT)-based target-speaker speech recognition (TS-RNNT) directly transcribes a target speaker's voice from a multi-talker mixture. It is a promising approach for streaming applications because it does not incur the extra…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-26 Takafumi Moriya , Hiroshi Sato , Tsubasa Ochiai , Marc Delcroix , Takanori Ashihara , Kohei Matsuura , Tomohiro Tanaka , Ryo Masumura , Atsunori Ogawa , Taichi Asami

Voice timbre attribute detection (vTAD) is the task of determining the relative intensity of timbre attributes between speech utterances. Voice timbre is a crucial yet inherently complex component of speech perception. While deep neural…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-06 Aemon Yat Fei Chiu , Yujia Xiao , Qiuqiang Kong , Tan Lee

Sound Event Detection (SED) plays a vital role in comprehending and perceiving acoustic scenes. Previous methods have demonstrated impressive capabilities. However, they are deficient in learning features of complex scenes from…

Sound · Computer Science 2024-09-12 Zehao Wang , Haobo Yue , Zhicheng Zhang , Da Mu , Jin Tang , Jianqin Yin

The word error rate (WER) of an automatic speech recognition (ASR) system increases when a mismatch occurs between the training and the testing conditions due to the noise, etc. In this case, the acoustic information can be less reliable.…

Computation and Language · Computer Science 2020-11-03 Dominique Fohr , Irina Illina

Deep Learning-based end-to-end Automatic Speech Recognition (ASR) has made significant strides but still struggles with performance on out-of-domain samples due to domain shifts in real-world scenarios. Test-Time Adaptation (TTA) methods…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-04 Guan-Ting Lin , Wei-Ping Huang , Hung-yi Lee

This work aims to investigate the use of deep neural network to detect commercial hobby drones in real-life environments by analyzing their sound data. The purpose of work is to contribute to a system for detecting drones used for malicious…

Sound · Computer Science 2017-01-23 Sungho Jeon , Jong-Woo Shin , Young-Jun Lee , Woong-Hee Kim , YoungHyoun Kwon , Hae-Yong Yang

The objective of deep learning methods based on encoder-decoder architectures for music source separation is to approximate either ideal time-frequency masks or spectral representations of the target music source(s). The spectral…

Multilingual ASR technology simplifies model training and deployment, but its accuracy is known to depend on the availability of language information at runtime. Since language identity is seldom known beforehand in real-world scenarios, it…

We propose a target driven adaptive (TDA) loss to enhance the performance of infrared small target detection (IRSTD). Prior works have used loss functions, such as binary cross-entropy loss and IoU loss, to train segmentation models for…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Yuho Shoji , Takahiro Toizumi , Atsushi Ito

The task of estimating the maximum number of concurrent speakers from single channel mixtures is important for various audio-based applications, such as blind source separation, speaker diarisation, audio surveillance or auditory scene…

Audio and Speech Processing · Electrical Eng. & Systems 2019-11-05 Fabian-Robert Stöter , Soumitro Chakrabarty , Bernd Edler , Emanuël A. P. Habets

Personal Voice Activity Detection (PVAD) is crucial for identifying target speaker segments in the mixture, yet its performance heavily depends on the quality of speaker embeddings. A key practical limitation is the short enrollment…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-21 Fuyuan Feng , Wenbin Zhang , Yu Gao , Longting Xu , Xiaofeng Mou , Yi Xu

Temporal difference learning (TD) is a foundational concept in reinforcement learning (RL), aimed at efficiently assessing a policy's value function. TD($\lambda$), a potent variant, incorporates a memory trace to distribute the prediction…

Machine Learning · Computer Science 2024-02-13 Jianfei Ma

End-to-end models for robust automatic speech recognition (ASR) have not been sufficiently well-explored in prior work. With end-to-end models, one could choose to preprocess the input speech using speech enhancement techniques and train…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-15 Archiki Prasad , Preethi Jyothi , Rajbabu Velmurugan

Deploying object detection on microcontrollers (MCUs) enables intelligent edge devices but current models cannot learn new object categories after deployment. Existing continual learning methods require storing raw images far exceeding MCU…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Bibin Wilson

Recent research in speaker verification has increasingly focused on achieving robust and reliable recognition under challenging channel conditions and noisy environments. Identifying speakers in radio communications is particularly…

Sound · Computer Science 2024-06-18 Wenhao Yang , Jianguo Wei , Wenhuan Lu , Lei Li , Xugang Lu