English
Related papers

Related papers: Spatio-Temporal Representation Learning Enhanced S…

200 papers

We consider the generative modeling of speech over multiple minutes, a requirement for long-form multimedia generation and audio-native voice assistants. However, textless spoken language models struggle to generate plausible speech past…

Computation and Language · Computer Science 2025-07-11 Se Jin Park , Julian Salazar , Aren Jansen , Keisuke Kinoshita , Yong Man Ro , RJ Skerry-Ryan

In this study, we propose the global context guided channel and time-frequency transformations to model the long-range, non-local time-frequency dependencies and channel variances in speaker representations. We use the global context…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-10 Wei Xia , John H. L. Hansen

Many of the existing TTS systems cannot accurately synthesize text containing a variety of numerical formats, resulting in reduced intelligibility of the synthesized speech. This research aims to develop a numerical format classifier that…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-03 Yaser Darwesh , Lit Wei Wern , Mumtaz Begum Mustafa

Data assimilation in models representing spatio-temporal phenomena poses a challenge, particularly if the spatial histogram of the variable appears with multiple modes. The traditional Kalman model is based on a Gaussian initial…

Methodology · Statistics 2020-06-26 Maxime Conjard , Henning Omre

3D Gaussian Splatting (3D-GS) has emerged as an efficient 3D representation and a promising foundation for semantic tasks like segmentation. However, existing 3D-GS-based segmentation methods typically rely on high-dimensional category…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 An Yang , Chenyu Liu , Jun Du , Jianqing Gao , Jia Pan , Jinshui Hu , Baocai Yin , Bing Yin , Cong Liu

Heart sound signals, phonocardiography (PCG) signals, allow for the automatic diagnosis of potential cardiovascular pathology. Such classification task can be tackled using the bidirectional long short-term memory (biLSTM) network, trained…

Sound · Computer Science 2026-04-16 Mahmoud Fakhry , Abeer FathAllah Brery

In this work, we focus on source tracing of synthetic speech generation systems (STSGS). Each source embeds distinctive paralinguistic features--such as pitch, tone, rhythm, and intonation--into their synthesized speech, reflecting the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-03 Girish , Mohd Mujtaba Akhtar , Orchid Chetia Phukan , Drishti Singh , Swarup Ranjan Behera , Pailla Balakrishna Reddy , Arun Balaji Buduru , Rajesh Sharma

Lip-reading aims to recognize speech content from videos via visual analysis of speakers' lip movements. This is a challenging task due to the existence of homophemes-words which involve identical or highly similar lip movements, as well as…

Computer Vision and Pattern Recognition · Computer Science 2019-09-04 Chenhao Wang

Cell detection is the task of detecting the approximate positions of cell centroids from microscopy images. Recently, convolutional neural network-based approaches have achieved promising performance. However, these methods require a…

Computer Vision and Pattern Recognition · Computer Science 2021-07-20 Kazuya Nishimura , Hyeonwoo Cho , Ryoma Bise

Existing visual object tracking usually learns a bounding-box based template to match the targets across frames, which cannot accurately learn a pixel-wise representation, thereby being limited in handling severe appearance variations. To…

Computer Vision and Pattern Recognition · Computer Science 2021-04-07 Fei Xie , Wankou Yang , Bo Liu , Kaihua Zhang , Wanli Xue , Wangmeng Zuo

We present prompt distribution learning for effectively adapting a pre-trained vision-language model to address downstream recognition tasks. Our method not only learns low-bias prompts from a few samples but also captures the distribution…

Computer Vision and Pattern Recognition · Computer Science 2022-05-09 Yuning Lu , Jianzhuang Liu , Yonggang Zhang , Yajing Liu , Xinmei Tian

Long Short Term Memory Connectionist Temporal Classification (LSTM-CTC) based end-to-end models are widely used in speech recognition due to its simplicity in training and efficiency in decoding. In conventional LSTM-CTC based models, a…

Computation and Language · Computer Science 2019-03-14 Yangyang Shi , Mei-Yuh Hwang , Xin Lei

In this work, a Bayesian approach to speaker normalization is proposed to compensate for the degradation in performance of a speaker independent speech recognition system. The speaker normalization method proposed herein uses the technique…

Sound · Computer Science 2016-10-20 Dhananjay Ram , Debasis Kundu , Rajesh M. Hegde

In mobile communication scenarios, the acquired channel state information (CSI) rapidly becomes outdated due to fast-changing channels. Opportunistic transmitter selection based on current CSI for secrecy improvement may be outdated during…

Signal Processing · Electrical Eng. & Systems 2024-05-02 Shashi Bhushan Kotwal , Chinmoy Kundu , Sudhakar Modem , Holger Claussen , Lester Ho

Current synthetic speech detection (SSD) methods perform well on certain datasets but still face issues of robustness and interpretability. A possible reason is that these methods do not analyze the deficiencies of synthetic speech. In this…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-02 Yuxiang Zhang , Zhuo Li , Jingze Lu , Wenchao Wang , Pengyuan Zhang

We compare using a PHOIBLE-based phone mapping method and using phonological features input in transfer learning for TTS in low-resource languages. We use diverse source languages (English, Finnish, Hindi, Japanese, and Russian) and target…

Computation and Language · Computer Science 2023-06-22 Phat Do , Matt Coler , Jelske Dijkstra , Esther Klabbers

Guided Source Separation (GSS) is a popular front-end for distant automatic speech recognition (ASR) systems using spatially distributed microphones. When considering spatially distributed microphones, the choice of reference microphone may…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-03 Anselm Lohmann , Tomohiro Nakatani , Rintaro Ikeshita , Marc Delcroix , Shoko Araki , Simon Doclo

A comparative study of the application of Gaussian Mixture Model (GMM) and Radial Basis Function (RBF) in biometric recognition of voice has been carried out and presented. The application of machine learning techniques to biometric…

Machine Learning · Computer Science 2012-11-13 Fatai Adesina Anifowose

Long Short-Term Memory (LSTM) is the primary recurrent neural networks architecture for acoustic modeling in automatic speech recognition systems. Residual learning is an efficient method to help neural networks converge easier and faster.…

Computation and Language · Computer Science 2017-08-21 Lu Huang , Jiasong Sun , Ji Xu , Yi Yang

Connectionist Temporal Classification has recently attracted a lot of interest as it offers an elegant approach to building acoustic models (AMs) for speech recognition. The CTC loss function maps an input sequence of observable feature…

Computation and Language · Computer Science 2017-08-16 Thomas Zenkel , Ramon Sanabria , Florian Metze , Jan Niehues , Matthias Sperber , Sebastian Stüker , Alex Waibel