English
Related papers

Related papers: Spatio-Temporal Representation Learning Enhanced S…

200 papers

Open-vocabulary panoptic reconstruction is crucial for advanced robotics and simulation. However, existing 3D reconstruction methods, such as NeRF or Gaussian Splatting variants, often struggle to achieve the real-time inference frequency…

Robotics · Computer Science 2026-04-14 Xuan Yu , Yuxuan Xie , Shichao Zhai , Shuhao Ye , Rong Xiong , Yue Wang

Content and style representations have been widely studied in the field of style transfer. In this paper, we propose a new loss function using speaker content representation for audio source separation, and we call it speaker representation…

Sound · Computer Science 2020-02-28 Seongkyu Mun , Soyeon Choe , Jaesung Huh , Joon Son Chung

Speaking Style Recognition (SSR) identifies a speaker's speaking style characteristics from speech. Existing style recognition approaches primarily rely on linguistic information, with limited integration of acoustic information, which…

Sound · Computer Science 2025-10-15 Guojian Li , Qijie Shao , Zhixian Zhao , Shuiyuan Wang , Zhonghua Fu , Lei Xie

Disentangled representation learning offers useful properties such as dimension reduction and interpretability, which are essential to modern deep learning approaches. Although deep learning techniques have been widely applied to…

Machine Learning · Computer Science 2022-04-11 Sichen Zhao , Wei Shao , Jeffrey Chan , Flora D. Salim

In this work, we explore a multimodal semi-supervised learning approach for punctuation prediction by learning representations from large amounts of unlabelled audio and text data. Conventional approaches in speech processing typically use…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-04 Monica Sunkara , Srikanth Ronanki , Dhanush Bekal , Sravan Bodapati , Katrin Kirchhoff

Self-attention (SA) based models have recently achieved significant performance improvements in hybrid and end-to-end automatic speech recognition (ASR) systems owing to their flexible context modeling capability. However, it is also known…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-19 Yosuke Kashiwagi , Emiru Tsunoo , Shinji Watanabe

In multiple-input multiple-output (MIMO) systems, it is crucial of utilizing the available channel state information (CSI) at the transmitter for precoding to improve the performance of frequency division duplex (FDD) networks. One of the…

Signal Processing · Electrical Eng. & Systems 2022-04-28 Xiangyi Li , Huaming Wu

In NLP, text language models based on words or subwords are known to outperform their character-based counterparts. Yet, in the speech community, the standard input of spoken LMs are 20ms or 40ms-long discrete units (shorter than a…

Computation and Language · Computer Science 2023-10-10 Robin Algayres , Yossi Adi , Tu Anh Nguyen , Jade Copet , Gabriel Synnaeve , Benoit Sagot , Emmanuel Dupoux

Cellular traffic prediction is of great importance for operators to manage network resources and make decisions. Traffic is highly dynamic and influenced by many exogenous factors, which would lead to the degradation of traffic prediction…

Machine Learning · Computer Science 2025-06-23 Hui Ma , Kai Yang , Man-On Pun

Reliable multimodal sensor fusion algorithms require accurate spatiotemporal calibration. Recently, targetless calibration techniques based on implicit neural representations have proven to provide precise and robust results. Nevertheless,…

Computer Vision and Pattern Recognition · Computer Science 2024-09-19 Quentin Herau , Moussab Bennehar , Arthur Moreau , Nathan Piasco , Luis Roldao , Dzmitry Tsishkou , Cyrille Migniot , Pascal Vasseur , Cédric Demonceaux

The use of spatial information with multiple microphones can improve far-field automatic speech recognition (ASR) accuracy. However, conventional microphone array techniques degrade speech enhancement performance when there is an array…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-23 Kenichi Kumatani , Minhua Wu , Shiva Sundaram , Nikko Strom , Bjorn Hoffmeister

In this work, we propose a novel straightforward method for medical volume and sequence segmentation with limited annotations. To avert laborious annotating, the recent success of self-supervised learning(SSL) motivates the pre-training on…

Computer Vision and Pattern Recognition · Computer Science 2023-10-04 Zejian Chen , Wei Zhuo , Tianfu Wang , Wufeng Xue , Dong Ni

Acoustic models based on long short-term memory recurrent neural networks (LSTM-RNNs) were applied to statistical parametric speech synthesis (SPSS) and showed significant improvements in naturalness and latency over those based on hidden…

Generating natural speech with a diverse and smooth prosody pattern is a challenging task. Although random sampling with phone-level prosody distribution has been investigated to generate different prosody patterns, the diversity of the…

Sound · Computer Science 2024-10-30 Chenpeng Du , Kai Yu

Previous studies demonstrated that a dynamic phone-informed compression of the input audio is beneficial for speech translation (ST). However, they required a dedicated model for phone recognition and did not test this solution for direct…

Computation and Language · Computer Science 2021-10-15 Marco Gaido , Mauro Cettolo , Matteo Negri , Marco Turchi

Continuous sign language recognition (CSLR) requires precise spatio-temporal modeling to accurately recognize sequences of gestures in videos. Existing frameworks often rely on CNN-based spatial backbones combined with temporal convolution…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Ahmed Abul Hasanaath , Hamzah Luqman

This paper proposes a new approach to duration modelling for statistical parametric speech synthesis in which a recurrent statistical model is trained to output a phone transition probability at each timestep (acoustic frame). Unlike…

Computation and Language · Computer Science 2020-07-28 Srikanth Ronanki , Oliver Watts , Simon King , Gustav Eje Henter

There has been increased interest in missing sensor data imputation, which is ubiquitous in the field of structural health monitoring (SHM) due to discontinuous sensing caused by sensor malfunction. To address this fundamental issue, this…

Machine Learning · Computer Science 2020-07-20 Pu Ren , Xinyu Chen , Lijun Sun , Hao Sun

Connectionist temporal classification (CTC) provides an end-to-end acoustic model (AM) training strategy. CTC learns accurate AMs without time-aligned phonetic transcription, but sometimes fails to converge, especially in…

Audio and Speech Processing · Electrical Eng. & Systems 2019-02-28 Di He , Xuesong Yang , Boon Pang Lim , Yi Liang , Mark Hasegawa-Johnson , Deming Chen

Most of the deep learning based speech enhancement (SE) methods rely on estimating the magnitude spectrum of the clean speech signal from the observed noisy speech signal, either by magnitude spectral masking or regression. These methods…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-28 Raktim Gautam Goswami , Sivaganesh Andhavarapu , K Sri Rama Murty
‹ Prev 1 4 5 6 7 8 10 Next ›