English
Related papers

Related papers: BFA: Real-time Multilingual Text-to-speech Forced …

200 papers

Pronunciation assessment is a major challenge in the computer-aided pronunciation training system, especially at the word (phoneme)-level. To obtain word (phoneme)-level scores, current methods usually rely on aligning components to obtain…

Computation and Language · Computer Science 2023-06-06 Yukang Liang , Kaitao Song , Shaoguang Mao , Huiqiang Jiang , Luna Qiu , Yuqing Yang , Dongsheng Li , Linli Xu , Lili Qiu

We propose an approach for learning critical articulators for phonemes through a machine learning approach. We formulate the learning with three models trained end to end. First, we use Acoustic to Articulatory Inversion (AAI) to predict…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-02 Jesuraj Bandekar , Sathvik Udupa , Prasanta Kumar Ghosh

Multi-modal emotion recognition in conversations is a challenging problem due to the complex and complementary interactions between different modalities. Audio and textual cues are particularly important for understanding emotions from a…

Sound · Computer Science 2025-04-02 Jiachen Luo , Huy Phan , Lin Wang , Joshua Reiss

Finding synthetic artifacts of spoofing data will help the anti-spoofing countermeasures (CMs) system discriminate between spoofed and real speech. The Conformer combines the best of convolutional neural network and the Transformer,…

Sound · Computer Science 2023-10-31 Yikang Wang , Hiromitsu Nishizaki , Ming Li

Modern large language models increasingly require long contexts for reasoning and multi-document tasks, but attention's quadratic complexity creates a severe computational bottleneck. We present Block-Sparse FlashAttention (BSFA), a drop-in…

Machine Learning · Computer Science 2025-12-09 Daniel Ohayon , Itay Lamprecht , Itay Hubara , Israel Cohen , Daniel Soudry , Noam Elata

Speech signal is constituted and contributed by various informative factors, such as linguistic content and speaker characteristic. There have been notable recent studies attempting to factorize speech signal into these individual factors…

Audio and Speech Processing · Electrical Eng. & Systems 2019-11-06 Zhiyuan Peng , Siyuan Feng , Tan Lee

The objective of this work is to develop a speaker recognition model to be used in diverse scenarios. We hypothesise that two components should be adequately configured to build such a model. First, adequate architecture would be required.…

Despite the impressive performance achieved by pre-trained language-and-vision models in downstream tasks, it remains an open question whether this reflects a proper understanding of image-text interaction. In this work, we explore to what…

Computation and Language · Computer Science 2024-01-22 Xinyi Chen , Raquel Fernández , Sandro Pezzelle

In noisy conditions, knowing speech contents facilitates listeners to more effectively suppress background noise components and to retrieve pure speech signals. Previous studies have also confirmed the benefits of incorporating phonetic…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-19 Yen-Ju Lu , Chien-Feng Liao , Xugang Lu , Jeih-weih Hung , Yu Tsao

Brain-computer interfaces (BCI) offer numerous human-centered application possibilities, particularly affecting people with neurological disorders. Text or speech decoding from brain activities is a relevant domain that could augment the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-10 Jihwan Lee , Tiantian Feng , Aditya Kommineni , Sudarsana Reddy Kadiri , Shrikanth Narayanan

The study of speech disorders can benefit greatly from time-aligned data. However, audio-text mismatches in disfluent speech cause rapid performance degradation for modern speech aligners, hindering the use of automatic approaches. In this…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-05 Theodoros Kouzelis , Georgios Paraskevopoulos , Athanasios Katsamanis , Vassilis Katsouros

Speech brain-computer interfaces (BCIs) aim to restore communication for people with paralysis by translating neural activity into text. Most systems use cascaded frameworks that decode phonemes before assembling sentences with an n-gram…

The recently proposed Conformer architecture which combines convolution with attention to capture both local and global dependencies has become the \textit{de facto} backbone model for Automatic Speech Recognition~(ASR). Inherited from the…

Recently, remote sensing image captioning has gained significant attention in the remote sensing community. Due to the significant differences in spatial resolution of remote sensing images, existing methods in this field have predominantly…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Cong Yang , Zuchao Li , Lefei Zhang

Binaural speech enhancement faces a severe trade-off challenge, where state-of-the-art performance is achieved by computationally intensive architectures, while lightweight solutions often come at the cost of significant performance…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-26 Xikun Lu , Yujian Ma , Xianquan Jiang , Xuelong Wang , Jinqiu Sang

Speech separation always faces the challenge of handling prolonged time sequences. Past methods try to reduce sequence lengths and use the Transformer to capture global information. However, due to the quadratic time complexity of the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-28 Haoxu Wang , Yiheng Jiang , Gang Qiao , Pengteng Shi , Biao Tian

Given the significant potential of large language models (LLMs) in sequence modeling, emerging studies have begun applying them to time-series forecasting. Despite notable progress, existing methods still face two critical challenges: 1)…

Artificial Intelligence · Computer Science 2025-01-09 Pengfei Wang , Huanran Zheng , Qi'ao Xu , Silong Dai , Yiqiao Wang , Wenjing Yue , Wei Zhu , Tianwen Qian , Xiaoling Wang

A Pascal challenge entitled monaural multi-talker speech recognition was developed, targeting the problem of robust automatic speech recognition against speech like noises which significantly degrades the performance of automatic speech…

Computation and Language · Computer Science 2016-10-06 Mahdi Khademian , Mohammad Mehdi Homayounpour

Continuous integrate-and-fire (CIF) based models, which use a soft and monotonic alignment mechanism, have been well applied in non-autoregressive (NAR) speech recognition with competitive performance compared with other NAR methods.…

The learning-from-observation (LfO) framework aims to map human demonstrations to a robot to reduce programming effort. To this end, an LfO system encodes a human demonstration into a series of execution units for a robot, which are…

Robotics · Computer Science 2021-03-25 Naoki Wake , Iori Yanokura , Kazuhiro Sasabuchi , Katsushi Ikeuchi