English
Related papers

Related papers: CIF: Continuous Integrate-and-Fire for End-to-End …

200 papers

Finding synthetic artifacts of spoofing data will help the anti-spoofing countermeasures (CMs) system discriminate between spoofed and real speech. The Conformer combines the best of convolutional neural network and the Transformer,…

Sound · Computer Science 2023-10-31 Yikang Wang , Hiromitsu Nishizaki , Ming Li

Cued Speech (CS) is an augmented lip reading complemented by hand coding, and it is very helpful to the deaf people. Automatic CS recognition can help communications between the deaf people and others. Due to the asynchronous nature of lips…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-25 Li Liu , Gang Feng , Denis Beautemps , Xiao-Ping Zhang

The performance of voice-controlled systems is usually influenced by accented speech. To make these systems more robust, the frontend accent recognition (AR) technologies have received increased attention in recent years. As accent is a…

Audio and Speech Processing · Electrical Eng. & Systems 2021-05-06 Zhan Zhang , Xi Chen , Yuehai Wang , Jianyi Yang

While prior research has proposed a plethora of methods that build neural classifiers robust against adversarial robustness, practitioners are still reluctant to adopt them due to their unacceptably severe clean accuracy penalties. This…

Machine Learning · Computer Science 2024-07-23 Yatong Bai , Brendon G. Anderson , Aerin Kim , Somayeh Sojoudi

LLM-based code interpreter agents are increasingly deployed in critical workflows, yet their robustness against risks introduced by their code execution capabilities remains underexplored. Existing benchmarks are limited to static datasets…

Cryptography and Security · Computer Science 2026-02-24 Lei Ba , Qinbin Li , Songze Li

Speech is one of the most effective ways of communication among humans. Even though audio is the most common way of transmitting speech, very important information can be found in other modalities, such as vision. Vision is particularly…

Computation and Language · Computer Science 2016-11-22 Ramon Sanabria , Florian Metze , Fernando De La Torre

Class-Incremental Learning (CIL) aims to train a reliable model with the streaming data, which emerges unknown classes sequentially. Different from traditional closed set learning, CIL has two main challenges: 1) Novel class detection. The…

Machine Learning · Computer Science 2020-09-01 Yang Yang , Zhen-Qiang Sun , HengShu Zhu , Yanjie Fu , Hui Xiong , Jian Yang

We propose Shallow Flow Matching (SFM), a novel mechanism that enhances flow matching (FM)-based text-to-speech (TTS) models within a coarse-to-fine generation paradigm. Unlike conventional FM modules, which use the coarse representations…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-24 Dong Yang , Yiyi Cai , Yuki Saito , Lixu Wang , Hiroshi Saruwatari

In real-time speech recognition applications, the latency is an important issue. We have developed a character-level incremental speech recognition (ISR) system that responds quickly even during the speech, where the hypotheses are…

Computation and Language · Computer Science 2016-06-29 Kyuyeon Hwang , Wonyong Sung

In end-to-end automatic speech recognition (ASR), a model is expected to implicitly learn representations suitable for recognizing a word-level sequence. However, the huge abstraction gap between input acoustic signals and output linguistic…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-09 Yosuke Higuchi , Keita Karube , Tetsuji Ogawa , Tetsunori Kobayashi

Self-supervised learning representation (SSLR) has demonstrated its significant effectiveness in automatic speech recognition (ASR), mainly with clean speech. Recent work pointed out the strength of integrating SSLR with single-channel…

Sound · Computer Science 2022-10-20 Yoshiki Masuyama , Xuankai Chang , Samuele Cornell , Shinji Watanabe , Nobutaka Ono

Scores from traditional confidence classifiers (CCs) in automatic speech recognition (ASR) systems lack universal interpretation and vary with updates to the underlying confidence or acoustic models (AMs). In this work, we build…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-02 Amber Afshan , Kshitiz Kumar , Jian Wu

Automatic speech recognition (ASR) on multi-talker recordings is challenging. Current methods using 3D spatial data from multi-channel audio and visual cues focus mainly on direct waves from the target speaker, overlooking reflection wave…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-13 Yiwen Shao , Shi-Xiong Zhang , Dong Yu

How to leverage dynamic contextual information in end-to-end speech recognition has remained an active research area. Previous solutions to this problem were either designed for specialized use cases that did not generalize well to…

Although deep learning and end-to-end models have been widely used and shown superiority in automatic speech recognition (ASR) and text-to-speech (TTS) synthesis, state-of-the-art forced alignment (FA) models are still based on hidden…

Sound · Computer Science 2022-04-01 Jingbei Li , Yi Meng , Zhiyong Wu , Helen Meng , Qiao Tian , Yuping Wang , Yuxuan Wang

We propose "Generative Fusion Decoding" (GFD), a novel shallow fusion framework designed to integrate large language models (LLMs) into cross-modal text recognition systems for automatic speech recognition (ASR) and optical character…

Computation and Language · Computer Science 2025-06-12 Chan-Jan Hsu , Yi-Chang Chen , Feng-Ting Liao , Pei-Chen Ho , Yu-Hsiang Wang , Po-Chun Hsu , Da-shan Shiu

Mobile-centric AI applications have high requirements for resource-efficiency of model inference. Input filtering is a promising approach to eliminate the redundancy so as to reduce the cost of inference. Previous efforts have tailored…

Artificial Intelligence · Computer Science 2023-06-08 Mu Yuan , Lan Zhang , Fengxiang He , Xueting Tong , Miao-Hui Song , Zhengyuan Xu , Xiang-Yang Li

Existing digital sensors capture images at fixed spatial and spectral resolutions (e.g., RGB, multispectral, and hyperspectral images), and each combination requires bespoke machine learning models. Neural Implicit Functions partially…

Computer Vision and Pattern Recognition · Computer Science 2023-10-03 Gengchen Mai , Ni Lao , Weiwei Sun , Yuchi Ma , Jiaming Song , Chenlin Meng , Hongxu Ma , Jinmeng Rao , Ziyuan Li , Stefano Ermon

This paper proposes a novel label-synchronous speech-to-text alignment technique for automatic speech recognition (ASR). The speech-to-text alignment is a problem of splitting long audio recordings with un-aligned transcripts into…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-22 Yusuke Kida , Tatsuya Komatsu , Masahito Togami

Despite the remarkable progress of deep learning in stereo matching, there exists a gap in accuracy between real-time models and slower state-of-the-art models which are suitable for practical applications. This paper presents an iterative…

Computer Vision and Pattern Recognition · Computer Science 2021-10-26 Kumail Raza , René Schuster , Didier Stricker