English
Related papers

Related papers: Exploring Wav2vec 2.0 fine-tuning for improved spe…

200 papers

Speech emotion recognition (SER) has been a challenging problem in spoken language processing research, because it is unclear how human emotions are connected to various components of sounds such as pitch, loudness, and energy. This paper…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-03 Tai Vu

State-of-the-art automatic speech recognition (ASR) systems perform well on healthy speech. However, the performance on impaired speech still remains an issue. The current study explores the usefulness of using Wav2Vec self-supervised…

Computation and Language · Computer Science 2022-04-05 Abner Hernandez , Paula Andrea Pérez-Toro , Elmar Nöth , Juan Rafael Orozco-Arroyave , Andreas Maier , Seung Hee Yang

In this paper, we use several techniques with conventional vocal feature extraction (MFCC, STFT), along with deep-learning approaches such as CNN, and also context-level analysis, by providing the textual data, and combining different…

Audio and Speech Processing · Electrical Eng. & Systems 2019-05-22 Andrew Huang , Puwei Bao

In this paper, we propose a single multi-task learning framework to perform End-to-End (E2E) speech recognition (ASR) and accent recognition (AR) simultaneously. The proposed framework is not only more compact but can also yield comparable…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-16 Jicheng Zhang , Yizhou Peng , Pham Van Tung , Haihua Xu , Hao Huang , Eng Siong Chng

Automatic Speech Recognition (ASR) systems have attained unprecedented performance with large speech models pre-trained based on self-supervised speech representation learning. However, these pre-trained speech models suffer from…

Computation and Language · Computer Science 2023-05-29 Eunseop Yoon , Hee Suk Yoon , John Harvill , Mark Hasegawa-Johnson , Chang D. Yoo

We propose the use of parameter-efficient fine-tuning (PEFT) of foundation models for cleft lip and palate (CLP) detection and severity classification. In CLP, nasalization increases with severity due to the abnormal passage between the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-22 Susmita Bhattacharjee , Jagabandhu Mishra , H. S. Shekhawat , S. R. Mahadeva Prasanna

In this paper, we propose to use pre-trained features from end-to-end ASR models to solve speech sentiment analysis as a down-stream task. We show that end-to-end ASR features, which integrate both acoustic and text information from speech,…

Computation and Language · Computer Science 2020-03-06 Zhiyun Lu , Liangliang Cao , Yu Zhang , Chung-Cheng Chiu , James Fan

New-age conversational agent systems perform both speech emotion recognition (SER) and automatic speech recognition (ASR) using two separate and often independent approaches for real-world application in noisy environments. In this paper,…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-29 Lokesh Bansal , S. Pavankumar Dubagunta , Malolan Chetlur , Pushpak Jagtap , Aravind Ganapathiraju

This paper is a study of performance-efficiency trade-offs in pre-trained models for automatic speech recognition (ASR). We focus on wav2vec 2.0, and formalize several architecture designs that influence both the model performance and its…

Computation and Language · Computer Science 2021-09-15 Felix Wu , Kwangyoun Kim , Jing Pan , Kyu Han , Kilian Q. Weinberger , Yoav Artzi

In recent years, speech emotion recognition (SER) has been used in wide ranging applications, from healthcare to the commercial sector. In addition to signal processing approaches, methods for SER now also use deep learning techniques which…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-29 Sneha Das , Nicole Nadine Lønfeldt , Anne Katrine Pagsberg , Line H. Clemmensen

In this work, we present our submission to the Speech Accessibility Project challenge for dysarthric speech recognition. We integrate parameter-efficient fine-tuning with latent audio representations to improve an encoder-decoder ASR…

Recently, we made available WeNet, a production-oriented end-to-end speech recognition toolkit, which introduces a unified two-pass (U2) framework and a built-in runtime to address the streaming and non-streaming decoding modes in a single…

Sound · Computer Science 2022-07-06 Binbin Zhang , Di Wu , Zhendong Peng , Xingchen Song , Zhuoyuan Yao , Hang Lv , Lei Xie , Chao Yang , Fuping Pan , Jianwei Niu

Building inclusive speech recognition systems is a crucial step towards developing technologies that speakers of all language varieties can use. Therefore, ASR systems must work for everybody independently of the way they speak. To…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-18 Alëna Aksënova , Zhehuai Chen , Chung-Cheng Chiu , Daan van Esch , Pavel Golik , Wei Han , Levi King , Bhuvana Ramabhadran , Andrew Rosenberg , Suzan Schwartz , Gary Wang

Parameter-efficient fine-tuning (PEFT) has emerged as a scalable solution for adapting large foundation models. While low-rank adaptation (LoRA) is widely used in speech applications, its state-of-the-art variants, e.g., VeRA, DoRA, PiSSA,…

Computation and Language · Computer Science 2025-09-04 Pu Wang , Shinji Watanabe , Hugo Van hamme

To address the performance gap of English ASR models on L2 English speakers, we evaluate fine-tuning of pretrained wav2vec 2.0 models (Baevski et al., 2020; Xu et al., 2021) on L2-ARCTIC, a non-native English speech corpus (Zhao et al.,…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-18 Toshiko Shibano , Xinyi Zhang , Mia Taige Li , Haejin Cho , Peter Sullivan , Muhammad Abdul-Mageed

End-to-end automatic speech recognition systems represent the state of the art, but they rely on thousands of hours of manually annotated speech for training, as well as heavyweight computation for inference. Of course, this impedes…

Computation and Language · Computer Science 2022-11-22 Raphael Tang , Karun Kumar , Gefei Yang , Akshat Pandey , Yajie Mao , Vladislav Belyaev , Madhuri Emmadi , Craig Murray , Ferhan Ture , Jimmy Lin

Automatic Speech Recognition (ASR) systems often struggle with transcribing child speech due to the lack of large child speech datasets required to accurately train child-friendly ASR models. However, there are huge amounts of annotated…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-26 Rishabh Jain , Andrei Barcovschi , Mariam Yiwere , Peter Corcoran , Horia Cucu

Recently, self-supervised learning (SSL) from unlabelled speech data has gained increased attention in the automatic speech recognition (ASR) community. Typical SSL methods include autoregressive predictive coding (APC), Wav2vec2.0, and…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-02 Ruchao Fan , Yunzheng Zhu , Jinhan Wang , Abeer Alwan

Despite the increasing research interest in end-to-end learning systems for speech emotion recognition, conventional systems either suffer from the overfitting due in part to the limited training data, or do not explicitly consider the…

Computation and Language · Computer Science 2019-04-01 Zixing Zhang , Bingwen Wu , Bjoern Schuller

In this paper, we propose to improve emotion recognition by combining acoustic information and conversation transcripts. On the one hand, an LSTM network was used to detect emotion from acoustic features like f0, shimmer, jitter, MFCC, etc.…

Audio and Speech Processing · Electrical Eng. & Systems 2019-11-04 Jaejin Cho , Raghavendra Pappagari , Purva Kulkarni , Jesus Villalba , Yishay Carmiel , Najim Dehak