English
Related papers

Related papers: Multilingual Phonological Feature Recognition with…

200 papers

Emotion recognition in speech is a challenging multimodal task that requires understanding both verbal content and vocal nuances. This paper introduces a novel approach to emotion detection using Large Language Models (LLMs), which have…

Computation and Language · Computer Science 2024-12-24 Zehui Wu , Ziwei Gong , Lin Ai , Pengyuan Shi , Kaan Donbekci , Julia Hirschberg

Speech foundation models trained with self-supervised learning produce generic speech representations that support a wide range of speech processing tasks. When further adapted with supervised learning, these models can achieve strong…

Computation and Language · Computer Science 2026-03-10 Maryem Bouziane , Salima Mdhaffar , Yannick Estève

Transformers have achieved state-of-the-art performance in morphological inflection tasks, yet their ability to generalize across languages and morphological rules remains limited. One possible explanation for this behavior can be the…

Computation and Language · Computer Science 2025-06-03 Gal Astrach , Yuval Pinter

What do deep neural speech models know about phonology? Existing work has examined the encoding of individual linguistic units such as phonemes in these models. Here we investigate interactions between units. Inspired by classic experiments…

Computation and Language · Computer Science 2024-07-04 Marianne de Heer Kloots , Willem Zuidema

Recently end-to-end neural audio/speech coding has shown its great potential to outperform traditional signal analysis based audio codecs. This is mostly achieved by following the VQ-VAE paradigm where blind features are learned,…

Sound · Computer Science 2023-02-28 Xue Jiang , Xiulian Peng , Yuan Zhang , Yan Lu

This paper describes an online algorithm for enhancing monaural noisy speech. Firstly, a novel phase-corrected low-delay gammatone filterbank is derived for signal subband decomposition and resynthesis; the subband signals are then analyzed…

Sound · Computer Science 2015-07-09 Zhangli Chen , Volker Hohmann

Self-supervised pre-trained features have consistently delivered state-of-art results in the field of natural language processing (NLP); however, their merits in the field of speech emotion recognition (SER) still need further…

Sound · Computer Science 2022-02-09 Edmilson Morais , Ron Hoory , Weizhong Zhu , Itai Gat , Matheus Damasceno , Hagai Aronowitz

Code-switching---the intra-utterance use of multiple languages---is prevalent across the world. Within text-to-speech (TTS), multilingual models have been found to enable code-switching. By modifying the linguistic input to…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-17 Marlene Staib , Tian Huey Teh , Alexandra Torresquintero , Devang S Ram Mohan , Lorenzo Foglianti , Raphael Lenain , Jiameng Gao

Fast Fourier convolution (FFC) is the recently proposed neural operator showing promising performance in several computer vision problems. The FFC operator allows employing large receptive field operations within early layers of the neural…

Sound · Computer Science 2022-04-08 Ivan Shchekotov , Pavel Andreev , Oleg Ivanov , Aibek Alanov , Dmitry Vetrov

Today, there have been many achievements in learning the association between voice and face. However, most previous work models rely on cosine similarity or L2 distance to evaluate the likeness of voices and faces following contrastive…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Chong Peng , Liqiang He , Dan Su

AI-synthesized voice technology has the potential to create realistic human voices for beneficial applications, but it can also be misused for malicious purposes. While existing AI-synthesized voice detection models excel in intra-domain…

Sound · Computer Science 2024-12-31 Hainan Ren , Li Lin , Chun-Hao Liu , Xin Wang , Shu Hu

While automatic subjective speech quality assessment has witnessed much progress, an open question is whether an automatic quality assessment at frame resolution is possible. This would be highly desirable, as it adds explainability to the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-10 Michael Kuhlmann , Fritz Seebauer , Petra Wagner , Reinhold Haeb-Umbach

Recent studies have shown that frame-level deep speaker features can be derived from a deep neural network with the training target set to discriminate speakers by a short speech segment. By pooling the frame-level features, utterance-level…

Audio and Speech Processing · Electrical Eng. & Systems 2018-11-09 Lantian Li , Zhiyuan Tang , Ying Shi , Dong Wang

Data2vec is a self-supervised learning (SSL) approach that employs a teacher-student architecture for contextual representation learning via masked prediction, demonstrating remarkable performance in monolingual ASR. Previous studies have…

Sound · Computer Science 2025-01-24 Qijie Shao , Linhao Dong , Kun Wei , Sining Sun , Lei Xie

Non-intrusive assessment of speech quality and intelligibility is essential when clean reference signals are unavailable. In this work, we propose a multimodal framework that integrates audio features and visual cues to predict PESQ and…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-12 Shafique Ahmed , Ryandhimas E. Zezario , Nasir Saleem , Amir Hussain , Hsin-Min Wang , Yu Tsao

Multilingual speech recognition with supervised learning has achieved great results as reflected in recent research. With the development of pretraining methods on audio and text data, it is imperative to transfer the knowledge from…

Computation and Language · Computer Science 2022-05-26 Ngoc-Quan Pham , Alex Waibel , Jan Niehues

The lack of labeled second language (L2) speech data is a major challenge in designing mispronunciation detection models. We introduce SpeechBlender - a fine-grained data augmentation pipeline for generating mispronunciation errors to…

Sound · Computer Science 2023-07-13 Yassine El Kheir , Shammur Absar Chowdhury , Ahmed Ali , Hamdy Mubarak , Shazia Afzal

Multilingual models can improve language processing, particularly for low resource situations, by sharing parameters across languages. Multilingual acoustic models, however, generally ignore the difference between phonemes (sounds that can…

Multimodal language models attempt to incorporate non-linguistic features for the language modeling task. In this work, we extend a standard recurrent neural network (RNN) language model with features derived from videos. We train our…

Computation and Language · Computer Science 2019-03-08 Antonios Anastasopoulos , Shankar Kumar , Hank Liao

This paper presents methods of making using of text supervision to improve the performance of sequence-to-sequence (seq2seq) voice conversion. Compared with conventional frame-to-frame voice conversion approaches, the seq2seq acoustic…

Sound · Computer Science 2020-01-14 Jing-Xuan Zhang , Zhen-Hua Ling , Yuan Jiang , Li-Juan Liu , Chen Liang , Li-Rong Dai