中文
相关论文

相关论文: Extracting speaker and emotion information from se…

200 篇论文

Estimating dimensional emotions, such as activation, valence and dominance, from acoustic speech signals has been widely explored over the past few years. While accurate estimation of activation and dominance from speech seem to be…

音频与语音处理 · 电气工程与系统科学 2022-07-08 Vikramjit Mitra , Hsiang-Yun Sherry Chien , Vasudha Kowtha , Joseph Yitan Cheng , Erdrin Azemi

Natural language processing (NLP) models trained on people-generated data can be unreliable because, without any constraints, they can learn from spurious correlations that are not relevant to the task. We hypothesize that enriching models…

计算与语言 · 计算机科学 2022-03-18 Alissa Ostapenko , Shuly Wintner , Melinda Fricke , Yulia Tsvetkov

There are individual differences in expressive behaviors driven by cultural norms and personality. This between-person variation can result in reduced emotion recognition performance. Therefore, personalization is an important step in…

音频与语音处理 · 电气工程与系统科学 2023-09-06 Minh Tran , Yufeng Yin , Mohammad Soleymani

Speech signals are inherently complex as they encompass both global acoustic characteristics and local semantic information. However, in the task of target speech extraction, certain elements of global and local semantic information in the…

声音 · 计算机科学 2024-08-27 Zhaoxi Mu , Xinyu Yang , Sining Sun , Qing Yang

Speech Emotion Recognition (SER) plays a pivotal role in enhancing human-computer interaction by enabling a deeper understanding of emotional states across a wide range of applications, contributing to more empathetic and effective…

音频与语音处理 · 电气工程与系统科学 2023-09-25 Amirali Soltani Tehrani , Niloufar Faridani , Ramin Toosi

Speech evaluation measures a learners oral proficiency using automatic models. Corpora for training such models often pose sparsity challenges given that there often is limited scored data from teachers, in addition to the score…

人工智能 · 计算机科学 2024-09-24 Huayun Zhang , Jeremy H. M. Wong , Geyu Lin , Nancy F. Chen

Traditional approaches to automatic emotion recognition are relying on the application of handcrafted features. More recently however the advent of deep learning enabled algorithms to learn meaningful representations of input data…

音频与语音处理 · 电气工程与系统科学 2020-10-01 Dominik Schiller , Silvan Mertes , Elisabeth André

Pooling is needed to aggregate frame-level features into utterance-level representations for speaker modeling. Given the success of statistics-based pooling methods, we hypothesize that speaker characteristics are well represented in the…

音频与语音处理 · 电气工程与系统科学 2022-06-28 Yusheng Tian , Jingyu Li , Tan Lee

Automatic speech emotion recognition (SER) is a challenging task that plays a crucial role in natural human-computer interaction. One of the main challenges in SER is data scarcity, i.e., insufficient amounts of carefully labeled data to…

声音 · 计算机科学 2021-08-17 Sarala Padi , Seyed Omid Sadjadi , Dinesh Manocha , Ram D. Sriram

In the era of advanced artificial intelligence and human-computer interaction, identifying emotions in spoken language is paramount. This research explores the integration of deep learning techniques in speech emotion recognition, offering…

声音 · 计算机科学 2023-10-20 Hanan Hamza , Fiza Gafoor , Fathima Sithara , Gayathri Anil , V. S. Anoop

A speaker extraction algorithm seeks to extract the speech of a target speaker from a multi-talker speech mixture when given a cue that represents the target speaker, such as a pre-enrolled speech utterance, or an accompanying video track.…

音频与语音处理 · 电气工程与系统科学 2022-01-25 Zexu Pan , Ruijie Tao , Chenglin Xu , Haizhou Li

Speaker diarization has gained considerable attention within speech processing research community. Mainstream speaker diarization rely primarily on speakers' voice characteristics extracted from acoustic signals and often overlook the…

声音 · 计算机科学 2024-02-06 Luyao Cheng , Siqi Zheng , Qinglin Zhang , Hui Wang , Yafeng Chen , Qian Chen , Shiliang Zhang

Speech Self-Supervised Learning (SSL) has demonstrated considerable efficacy in various downstream tasks. Nevertheless, prevailing self-supervised models often overlook the incorporation of emotion-related prior information, thereby…

音频与语音处理 · 电气工程与系统科学 2024-06-12 Rui Liu , Zening Ma

Speech-driven 3D facial animation technology has been developed for years, but its practical application still lacks expectations. The main challenges lie in data limitations, lip alignment, and the naturalness of facial expressions.…

计算机视觉与模式识别 · 计算机科学 2024-04-30 Xiangyu Liang , Wenlin Zhuang , Tianyong Wang , Guangxing Geng , Guangyue Geng , Haifeng Xia , Siyu Xia

Training machines to understand natural language and interact with humans is one of the major goals of artificial intelligence. Recent years have witnessed an evolution from matching networks to pre-trained language models (PrLMs). In…

计算与语言 · 计算机科学 2023-01-12 Zhuosheng Zhang , Hai Zhao , Longxiang Liu

The prediction of valence from speech is an important, but challenging problem. The externalization of valence in speech has speaker-dependent cues, which contribute to performances that are often significantly lower than the prediction of…

声音 · 计算机科学 2023-05-15 Kusha Sridhar , Carlos Busso

It was shown that pre-trained models with self-supervised learning (SSL) techniques are effective in various downstream speech tasks. However, most such models are trained on single-speaker speech data, limiting their effectiveness in…

音频与语音处理 · 电气工程与系统科学 2024-07-04 Jingru Lin , Meng Ge , Junyi Ao , Liqun Deng , Haizhou Li

Large speech models-derived features have recently shown increased performance over signal-based features across multiple downstream tasks, even when the networks are not finetuned towards the target task. In this paper we show the results…

音频与语音处理 · 电气工程与系统科学 2023-11-02 Adrian Bogdan Stânea , Vlad Striletchi , Cosmin Striletchi , Adriana Stan

Self-supervised speech representations are known to encode both speaker and phonetic information, but how they are distributed in the high-dimensional space remains largely unexplored. We hypothesize that they are encoded in orthogonal…

计算与语言 · 计算机科学 2023-12-12 Oli Liu , Hao Tang , Sharon Goldwater

Automatically assessing emotional valence in human speech has historically been a difficult task for machine learning algorithms. The subtle changes in the voice of the speaker that are indicative of positive or negative emotional states…

计算与语言 · 计算机科学 2017-05-09 Jonathan Chang , Stefan Scherer