中文
相关论文

相关论文: Speaker Verification in Emotional Talking Environm…

200 篇论文

Silent speech interfaces have been recently proposed as a way to enable communication when the acoustic signal is not available. This introduces the need to build visual speech recognition systems for silent and whispered speech. However,…

计算机视觉与模式识别 · 计算机科学 2018-02-20 Stavros Petridis , Jie Shen , Doruk Cetin , Maja Pantic

Emotional voice conversion (EVC) aims to change the emotional state of an utterance while preserving the linguistic content and speaker identity. In this paper, we propose a novel 2-stage training strategy for sequence-to-sequence emotional…

计算与语言 · 计算机科学 2021-06-10 Kun Zhou , Berrak Sisman , Haizhou Li

We present a voice conversion framework that converts normal speech into dysarthric speech while preserving the speaker identity. Such a framework is essential for (1) clinical decision making processes and alleviation of patient stress,…

Expressive voice conversion performs identity conversion for emotional speakers by jointly converting speaker identity and emotional style. Due to the hierarchical structure of speech emotion, it is challenging to disentangle the emotional…

音频与语音处理 · 电气工程与系统科学 2022-07-22 Zongyang Du , Berrak Sisman , Kun Zhou , Haizhou Li

Research on hate speech has predominantly revolved around detection and interpretation from textual inputs, leaving verbal content largely unexplored. While there has been limited exploration into hate speech detection within verbal…

计算与语言 · 计算机科学 2024-08-13 Jinmyeong An , Wonjun Lee , Yejin Jeon , Jungseul Ok , Yunsu Kim , Gary Geunbae Lee

Unlike other data modalities such as text and vision, speech does not lend itself to easy interpretation. While lay people can understand how to describe an image or sentence via perception, non-expert descriptions of speech often end at…

声音 · 计算机科学 2023-10-05 Robin Netzorg , Bohan Yu , Andrea Guzman , Peter Wu , Luna McNulty , Gopala Anumanchipalli

Speaker verification based on ad-hoc microphone arrays has the potential of reducing the error significantly in adverse acoustic environments. However, existing approaches extract utterance-level speaker embeddings from each channel of an…

声音 · 计算机科学 2022-03-29 Chengdong Liang , Yijiang Chen , Jiadi Yao , Xiao-Lei Zhang

Several works have developed end-to-end pipelines for generating lip-synced talking faces with various real-world applications, such as teaching and language translation in videos. However, these prior works fail to create realistic-looking…

计算机视觉与模式识别 · 计算机科学 2023-03-28 Sahil Goyal , Shagun Uppal , Sarthak Bhagat , Yi Yu , Yifang Yin , Rajiv Ratn Shah

Identifying multiple speakers without knowing where a speaker's voice is in a recording is a challenging task. In this paper, a hierarchical attention network is proposed to solve a weakly labelled speaker identification problem. The use of…

音频与语音处理 · 电气工程与系统科学 2020-08-28 Yanpei Shi , Qiang Huang , Thomas Hain

The human brain has the capability to associate the unknown person's voice and face by leveraging their general relationship, referred to as ``cross-modal speaker verification''. This task poses significant challenges due to the complex…

音频与语音处理 · 电气工程与系统科学 2024-07-26 Ruijie Tao , Zhan Shi , Yidi Jiang , Duc-Tuan Truong , Eng-Siong Chng , Massimo Alioto , Haizhou Li

In this paper, we propose "personal VAD", a system to detect the voice activity of a target speaker at the frame level. This system is useful for gating the inputs to a streaming on-device speech recognition system, such that it only…

音频与语音处理 · 电气工程与系统科学 2020-04-09 Shaojin Ding , Quan Wang , Shuo-yiin Chang , Li Wan , Ignacio Lopez Moreno

Speech emotion recognition (SER) has made significant strides with the advent of powerful self-supervised learning (SSL) models. However, the generalization of these models to diverse languages and emotional expressions remains a challenge.…

计算与语言 · 计算机科学 2024-08-16 Mohamed Osman , Daniel Z. Kaplan , Tamer Nadeem

Spoken Language Understanding (SLU) has progressed from traditional single-task methods to large audio language model (LALM) solutions. Yet, most existing speech benchmarks focus on single-speaker or isolated tasks, overlooking the…

音频与语音处理 · 电气工程与系统科学 2025-08-12 Shuai Wang , Zhaokai Sun , Zhennan Lin , Chengyou Wang , Zhou Pan , Lei Xie

Speaker identification in the household scenario (e.g., for smart speakers) is typically based on only a few enrollment utterances but a much larger set of unlabeled data, suggesting semisupervised learning to improve speaker profiles. We…

声音 · 计算机科学 2022-02-22 Long Chen , Venkatesh Ravichandran , Andreas Stolcke

Research on non-verbal behavior generation for social interactive agents focuses mainly on the believability and synchronization of non-verbal cues with speech. However, existing models, predominantly based on deep learning architectures,…

计算机视觉与模式识别 · 计算机科学 2024-10-11 Alice Delbosc , Magalie Ochs , Nicolas Sabouret , Brian Ravenet , Stephane Ayache

This paper addresses spoken language identification (SLI) and speech recognition of multilingual broadcast and institutional speech, real application scenarios that have been rarely addressed in the SLI literature. Observing that in these…

音频与语音处理 · 电气工程与系统科学 2024-06-14 Martina Valente , Fabio Brugnara , Giovanni Morrone , Enrico Zovato , Leonardo Badino

In this dissertation the practical speech emotion recognition technology is studied, including several cognitive related emotion types, namely fidgetiness, confidence and tiredness. The high quality of naturalistic emotional speech data is…

声音 · 计算机科学 2017-09-28 Chengwei Huang

Human emotions are inherently ambiguous and impure. When designing systems to anticipate human emotions based on speech, the lack of emotional purity must be considered. However, most of the current methods for speech emotion classification…

音频与语音处理 · 电气工程与系统科学 2020-08-13 Shuiyang Mao , P. C. Ching , Tan Lee

Speech emotion recognition (SER) systems aim to recognize human emotional state during human-computer interaction. Most existing SER systems are trained based on utterance-level labels. However, not all frames in an audio have affective…

声音 · 计算机科学 2023-12-29 Qifei Li , Yingming Gao , Cong Wang , Yayue Deng , Jinlong Xue , Yichen Han , Ya Li

This paper presents a visual passwords system to increase security. The system depends mainly on recognizing the speaker using the visual speech signal alone. The proposed scheme works in two stages: setting the visual password stage and…

计算机视觉与模式识别 · 计算机科学 2014-09-04 Ahmad Basheer Hassanat
‹ 上一页 1 8 9 10 下一页 ›