中文
相关论文

相关论文: Haha-Pod: An Attempt for Laughter-based Non-Verbal…

200 篇论文

We introduce a method to identify speakers by computing with high-dimensional random vectors. Its strengths are simplicity and speed. With only 1.02k active parameters and a 128-minute pass through the training data we achieve Top-1 and…

声音 · 计算机科学 2022-08-30 Ping-Chen Huang , Denis Kleyko , Jan M. Rabaey , Bruno A. Olshausen , Pentti Kanerva

Voice conversion (VC) using deep learning technologies can now generate high quality one-to-many voices and thus has been used in some practical application fields, such as entertainment and healthcare. However, voice conversion can pose…

声音 · 计算机科学 2024-05-02 Qiang Huang

Sarcasm detection and humor classification are inherently subtle problems, primarily due to their dependence on the contextual and non-verbal information. Furthermore, existing studies in these two topics are usually constrained in…

计算与语言 · 计算机科学 2021-06-01 Manjot Bedi , Shivani Kumar , Md Shad Akhtar , Tanmoy Chakraborty

The goal of this work is to reconstruct speech from a silent talking face video. Recent studies have shown impressive performance on synthesizing speech from silent talking face videos. However, they have not explicitly considered on…

计算机视觉与模式识别 · 计算机科学 2022-07-21 Joanna Hong , Minsu Kim , Yong Man Ro

Verifying if two audio segments belong to the same speaker has been recently put forward as a flexible way to carry out speaker identification, since it does not require to be re-trained when new speakers appear on the auditory scene.…

音频与语音处理 · 电气工程与系统科学 2020-10-26 Ivette Velez , Caleb Rascon , Gibran Fuentes-Pineda

Optimization of a trade-off between the number of speakers and their temporal variability (or session diversity) is crucial for the development of a speaker recognition system together with making the data collection process feasible from a…

音频与语音处理 · 电气工程与系统科学 2024-11-13 Anton Okhotnikov , Nikita Torgashov , Ivan Yakovlev , Pavel Malov , Rostislav Makarov

The short duration of an input utterance is one of the most critical threats that degrade the performance of speaker verification systems. This study aimed to develop an integrated text-independent speaker verification system that inputs…

音频与语音处理 · 电气工程与系统科学 2019-04-11 Jee-weon Jung , Hee-soo Heo , Hye-jin Shim , Ha-jin Yu

Deep speaker embeddings have become the leading method for encoding speaker identity in speaker recognition tasks. The embedding space should ideally capture the variations between all possible speakers, encoding the multiple acoustic…

声音 · 计算机科学 2021-04-26 Chau Luu , Peter Bell , Steve Renals

In this paper, an architecture based on Long Short-Term Memory Networks has been proposed for the text-independent scenario which is aimed to capture the temporal speaker-related information by operating over traditional speech features.…

音频与语音处理 · 电气工程与系统科学 2018-09-10 Aryan Mobiny , Mohammad Najarian

Audio-visual speech enhancement aims to extract clean speech from a noisy environment by leveraging not only the audio itself but also the target speaker's lip movements. This approach has been shown to yield improvements over audio-only…

The difficulty of acquiring abundant, high-quality data, especially in multi-lingual contexts, has sparked interest in addressing low-resource scenarios. Moreover, current literature rely on fixed expressions from language IDs, which…

声音 · 计算机科学 2024-09-30 Youngjae Kim , Yejin Jeon , Gary Geunbae Lee

When video is shot in noisy environment, the voice of a speaker seen in the video can be enhanced using the visible mouth movements, reducing background noise. While most existing methods use audio-only inputs, improved performance is…

计算机视觉与模式识别 · 计算机科学 2018-06-14 Aviv Gabbay , Asaph Shamir , Shmuel Peleg

Podcasts are a popular medium on the web, featuring diverse and multilingual content that often includes unverified claims. Fact-checking podcasts is a challenging task, requiring transcription, annotation, and claim verification, all while…

计算与语言 · 计算机科学 2025-02-04 Vinay Setty , Adam James Becker

In active speaker detection (ASD), we would like to detect whether an on-screen person is speaking based on audio-visual cues. Previous studies have primarily focused on modeling audio-visual synchronization cue, which depends on the video…

音频与语音处理 · 电气工程与系统科学 2023-06-13 Yidi Jiang , Ruijie Tao , Zexu Pan , Haizhou Li

This paper introduces VoxSim, a dataset of perceptual voice similarity ratings. Recent efforts to automate the assessment of speech synthesis technologies have primarily focused on predicting mean opinion score of naturalness, leaving…

音频与语音处理 · 电气工程与系统科学 2024-07-29 Junseok Ahn , Youkyum Kim , Yeunju Choi , Doyeop Kwak , Ji-Hoon Kim , Seongkyu Mun , Joon Son Chung

Whether it be for results summarization, or the analysis of classifier fusion, some means to compare different classifiers can often provide illuminating insight into their behaviour, (dis)similarity or complementarity. We propose a simple…

The growing capabilities of large language models and multimodal systems have spurred interest in voice-first AI assistants, yet existing benchmarks are inadequate for evaluating the full range of these systems' capabilities. We introduce…

计算与语言 · 计算机科学 2025-09-29 Ke Wang , Houxing Ren , Zimu Lu , Mingjie Zhan , Hongsheng Li

This paper investigates the application of environmental feature representations for room verification tasks and acoustic meta-data estimation. Audio recordings contain both speaker and non-speaker information. We refer to the…

声音 · 计算机科学 2022-03-10 Desmond Caulley

An important step in speaker verification is extracting features that best characterize the speaker voice. This paper investigates a front-end processing that aims at improving the performance of speaker verification based on the SVMs…

机器学习 · 计算机科学 2013-06-13 Kawthar Yasmine Zergat , Abderrahmane Amrouche

Active authentication refers to a new mode of identity verification in which biometric indicators are continuously tested to provide real-time or near real-time monitoring of an authorized access to a service or use of a device. This is in…

音频与语音处理 · 电气工程与系统科学 2020-04-28 Zhong Meng , M Umair Bin Altaf , Biing-Hwang , Juang