中文
相关论文

相关论文: Visual gesture variability between talkers in cont…

200 篇论文

Learning continually from a stream of non-i.i.d. data is an open challenge in deep learning, even more so when working in resource-constrained environments such as embedded devices. Visual models that are continually updated through…

人工智能 · 计算机科学 2025-07-30 Clea Rebillard , Julio Hurtado , Andrii Krutsylo , Lucia Passaro , Vincenzo Lomonaco

Speaker verification (SV) systems are currently being used to make sensitive decisions like giving access to bank accounts or deciding whether the voice of a suspect coincides with that of the perpetrator of a crime. Ensuring that these…

音频与语音处理 · 电气工程与系统科学 2025-11-18 Mariel Estevez , Luciana Ferrer

The goal of this paper is to learn robust speaker representation for bilingual speaking scenario. The majority of the world's population speak at least two languages; however, most speaker recognition systems fail to recognise the same…

音频与语音处理 · 电气工程与系统科学 2023-06-08 Kihyun Nam , Youkyum Kim , Jaesung Huh , Hee Soo Heo , Jee-weon Jung , Joon Son Chung

In this work, we re-think the task of speech enhancement in unconstrained real-world environments. Current state-of-the-art methods use only the audio stream and are limited in their performance in a wide range of real-world noises. Recent…

计算机视觉与模式识别 · 计算机科学 2020-12-22 Sindhu B Hegde , K R Prajwal , Rudrabha Mukhopadhyay , Vinay Namboodiri , C. V. Jawahar

Lip reading involves interpreting a speaker's speech by analyzing sequences of lip movements. Currently, most models regard the left and right halves of the lips as a symmetrical whole, lacking a thorough investigation of their differences.…

计算机视觉与模式识别 · 计算机科学 2024-09-10 Zejun gu , Junxia jiang

In this project, we worked on speech recognition, specifically predicting individual words based on both the video frames and audio. Empowered by convolutional neural networks, the recent speech recognition and lip reading models are…

计算机视觉与模式识别 · 计算机科学 2018-12-27 Devesh Walawalkar , Yihui He , Rohit Pillai

Speech production is a dynamic procedure, which involved multi human organs including the tongue, jaw and lips. Modeling the dynamics of the vocal tract deformation is a fundamental problem to understand the speech, which is the most common…

音频与语音处理 · 电气工程与系统科学 2021-06-23 Haiyang Liu , Jihan Zhang

To join the advantages of classical and end-to-end approaches for speech recognition, we present a simple, novel and competitive approach for phoneme-based neural transducer modeling. Different alignment label topologies are compared and…

计算与语言 · 计算机科学 2021-04-21 Wei Zhou , Simon Berger , Ralf Schlüter , Hermann Ney

Active speaker detection requires a solid integration of multi-modal cues. While individual modalities can approximate a solution, accurate predictions can only be achieved by explicitly fusing the audio and visual features and modeling…

计算机视觉与模式识别 · 计算机科学 2021-10-06 Juan León-Alcázar , Fabian Caba Heilbron , Ali Thabet , Bernard Ghanem

The goal of this work is to train discriminative cross-modal embeddings without access to manually annotated data. Recent advances in self-supervised learning have shown that effective representations can be learnt from natural cross-modal…

声音 · 计算机科学 2020-11-05 Soo-Whan Chung , Hong Goo Kang , Joon Son Chung

Lipreading or visually recognizing speech from the mouth movements of a speaker is a challenging and mentally taxing task. Unfortunately, multiple medical conditions force people to depend on this skill in their day-to-day lives for…

计算机视觉与模式识别 · 计算机科学 2021-11-04 Bipasha Sen , Aditya Agarwal , Rudrabha Mukhopadhyay , Vinay Namboodiri , C V Jawahar

Audio-visual learning has demonstrated promising results in many classical speech tasks (e.g., speech separation, automatic speech recognition, wake-word spotting). We believe that introducing visual modality will also benefit speaker…

音频与语音处理 · 电气工程与系统科学 2025-08-01 Ming Cheng , Ming Li

Conventional audio-visual methods for speaker verification rely on large amounts of labeled data and separate modality-specific architectures, which is computationally expensive, limiting their scalability. To address these problems, we…

计算机视觉与模式识别 · 计算机科学 2025-06-25 Gnana Praveen Rajasekhar , Jahangir Alam

We present a transformer-based architecture for voice separation of a target speaker from multiple other speakers and ambient noise. We achieve this by using two separate neural networks: (A) An enrolment network designed to craft…

音频与语音处理 · 电气工程与系统科学 2025-01-03 Akam Rahimi , Triantafyllos Afouras , Andrew Zisserman

It is common in everyday spoken communication that we look at the turning head of a talker to listen to his/her voice. Humans see the talker to listen better, so do machines. However, previous studies on audio-visual speaker extraction have…

声音 · 计算机科学 2023-09-14 Qinghua Liu , Meng Ge , Zhizheng Wu , Haizhou Li

During language acquisition, infants have the benefit of visual cues to ground spoken language. Robots similarly have access to audio and visual sensors. Recent work has shown that images and spoken captions can be mapped into a meaningful…

计算与语言 · 计算机科学 2017-05-29 Herman Kamper , Shane Settle , Gregory Shakhnarovich , Karen Livescu

Recent years have seen a surge in finding association between faces and voices within a cross-modal biometric application along with speaker recognition. Inspired from this, we introduce a challenging task in establishing association…

计算机视觉与模式识别 · 计算机科学 2021-04-23 Muhammad Saad Saeed , Shah Nawaz , Pietro Morerio , Arif Mahmood , Ignazio Gallo , Muhammad Haroon Yousaf , Alessio Del Bue

Lip-to-speech involves generating a natural-sounding speech synchronized with a soundless video of a person talking. Despite recent advances, current methods still cannot produce high-quality speech with high levels of intelligibility for…

音频与语音处理 · 电气工程与系统科学 2024-03-29 Yochai Yemini , Aviv Shamsian , Lior Bracha , Sharon Gannot , Ethan Fetaya

As experts in voice modification, trans-feminine gender-affirming voice teachers have unique perspectives on voice that confound current understandings of speaker identity. To demonstrate this, we present the Versatile Voice Dataset (VVD),…

Speech recognition is very challenging in student learning environments that are characterized by significant cross-talk and background noise. To address this problem, we present a bilingual speech recognition system that uses an…