中文
相关论文

相关论文: Is Someone Speaking? Exploring Long-term Temporal …

200 篇论文

Speech applications dealing with conversations require not only recognizing the spoken words, but also determining who spoke when. The task of assigning words to speakers is typically addressed by merging the outputs of two separate…

计算与语言 · 计算机科学 2019-07-12 Laurent El Shafey , Hagen Soltau , Izhak Shafran

While speaking at different rates, articulators (like tongue, lips) tend to move differently and the enunciations are also of different durations. In the past, affine transformation and DNN have been used to transform articulatory movements…

音频与语音处理 · 电气工程与系统科学 2020-08-21 Abhayjeet Singh , Aravind Illa , Prasanta Kumar Ghosh

Turn-taking has played an essential role in structuring the regulation of a conversation. The task of identifying the main speaker (who is properly taking his/her turn of speaking) and the interrupters (who are interrupting or reacting to…

计算机视觉与模式识别 · 计算机科学 2021-08-29 Thanh-Dat Truong , Chi Nhan Duong , The De Vu , Hoang Anh Pham , Bhiksha Raj , Ngan Le , Khoa Luu

In speaker verification systems, the utilization of short utterances presents a persistent challenge, leading to performance degradation primarily due to insufficient phonetic information to characterize the speakers. To overcome this…

音频与语音处理 · 电气工程与系统科学 2024-06-12 Seung-bin Kim , Chan-yeong Lim , Jungwoo Heo , Ju-ho Kim , Hyun-seo Shin , Kyo-Won Koo , Ha-Jin Yu

Traditional speaker diarization seeks to detect ``who spoke when'' according to speaker characteristics. Extending to target speech diarization, we detect ``when target event occurs'' according to the semantic characteristics of speech. We…

音频与语音处理 · 电气工程与系统科学 2024-06-12 Yidi Jiang , Ruijie Tao , Zhengyang Chen , Yanmin Qian , Haizhou Li

Humans exhibit a remarkable ability to focus auditory attention in complex acoustic environments, such as cocktail parties. Auditory attention detection (AAD) aims to identify the attended speaker by analyzing brain signals, such as…

信号处理 · 电气工程与系统科学 2025-03-07 Yuan Liao , Yuhong Zhang , Qiushi Han , Yuhang Yang , Weiwei Ding , Yuzhe Gu , Hengxin Yang , Liya Huang

Active speaker detection plays a vital role in human-machine interaction. Recently, a few end-to-end audiovisual frameworks emerged. However, these models' inference time was not explored and are not applicable for real-time applications…

声音 · 计算机科学 2022-11-24 Fiseha B. Tesema , Zheyuan Lin , Shiqiang Zhu , Wei Song , Jason Gu , Hong Wu

Our goal is to isolate individual speakers from multi-talker simultaneous speech in videos. Existing works in this area have focussed on trying to separate utterances from known speakers in controlled environments. In this paper, we propose…

计算机视觉与模式识别 · 计算机科学 2018-06-20 Triantafyllos Afouras , Joon Son Chung , Andrew Zisserman

Speaker identification, determining which character said each utterance in literary text, benefits many downstream tasks. Most existing approaches use expert-defined rules or rule-based features to directly approach this task, but these…

计算与语言 · 计算机科学 2022-10-13 Ben Zhou , Dian Yu , Dong Yu , Dan Roth

In recent years, speech processing algorithms have seen tremendous progress primarily due to the deep learning renaissance. This is especially true for speech separation where the time-domain audio separation network (TasNet) has led to…

声音 · 计算机科学 2021-03-30 Morten Kolbæk , Zheng-Hua Tan , Søren Holdt Jensen , Jesper Jensen

Audio-Visual Speaker Detection (AVSD) hinges on modeling both individual temporal continuity and inter-personal social context. Existing coupled architectures struggle to reconcile these tasks in shared representation spaces due to…

多媒体 · 计算机科学 2026-04-17 Junhao Xiao , Shun Feng , Zhiyu Wu , Jinghan Yu , Haibiao Yao , Zhiyuan Ma , Jianjun Li , Youjun Bao , Yi Chen

Speech recognition (ASR) and speaker diarization (SD) models have traditionally been trained separately to produce rich conversation transcripts with speaker labels. Recent advances have shown that joint ASR and SD models can learn to…

音频与语音处理 · 电气工程与系统科学 2020-11-06 Huanru Henry Mao , Shuyang Li , Julian McAuley , Garrison Cottrell

Audio-visual speech enhancement (AVSE) methods use both audio and visual features for the task of speech enhancement and the use of visual features has been shown to be particularly effective in multi-speaker scenarios. In the majority of…

音频与语音处理 · 电气工程与系统科学 2022-02-18 Shrishti Saha Shetu , Soumitro Chakrabarty , Emanuël A. P. Habets

We study the problem of detecting talking activities in collaborative learning videos. Our approach uses head detection and projections of the log-magnitude of optical flow vectors to reduce the problem to a simple classification of small…

计算机视觉与模式识别 · 计算机科学 2021-10-18 Wenjing Shi , Marios S. Pattichis , Sylvia Celedón-Pattichis , Carlos LópezLeiva

In this paper, we propose "personal VAD", a system to detect the voice activity of a target speaker at the frame level. This system is useful for gating the inputs to a streaming on-device speech recognition system, such that it only…

音频与语音处理 · 电气工程与系统科学 2020-04-09 Shaojin Ding , Quan Wang , Shuo-yiin Chang , Li Wan , Ignacio Lopez Moreno

Lipreading is the task of decoding text from the movement of a speaker's mouth. Traditional approaches separated the problem into two stages: designing or learning visual features, and prediction. More recent deep lipreading approaches are…

机器学习 · 计算机科学 2016-12-19 Yannis M. Assael , Brendan Shillingford , Shimon Whiteson , Nando de Freitas

With the recent advancements in AI, Intelligent Virtual Assistants (IVA) have become a ubiquitous part of every home. Going forward, we are witnessing a confluence of vision, speech and dialog system technologies that are enabling the IVAs…

计算与语言 · 计算机科学 2018-12-21 Shachi H Kumar , Eda Okur , Saurav Sahay , Juan Jose Alvarado Leanos , Jonathan Huang , Lama Nachman

While the use of deep neural networks has significantly boosted speaker recognition performance, it is still challenging to separate speakers in poor acoustic environments. To improve robustness of speaker recognition system performance in…

音频与语音处理 · 电气工程与系统科学 2020-05-19 Yanpei Shi , Qiang Huang , Thomas Hain

Human auditory cortex excels at selectively suppressing background noise to focus on a target speaker. The process of selective attention in the brain is known to contextually exploit the available audio and visual cues to better focus on…

声音 · 计算机科学 2018-09-12 Mandar Gogate , Ahsan Adeel , Ricard Marxer , Jon Barker , Amir Hussain

Audio-visual target speaker extraction (AV-TSE) aims to extract the specific person's speech from the audio mixture given auxiliary visual cues. Previous methods usually search for the target voice through speech-lip synchronization.…

音频与语音处理 · 电气工程与系统科学 2025-03-04 Ruijie Tao , Xinyuan Qian , Yidi Jiang , Junjie Li , Jiadong Wang , Haizhou Li