English
Related papers

Related papers: Is Someone Speaking? Exploring Long-term Temporal …

200 papers

Speech applications dealing with conversations require not only recognizing the spoken words, but also determining who spoke when. The task of assigning words to speakers is typically addressed by merging the outputs of two separate…

Computation and Language · Computer Science 2019-07-12 Laurent El Shafey , Hagen Soltau , Izhak Shafran

While speaking at different rates, articulators (like tongue, lips) tend to move differently and the enunciations are also of different durations. In the past, affine transformation and DNN have been used to transform articulatory movements…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-21 Abhayjeet Singh , Aravind Illa , Prasanta Kumar Ghosh

Turn-taking has played an essential role in structuring the regulation of a conversation. The task of identifying the main speaker (who is properly taking his/her turn of speaking) and the interrupters (who are interrupting or reacting to…

Computer Vision and Pattern Recognition · Computer Science 2021-08-29 Thanh-Dat Truong , Chi Nhan Duong , The De Vu , Hoang Anh Pham , Bhiksha Raj , Ngan Le , Khoa Luu

In speaker verification systems, the utilization of short utterances presents a persistent challenge, leading to performance degradation primarily due to insufficient phonetic information to characterize the speakers. To overcome this…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-12 Seung-bin Kim , Chan-yeong Lim , Jungwoo Heo , Ju-ho Kim , Hyun-seo Shin , Kyo-Won Koo , Ha-Jin Yu

Traditional speaker diarization seeks to detect ``who spoke when'' according to speaker characteristics. Extending to target speech diarization, we detect ``when target event occurs'' according to the semantic characteristics of speech. We…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-12 Yidi Jiang , Ruijie Tao , Zhengyang Chen , Yanmin Qian , Haizhou Li

Humans exhibit a remarkable ability to focus auditory attention in complex acoustic environments, such as cocktail parties. Auditory attention detection (AAD) aims to identify the attended speaker by analyzing brain signals, such as…

Signal Processing · Electrical Eng. & Systems 2025-03-07 Yuan Liao , Yuhong Zhang , Qiushi Han , Yuhang Yang , Weiwei Ding , Yuzhe Gu , Hengxin Yang , Liya Huang

Active speaker detection plays a vital role in human-machine interaction. Recently, a few end-to-end audiovisual frameworks emerged. However, these models' inference time was not explored and are not applicable for real-time applications…

Sound · Computer Science 2022-11-24 Fiseha B. Tesema , Zheyuan Lin , Shiqiang Zhu , Wei Song , Jason Gu , Hong Wu

Our goal is to isolate individual speakers from multi-talker simultaneous speech in videos. Existing works in this area have focussed on trying to separate utterances from known speakers in controlled environments. In this paper, we propose…

Computer Vision and Pattern Recognition · Computer Science 2018-06-20 Triantafyllos Afouras , Joon Son Chung , Andrew Zisserman

Speaker identification, determining which character said each utterance in literary text, benefits many downstream tasks. Most existing approaches use expert-defined rules or rule-based features to directly approach this task, but these…

Computation and Language · Computer Science 2022-10-13 Ben Zhou , Dian Yu , Dong Yu , Dan Roth

In recent years, speech processing algorithms have seen tremendous progress primarily due to the deep learning renaissance. This is especially true for speech separation where the time-domain audio separation network (TasNet) has led to…

Sound · Computer Science 2021-03-30 Morten Kolbæk , Zheng-Hua Tan , Søren Holdt Jensen , Jesper Jensen

Audio-Visual Speaker Detection (AVSD) hinges on modeling both individual temporal continuity and inter-personal social context. Existing coupled architectures struggle to reconcile these tasks in shared representation spaces due to…

Multimedia · Computer Science 2026-04-17 Junhao Xiao , Shun Feng , Zhiyu Wu , Jinghan Yu , Haibiao Yao , Zhiyuan Ma , Jianjun Li , Youjun Bao , Yi Chen

Speech recognition (ASR) and speaker diarization (SD) models have traditionally been trained separately to produce rich conversation transcripts with speaker labels. Recent advances have shown that joint ASR and SD models can learn to…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-06 Huanru Henry Mao , Shuyang Li , Julian McAuley , Garrison Cottrell

Audio-visual speech enhancement (AVSE) methods use both audio and visual features for the task of speech enhancement and the use of visual features has been shown to be particularly effective in multi-speaker scenarios. In the majority of…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-18 Shrishti Saha Shetu , Soumitro Chakrabarty , Emanuël A. P. Habets

We study the problem of detecting talking activities in collaborative learning videos. Our approach uses head detection and projections of the log-magnitude of optical flow vectors to reduce the problem to a simple classification of small…

Computer Vision and Pattern Recognition · Computer Science 2021-10-18 Wenjing Shi , Marios S. Pattichis , Sylvia Celedón-Pattichis , Carlos LópezLeiva

In this paper, we propose "personal VAD", a system to detect the voice activity of a target speaker at the frame level. This system is useful for gating the inputs to a streaming on-device speech recognition system, such that it only…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-09 Shaojin Ding , Quan Wang , Shuo-yiin Chang , Li Wan , Ignacio Lopez Moreno

Lipreading is the task of decoding text from the movement of a speaker's mouth. Traditional approaches separated the problem into two stages: designing or learning visual features, and prediction. More recent deep lipreading approaches are…

Machine Learning · Computer Science 2016-12-19 Yannis M. Assael , Brendan Shillingford , Shimon Whiteson , Nando de Freitas

With the recent advancements in AI, Intelligent Virtual Assistants (IVA) have become a ubiquitous part of every home. Going forward, we are witnessing a confluence of vision, speech and dialog system technologies that are enabling the IVAs…

Computation and Language · Computer Science 2018-12-21 Shachi H Kumar , Eda Okur , Saurav Sahay , Juan Jose Alvarado Leanos , Jonathan Huang , Lama Nachman

While the use of deep neural networks has significantly boosted speaker recognition performance, it is still challenging to separate speakers in poor acoustic environments. To improve robustness of speaker recognition system performance in…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-19 Yanpei Shi , Qiang Huang , Thomas Hain

Human auditory cortex excels at selectively suppressing background noise to focus on a target speaker. The process of selective attention in the brain is known to contextually exploit the available audio and visual cues to better focus on…

Sound · Computer Science 2018-09-12 Mandar Gogate , Ahsan Adeel , Ricard Marxer , Jon Barker , Amir Hussain

Audio-visual target speaker extraction (AV-TSE) aims to extract the specific person's speech from the audio mixture given auxiliary visual cues. Previous methods usually search for the target voice through speech-lip synchronization.…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-04 Ruijie Tao , Xinyuan Qian , Yidi Jiang , Junjie Li , Jiadong Wang , Haizhou Li
‹ Prev 1 3 4 5 6 7 10 Next ›