中文
相关论文

相关论文: BIAS: A Body-based Interpretable Active Speaker Ap…

200 篇论文

Speech-driven 3D facial animation with accurate lip synchronization has been widely studied. However, synthesizing realistic motions for the entire face during speech has rarely been explored. In this work, we present a joint audio-text…

计算机视觉与模式识别 · 计算机科学 2021-12-08 Yingruo Fan , Zhaojiang Lin , Jun Saito , Wenping Wang , Taku Komura

For speech interaction, voice activity detection (VAD) is often used as a front-end. However, traditional VAD algorithms usually need to wait for a continuous tail silence to reach a preset maximum duration before segmentation, resulting in…

音频与语音处理 · 电气工程与系统科学 2023-05-23 Mohan Shi , Yuchun Shu , Lingyun Zuo , Qian Chen , Shiliang Zhang , Jie Zhang , Li-Rong Dai

We introduce a new approach for audio-visual speech separation. Given a video, the goal is to extract the speech associated with a face in spite of simultaneous background sounds and/or other human speakers. Whereas existing methods focus…

计算机视觉与模式识别 · 计算机科学 2021-04-07 Ruohan Gao , Kristen Grauman

In this paper, we present a bias and sustainability focused investigation of Automatic Speech Recognition (ASR) systems, namely Whisper and Massively Multilingual Speech (MMS), which have achieved state-of-the-art (SOTA) performances.…

计算与语言 · 计算机科学 2025-03-04 Ajinkya Kulkarni , Atharva Kulkarni , Miguel Couceiro , Isabel Trancoso

With the widespread use of intelligent systems, such as smart speakers, addressee recognition has become a concern in human-computer interaction, as more and more people expect such systems to understand complicated social scenes, including…

人工智能 · 计算机科学 2018-09-13 Thao Minh Le , Nobuyuki Shimizu , Takashi Miyazaki , Koichi Shinoda

Lip reading aims to predict spoken language by analyzing lip movements. Despite advancements in lip reading technologies, performance degrades when models are applied to unseen speakers due to their sensitivity to variations in visual…

计算机视觉与模式识别 · 计算机科学 2025-01-03 Jeong Hun Yeo , Chae Won Kim , Hyunjun Kim , Hyeongseop Rha , Seunghee Han , Wen-Huang Cheng , Yong Man Ro

Autism Spectrum Disorder (ASD) is characterized by challenges with social interaction and communication and by restricted or repetitive patterns of thought and behavior, with significant variability in presentation. Approximately a quarter…

Large-scale multimodal models achieve strong results on tasks like Visual Question Answering (VQA), but they are often limited when queries require cultural and visual information, everyday knowledge, particularly in low-resource and…

The diagnosis of autism spectrum disorder (ASD) is a complex, challenging task as it depends on the analysis of interactional behaviors by psychologists rather than the use of biochemical diagnostics. In this paper, we present a modeling…

音频与语音处理 · 电气工程与系统科学 2024-01-19 Tahiya Chowdhury , Veronica Romero , Amanda Stent

Speaker-independent VSR is a complex task that involves identifying spoken words or phrases from video recordings of a speaker's facial movements. Over the years, there has been a considerable amount of research in the field of VSR…

计算机视觉与模式识别 · 计算机科学 2023-12-13 Praneeth Nemani , G. Sai Krishna , Supriya Kundrapu

Speech applications are expected to be low-power and robust under noisy conditions. An effective Voice Activity Detection (VAD) front-end lowers the computational need. Spiking Neural Networks (SNNs) are known to be biologically plausible…

声音 · 计算机科学 2024-03-12 Qu Yang , Qianhui Liu , Nan Li , Meng Ge , Zeyang Song , Haizhou Li

Speech enhancement plays an essential role in various applications, and the integration of visual information has been demonstrated to bring substantial advantages. However, the majority of current research concentrates on the examination…

声音 · 计算机科学 2025-04-03 Xinyuan Qian , Jiaran Gao , Yaodan Zhang , Qiquan Zhang , Hexin Liu , Leibny Paola Garcia , Haizhou Li

Attention-based encoder-decoder (AED) models have shown impressive performance in ASR. However, most existing AED methods neglect to simultaneously leverage both acoustic and semantic features in decoder, which is crucial for generating…

计算与语言 · 计算机科学 2023-05-24 Tian-Hao Zhang , Hai-Bo Qin , Zhi-Hao Lai , Song-Lu Chen , Qi Liu , Feng Chen , Xinyuan Qian , Xu-Cheng Yin

asya is a mobile application that consists of deep learning models which analyze spectra of a human voice and do noise detection, speaker diarization, gender detection, tempo estimation, and classification of emotions using only voice. All…

音频与语音处理 · 电气工程与系统科学 2020-08-21 Evalds Urtans , Ariel Tabaks

Vocal entrainment is a social adaptation mechanism in human interaction, knowledge of which can offer useful insights to an individual's cognitive-behavioral characteristics. We propose a context-aware approach for measuring vocal…

音频与语音处理 · 电气工程与系统科学 2022-11-08 Rimita Lahiri , Md Nasir , Catherine Lord , So Hyun Kim , Shrikanth Narayanan

Current speaker diarization systems rely on an external voice activity detection model prior to speaker embedding extraction on the detected speech segments. In this paper, we establish that the attention system of a speaker embedding…

音频与语音处理 · 电气工程与系统科学 2024-05-16 Jenthe Thienpondt , Kris Demuynck

Deep learning systems often struggle with processing long sequences, where computational complexity can become a bottleneck. Current methods for automated dementia detection using speech frequently rely on static, time-agnostic features or…

声音 · 计算机科学 2025-10-02 Chukwuemeka Ugwu , Oluwafemi Oyeleke

Methods that can generate synthetic speech which is perceptually indistinguishable from speech recorded by a human speaker, are easily available. Several incidents report misuse of synthetic speech generated from these methods to commit…

计算机视觉与模式识别 · 计算机科学 2024-04-18 Amit Kumar Singh Yadav , Kratika Bhagtani , Davide Salvi , Paolo Bestagini , Edward J. Delp

Speech activity detection (or endpointing) is an important processing step for applications such as speech recognition, language identification and speaker diarization. Both audio- and vision-based approaches have been used for this task in…

Recent advances in Visual Anomaly Detection (VAD) have introduced sophisticated algorithms leveraging embeddings generated by pre-trained feature extractors. Inspired by these developments, we investigate the adaptation of such algorithms…