中文
相关论文

相关论文: FabuLight-ASD: Unveiling Speech Activity via Body …

200 篇论文

In this study, we aim to explore Multitask Speech Language Model (SpeechLM) efficient inference via token reduction. Unlike other modalities such as vision or text, speech has unique temporal dependencies, making previous efficient…

音频与语音处理 · 电气工程与系统科学 2024-10-07 Yichen Lu , Jiaqi Song , Chao-Han Huck Yang , Shinji Watanabe

The growing applications of AR/VR increase the demand for real-time full-body pose estimation from Head-Mounted Displays (HMDs). Although HMDs provide joint signals from the head and hands, reconstructing a full-body pose remains…

计算机视觉与模式识别 · 计算机科学 2025-04-28 Shuting Zhao , Linxin Bai , Liangjing Shao , Ye Zhang , Xinrong Chen

In this paper, we propose a visual embedding approach to improving embedding aware speech enhancement (EASE) by synchronizing visual lip frames at the phone and place of articulation levels. We first extract visual embedding from lip frames…

声音 · 计算机科学 2020-09-22 Hang Chen , Jun Du , Yu Hu , Li-Rong Dai , Bao-Cai Yin , Chin-Hui Lee

Active speaker detection (ASD) and virtual cinematography (VC) can significantly improve the remote user experience of a video conference by automatically panning, tilting and zooming of a video conferencing camera: users subjectively rate…

音频与语音处理 · 电气工程与系统科学 2022-05-26 Ross Cutler , Ramin Mehran , Sam Johnson , Cha Zhang , Adam Kirk , Oliver Whyte , Adarsh Kowdle

We propose a dataset, AVASpeech-SMAD, to assist speech and music activity detection research. With frame-level music labels, the proposed dataset extends the existing AVASpeech dataset, which originally consists of 45 hours of audio and…

音频与语音处理 · 电气工程与系统科学 2021-11-03 Yun-Ning Hung , Karn N. Watcharasupat , Chih-Wei Wu , Iroro Orife , Kelian Li , Pavan Seshadri , Junyoung Lee

Successful active speaker detection requires a three-stage pipeline: (i) audio-visual encoding for all speakers in the clip, (ii) inter-speaker relation modeling between a reference speaker and the background speakers within each frame, and…

计算机视觉与模式识别 · 计算机科学 2021-09-08 Okan Köpüklü , Maja Taseska , Gerhard Rigoll

Overlapped speech detection (OSD) is critical for speech applications in scenario of multi-party conversion. Despite numerous research efforts and progresses, comparing with speech activity detection (VAD), OSD remains an open challenge and…

声音 · 计算机科学 2022-09-27 Ziqing Du , Kai Liu , Xucheng Wan , Huan Zhou

Voice Activity Detection (VAD) is not easy task when the input audio signal is noisy, and it is even more complicated when the input is not even an audio recording. This is the case with Silent Speech Interfaces (SSI) where we record the…

声音 · 计算机科学 2021-09-21 Amin Honarmandi Shandiz , László Tóth

Passive human speed estimation plays a critical role in acoustic sensing. Despite extensive study, existing systems, however, suffer from various limitations: First, the channel measurement rate proves inadequate to estimate high moving…

人机交互 · 计算机科学 2025-10-16 Sheng Lyu , Chenshu Wu

Autism Spectrum Disorder (ASD) is a severe neuropsychiatric disorder that affects intellectual development, social behavior, and facial features, and the number of cases is still significantly increasing. Due to the variety of symptoms ASD…

图像与视频处理 · 电气工程与系统科学 2021-10-11 Ryan Liu , Spencer He

Robots are becoming everyday devices, increasing their interaction with humans. To make human-machine interaction more natural, cognitive features like Visual Voice Activity Detection (VVAD), which can detect whether a person is speaking or…

计算机视觉与模式识别 · 计算机科学 2024-01-03 Adrian Lubitz , Matias Valdenegro-Toro , Frank Kirchner

Deep learning and contactless sensing technologies have significantly advanced the automated assessment of human behaviors in healthcare. In the context of autism spectrum disorder (ASD), repetitive motor behaviors such as spinning, head…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Amit Kumar Singh , Vrijendra Singh

Voice activity detection (VAD) improves the performance of speaker verification (SV) by preserving speech segments and attenuating the effects of non-speech. However, this scheme is not ideal: (1) it fails in noisy environments or…

声音 · 计算机科学 2023-06-01 Zuheng Kang , Jianzong Wang , Junqing Peng , Jing Xiao

This paper presents an audio visual automatic speech recognition (AV-ASR) system using a Transformer-based architecture. We particularly focus on the scene context provided by the visual information, to ground the ASR. We extract…

音频与语音处理 · 电气工程与系统科学 2020-05-01 Georgios Paraskevopoulos , Srinivas Parthasarathy , Aparna Khare , Shiva Sundaram

In recent years, video conferencing applications have become increasingly prevalent, relying heavily on high-speed internet connectivity. When such connectivity is lacking, users often default to audio-only communication, a mode that…

多媒体 · 计算机科学 2025-11-12 Panneer Selvam Santhalingam , Swann Thantsin , Ahmad Kamari , Parth Pathak , Kenneth DeHaan

Most existing sound event detection~(SED) algorithms operate under a closed-set assumption, restricting their detection capabilities to predefined classes. While recent efforts have explored language-driven zero-shot SED by exploiting…

声音 · 计算机科学 2025-10-28 Pengfei Cai , Yan Song , Qing Gu , Nan Jiang , Haoyu Song , Ian McLoughlin

Recent multimodal large language models (MLLMs) such as GPT-4o and Qwen3-Omni show strong perception but struggle in multi-speaker, dialogue-centric settings that demand agentic reasoning tracking who speaks, maintaining roles, and…

Accurate and efficient auscultation-based diagnostics are vital for early disease detection, especially in resource-limited settings where specialized clinical expertise is scarce. Traditional auscultation, which heavily depends on…

声音 · 计算机科学 2025-03-26 Pingjie Wang , Liudan Zhao , Zihan Zhao , Miao He , Xin Sun , Ya Zhang , Kun Sun , Yanfeng Wang , Yu Wang

Robust voice activity detection (VAD) is a challenging task in low signal-to-noise (SNR) environments. Recent studies show that speech enhancement is helpful to VAD, but the performance improvement is limited. To address this issue, here we…

音频与语音处理 · 电气工程与系统科学 2021-04-14 Xu Tan , Xiao-Lei Zhang

We present a cross-modal unsupervised framework for active speaker detection in media content such as TV shows and movies. Machine learning advances have enabled impressive performance in identifying individuals from speech and facial…

图像与视频处理 · 电气工程与系统科学 2022-09-27 Rahul Sharma , Shrikanth Narayanan