中文
相关论文

相关论文: No-audio speaking status detection in crowded sett…

200 篇论文

We introduce a seemingly impossible task: given only an audio clip of someone speaking, decide which of two face images is the speaker. In this paper we study this, and a number of related cross-modal tasks, aimed at answering the question:…

计算机视觉与模式识别 · 计算机科学 2018-04-04 Arsha Nagrani , Samuel Albanie , Andrew Zisserman

The existing state-of-the-art method for audio-visual conditioned video prediction uses the latent codes of the audio-visual frames from a multimodal stochastic network and a frame encoder to predict the next visual frame. However, a direct…

计算机视觉与模式识别 · 计算机科学 2023-09-21 Yating Xu , Conghui Hu , Gim Hee Lee

Amodal recognition is the ability of the system to detect occluded objects. Most SOTA Visual Recognition systems lack the ability to perform amodal recognition. Few studies have achieved amodal recognition through passive prediction or…

计算机视觉与模式识别 · 计算机科学 2023-03-09 Venkatraman Narayanan , Bala Murali Manoghar , Rama Prashanth RV , Phu Pham , Aniket Bera

The body movements accompanying speech aid speakers in expressing their ideas. Co-speech motion generation is one of the important approaches for synthesizing realistic avatars. Due to the intricate correspondence between speech and motion,…

多媒体 · 计算机科学 2024-08-28 Sen Wang , Jiangning Zhang , Xin Tan , Zhifeng Xie , Chengjie Wang , Lizhuang Ma

State-of-the-art multi-object tracking~(MOT) methods follow the tracking-by-detection paradigm, where object trajectories are obtained by associating per-frame outputs of object detectors. In crowded scenes, however, detectors often fail to…

计算机视觉与模式识别 · 计算机科学 2021-02-03 Weihong Ren , Xinchao Wang , Jiandong Tian , Yandong Tang , Antoni B. Chan

The thud of a bouncing ball, the onset of speech as lips open -- when visual and audio events occur together, it suggests that there might be a common, underlying event that produced both signals. In this paper, we argue that the visual and…

计算机视觉与模式识别 · 计算机科学 2018-10-10 Andrew Owens , Alexei A. Efros

Identification and localization of sounds are both integral parts of computational auditory scene analysis. Although each can be solved separately, the goal of forming coherent auditory objects and achieving a comprehensive spatial scene…

声音 · 计算机科学 2019-12-24 Ivo Trowitzsch , Christopher Schymura , Dorothea Kolossa , Klaus Obermayer

Speech is understood better by using visual context; for this reason, there have been many attempts to use images to adapt automatic speech recognition (ASR) systems. Current work, however, has shown that visually adapted ASR models only…

计算与语言 · 计算机科学 2020-02-19 Tejas Srinivasan , Ramon Sanabria , Florian Metze

Active speaker detection (ASD) is a multi-modal task that aims to identify who, if anyone, is speaking from a set of candidates. Current audio-visual approaches for ASD typically rely on visually pre-extracted face tracks (sequences of…

音频与语音处理 · 电气工程与系统科学 2022-03-08 Davide Berghi , Adrian Hilton , Philip J. B. Jackson

Our goal is to isolate individual speakers from multi-talker simultaneous speech in videos. Existing works in this area have focussed on trying to separate utterances from known speakers in controlled environments. In this paper, we propose…

计算机视觉与模式识别 · 计算机科学 2018-06-20 Triantafyllos Afouras , Joon Son Chung , Andrew Zisserman

Smooth and seamless robot navigation while interacting with humans depends on predicting human movements. Forecasting such human dynamics often involves modeling human trajectories (global motion) or detailed body joint movements (local…

计算机视觉与模式识别 · 计算机科学 2020-07-15 Vida Adeli , Ehsan Adeli , Ian Reid , Juan Carlos Niebles , Hamid Rezatofighi

Speech activity detection (SAD), which often rests on the fact that the noise is "more" stationary than speech, is particularly challenging in non-stationary environments, because the time variance of the acoustic scene makes it difficult…

音频与语音处理 · 电气工程与系统科学 2020-07-29 Jens Heitkaemper , Joerg Schmalenstroeer , Reinhold Haeb-Umbach

It is now well established from a variety of studies that there is a significant benefit from combining video and audio data in detecting active speakers. However, either of the modalities can potentially mislead audiovisual fusion by…

Speaker diarization, the process of segmenting an audio stream or transcribed speech content into homogenous partitions based on speaker identity, plays a crucial role in the interpretation and analysis of human speech. Most existing…

机器学习 · 计算机科学 2024-08-23 Luyao Cheng , Hui Wang , Siqi Zheng , Yafeng Chen , Rongjie Huang , Qinglin Zhang , Qian Chen , Xihao Li

We present a review on the current state of publicly available datasets within the human action recognition community; highlighting the revival of pose based methods and recent progress of understanding person-person interaction modeling.…

计算机视觉与模式识别 · 计算机科学 2015-11-20 Michael Edwards , Jingjing Deng , Xianghua Xie

Condescending language use is caustic; it can bring dialogues to an end and bifurcate communities. Thus, systems for condescension detection could have a large positive impact. A challenge here is that condescension is often impossible to…

计算与语言 · 计算机科学 2019-09-26 Zijian Wang , Christopher Potts

In challenging real-life conditions such as extreme head-pose, occlusions, and low-resolution images where the visual information fails to estimate visual attention/gaze direction, audio signals could provide important and complementary…

计算机视觉与模式识别 · 计算机科学 2022-08-15 Shreya Ghosh , Abhinav Dhall , Munawar Hayat , Jarrod Knibbe

The raise of collaborative robotics has led to wide range of sensor technologies to detect human-machine interactions: at short distances, proximity sensors detect nontactile gestures virtually occlusion-free, while at medium distances,…

计算机视觉与模式识别 · 计算机科学 2019-10-17 Christoph Heindl , Markus Ikeda , Gernot Stübl , Andreas Pichler , Josef Scharinger

In this paper, we present a novel training method for speaker change detection models. Speaker change detection is often viewed as a binary sequence labelling problem. The main challenges with this approach are the vagueness of annotated…

音频与语音处理 · 电气工程与系统科学 2022-05-17 Joonas Kalda , Tanel Alumäe

Prosody plays a vital role in verbal communication. Acoustic cues of prosody have been examined extensively. However, prosodic characteristics are not only perceived auditorily, but also visually based on head and facial movements. The…

计算与语言 · 计算机科学 2022-09-14 Hartmut Meister , Isa Samira Winter , Moritz Waeachtler , Pascale Sandmann , Khaled Abdellatif
‹ 上一页 1 8 9 10 下一页 ›