English
Related papers

Related papers: No-audio speaking status detection in crowded sett…

200 papers

Public speaking and presentation competence plays an essential role in many areas of social interaction in our educational, professional, and everyday life. Since our intention during a speech can differ from what is actually understood by…

Computer Vision and Pattern Recognition · Computer Science 2021-05-07 Ömer Sümer , Cigdem Beyan , Fabian Ruth , Olaf Kramer , Ulrich Trautwein , Enkelejda Kasneci

With wearable IMU sensors, one can estimate human poses from wearable devices without requiring visual input~\cite{von2017sparse}. In this work, we pose the question: Can we reason about object structure in real-world environments solely…

Robotics · Computer Science 2022-07-15 Yinyu Nie , Angela Dai , Xiaoguang Han , Matthias Nießner

Humans can robustly recognize and localize objects by using visual and/or auditory cues. While machines are able to do the same with visual data already, less work has been done with sounds. This work develops an approach for scene…

Sound · Computer Science 2022-03-01 Dengxin Dai , Arun Balajee Vasudevan , Jiri Matas , Luc Van Gool

Multi-person pose estimation is fundamental to many computer vision tasks and has made significant progress in recent years. However, few previous methods explored the problem of pose estimation in crowded scenes while it remains…

Computer Vision and Pattern Recognition · Computer Science 2019-01-24 Jiefeng Li , Can Wang , Hao Zhu , Yihuan Mao , Hao-Shu Fang , Cewu Lu

Active speaker detection requires a solid integration of multi-modal cues. While individual modalities can approximate a solution, accurate predictions can only be achieved by explicitly fusing the audio and visual features and modeling…

Computer Vision and Pattern Recognition · Computer Science 2021-10-06 Juan León-Alcázar , Fabian Caba Heilbron , Ali Thabet , Bernard Ghanem

While most conversational AI systems focus on textual dialogue only, conditioning utterances on visual context (when it's available) can lead to more realistic conversations. Unfortunately, a major challenge for incorporating visual context…

Computer Vision and Pattern Recognition · Computer Science 2021-03-30 Paul Hongsuck Seo , Arsha Nagrani , Cordelia Schmid

Human speech is often accompanied by hand and arm gestures. Given audio speech input, we generate plausible gestures to go along with the sound. Specifically, we perform cross-modal translation from "in-the-wild'' monologue speech of a…

Computer Vision and Pattern Recognition · Computer Science 2019-06-11 Shiry Ginosar , Amir Bar , Gefen Kohavi , Caroline Chan , Andrew Owens , Jitendra Malik

In this work, we present a novel audio-visual dataset for active speaker detection in the wild. A speaker is considered active when his or her face is visible and the voice is audible simultaneously. Although active speaker detection is a…

Computer Vision and Pattern Recognition · Computer Science 2021-08-18 You Jin Kim , Hee-Soo Heo , Soyeon Choe , Soo-Whan Chung , Yoohwan Kwon , Bong-Jin Lee , Youngki Kwon , Joon Son Chung

Gestures are inherent to human interaction and often complement speech in face-to-face communication, forming a multimodal communication system. An important task in gesture analysis is detecting a gesture's beginning and end. Research on…

Computer Vision and Pattern Recognition · Computer Science 2024-04-24 Esam Ghaleb , Ilya Burenko , Marlou Rasenberg , Wim Pouw , Ivan Toni , Peter Uhrig , Anna Wilson , Judith Holler , Aslı Özyürek , Raquel Fernández

Multi-modal based speech separation has exhibited a specific advantage on isolating the target character in multi-talker noisy environments. Unfortunately, most of current separation strategies prefer a straightforward fusion based on…

Sound · Computer Science 2022-03-08 Junwen Xiong , Peng Zhang , Lei Xie , Wei Huang , Yufei Zha , Yanning Zhang

Dense video captioning is a task of localizing interesting events from an untrimmed video and producing textual description (captions) for each localized event. Most of the previous works in dense video captioning are solely based on visual…

Computer Vision and Pattern Recognition · Computer Science 2020-05-07 Vladimir Iashin , Esa Rahtu

State-of-the-art Active Speaker Detection (ASD) approaches mainly use audio and facial features as input. However, the main hypothesis in this paper is that body dynamics is also highly correlated to "speaking" (and "listening") actions and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-12 Tiago Roxo , Joana C. Costa , Pedro Inácio , Hugo Proença

In this paper, we present a system that associates faces with voices in a video by fusing information from the audio and visual signals. The thesis underlying our work is that an extremely simple approach to generating (weak) speech…

Multimedia · Computer Science 2017-06-02 Ken Hoover , Sourish Chaudhuri , Caroline Pantofaru , Malcolm Slaney , Ian Sturdy

This paper presents our solution to ACM MM challenge: Large-scale Human-centric Video Analysis in Complex Events\cite{lin2020human}; specifically, here we focus on Track3: Crowd Pose Tracking in Complex Events. Remarkable progress has been…

Computer Vision and Pattern Recognition · Computer Science 2020-10-22 Li Yuan , Shuning Chang , Ziyuan Huang , Yichen Zhou , Yunpeng Chen , Xuecheng Nie , Francis E. H. Tay , Jiashi Feng , Shuicheng Yan

Tracking body and hand motions in the 3D space is essential for social and self-presence in augmented and virtual environments. Unlike the popular 3D pose estimation setting, the problem is often formulated as inside-out tracking based on…

Computer Vision and Pattern Recognition · Computer Science 2020-12-08 Mathias Parger , Chengcheng Tang , Yuanlu Xu , Christopher Twigg , Lingling Tao , Yijing Li , Robert Wang , Markus Steinberger

Stance detection plays a pivotal role in enabling an extensive range of downstream applications, from discourse parsing to tracing the spread of fake news and the denial of scientific facts. While most stance classification models rely on…

Computation and Language · Computer Science 2024-12-13 Guy Barel , Oren Tsur , Dan Vilenchik

With the widespread use of intelligent systems, such as smart speakers, addressee recognition has become a concern in human-computer interaction, as more and more people expect such systems to understand complicated social scenes, including…

Artificial Intelligence · Computer Science 2018-09-13 Thao Minh Le , Nobuyuki Shimizu , Takashi Miyazaki , Koichi Shinoda

Human-centric visual understanding is an important desideratum for effective human-robot interaction. In order to navigate crowded public places, social robots must be able to interpret the activity of the surrounding humans. This paper…

Computer Vision and Pattern Recognition · Computer Science 2023-07-28 Shengnan Hu , Ce Zheng , Zixiang Zhou , Chen Chen , Gita Sukthankar

Under noisy conditions, automatic speech recognition (ASR) can greatly benefit from the addition of visual signals coming from a video of the speaker's face. However, when multiple candidate speakers are visible this traditionally requires…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-12 Otavio Braga , Olivier Siohan

Visual identification of gunmen in a crowd is a challenging problem, that requires resolving the association of a person with an object (firearm). We present a novel approach to address this problem, by defining human-object interaction…

Computer Vision and Pattern Recognition · Computer Science 2020-05-21 Abdul Basit , Muhammad Akhtar Munir , Mohsen Ali , Arif Mahmood