English
Related papers

Related papers: No-audio speaking status detection in crowded sett…

200 papers

Speech Activity Detection (SAD) systems often misclassify singing as speech, leading to degraded performance in applications such as dialogue enhancement and automatic speech recognition. We introduce Singing-Robust Speech Activity…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-11 Philipp Grundhuber , Mhd Modar Halimeh , Martin Strauß , Emanuël A. P. Habets

Audio-visual speaker tracking aims to determine the location of human targets in a scene using signals captured by a multi-sensor platform, whose accuracy and robustness can be improved by multi-modal fusion methods. Recently, several…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Yidi Li , Hong Liu , Bing Yang

We present a joint audio-visual model for isolating a single speech signal from a mixture of sounds such as other speakers and background noise. Solving this task using only audio as input is extremely challenging and does not provide an…

Understanding crowd behaviors in a large social event is crucial for event management. Passive WiFi sensing, by collecting WiFi probe requests sent from mobile devices, provides a better way to monitor crowds compared with people counters…

Social and Information Networks · Computer Science 2020-02-12 Yuren Zhou , Billy Pik Lik Lau , Zann Koh , Chau Yuen , Benny Kai Kiat Ng

The objective of this work is person-clustering in videos -- grouping characters according to their identity. Previous methods focus on the narrower task of face-clustering, and for the most part ignore other cues such as the person's…

Computer Vision and Pattern Recognition · Computer Science 2021-05-21 Andrew Brown , Vicky Kalogeiton , Andrew Zisserman

Visually grounded speech models learn from images paired with spoken captions. By tagging images with soft text labels using a trained visual classifier with a fixed vocabulary, previous work has shown that it is possible to train a model…

Computation and Language · Computer Science 2021-06-24 Kayode Olaleye , Herman Kamper

Humans are able to localize objects in the environment using both visual and auditory cues, integrating information from multiple modalities into a common reference frame. We introduce a system that can leverage unlabeled audio-visual data…

Computer Vision and Pattern Recognition · Computer Science 2019-10-28 Chuang Gan , Hang Zhao , Peihao Chen , David Cox , Antonio Torralba

Object referring has important applications, especially for human-machine interaction. While having received great attention, the task is mainly attacked with written language (text) as input rather than spoken language (speech), which is…

Computer Vision and Pattern Recognition · Computer Science 2017-12-06 Arun Balajee Vasudevan , Dengxin Dai , Luc Van Gool

We consider the navigation of mobile robots in crowded environments, for which onboard sensing of the crowd is typically limited by occlusions. We address the problem of inferring the human occupancy in the space around the robot, in blind…

Robotics · Computer Science 2021-09-20 Javad Amirian , Jean-Bernard Hayet , Julien Pettre

In the task of audio-visual sound source separation, which leverages visual information for sound source separation, identifying objects in an image is a crucial step prior to separating the sound source. However, existing methods that…

Computer Vision and Pattern Recognition · Computer Science 2022-03-31 Takashi Oya , Shohei Iwase , Shigeo Morishima

Activity recognition computer vision algorithms can be used to detect the presence of autism-related behaviors, including what are termed "restricted and repetitive behaviors", or stimming, by diagnostic instruments. The limited data that…

Computer Vision and Pattern Recognition · Computer Science 2021-01-12 Peter Washington , Aaron Kline , Onur Cezmi Mutlu , Emilie Leblanc , Cathy Hou , Nate Stockham , Kelley Paskov , Brianna Chrisman , Dennis P. Wall

Modern mobile devices are able to provide context-aware and personalized services to the users, by leveraging on their sensing capabilities to infer the activity and situation in which a person is currently involved. Current solutions for…

Machine Learning · Computer Science 2023-07-10 Mattia Giovanni Campana , Franca Delmastro

Embodiment can enhance conversational agents, such as increasing their perceived presence. This is typically achieved through visual representations of a virtual body; however, visual modalities are not always available, such as when users…

Human-Computer Interaction · Computer Science 2026-01-30 Yi Fei Cheng , Jarod Bloch , Alexander Wang , Andrea Bianchi , Anusha Withana , Anhong Guo , Laurie M. Heller , David Lindlbauer

Human activity recognition has become an attractive research area with the development of on-body wearable sensing technology. With comfortable electronic-textiles, sensors can be embedded into clothing so that it is possible to record…

Robotics · Computer Science 2022-09-26 Tianchen Shen , Irene Di Giulio , Matthew Howard

The availability of digital devices operated by voice is expanding rapidly. However, the applications of voice interfaces are still restricted. For example, speaking in public places becomes an annoyance to the surrounding people, and…

Human-Computer Interaction · Computer Science 2023-03-06 Naoki Kimura , Michinari Kono , Jun Rekimoto

Pose estimation in the wild is a challenging problem, particularly in situations of (i) occlusions of varying degrees and (ii) crowded outdoor scenes. Most of the existing studies of pose estimation did not report the performance in similar…

Computer Vision and Pattern Recognition · Computer Science 2020-02-18 Sudip Das , Perla Sai Raj Kishore , Ujjwal Bhattacharya

Videoconferencing is now a frequent mode of communication in both professional and informal settings, yet it often lacks the fluidity and enjoyment of in-person conversation. This study leverages multimodal machine learning to predict…

Machine Learning · Computer Science 2025-03-11 Andrew Chang , Viswadruth Akkaraju , Ray McFadden Cogliano , David Poeppel , Dustin Freeman

This paper proposes a novel approach for crowd counting in low to high density scenarios in static images. Current approaches cannot handle huge crowd diversity well and thus perform poorly in extreme cases, where the crowd density in…

Computer Vision and Pattern Recognition · Computer Science 2020-02-28 Usman Sajid , Hasan Sajid , Hongcheng Wang , Guanghui Wang

Current state-of-the-art speech recognition models are trained to map acoustic signals into sub-lexical units. While these models demonstrate superior performance, they remain vulnerable to out-of-distribution conditions such as background…

Sound · Computer Science 2024-10-10 Sagarika Alavilli , Annesya Banerjee , Gasser Elbanna , Annika Magaro

Frequent interactions between individuals are a fundamental challenge for pose estimation algorithms. Current pipelines either use an object detector together with a pose estimator (top-down approach), or localize all body parts first and…

Computer Vision and Pattern Recognition · Computer Science 2023-10-10 Mu Zhou , Lucas Stoffl , Mackenzie Weygandt Mathis , Alexander Mathis