中文
相关论文

相关论文: No-audio speaking status detection in crowded sett…

200 篇论文

The embedded sensors in widely used smartphones and other wearable devices make the data of human activities more accessible. However, recognizing different human activities from the wearable sensor data remains a challenging research…

机器学习 · 计算机科学 2023-07-25 Taoran Sheng , Manfred Huber

Detecting individual pedestrians in a crowd remains a challenging problem since the pedestrians often gather together and occlude each other in real-world scenarios. In this paper, we first explore how a state-of-the-art pedestrian detector…

计算机视觉与模式识别 · 计算机科学 2018-03-28 Xinlong Wang , Tete Xiao , Yuning Jiang , Shuai Shao , Jian Sun , Chunhua Shen

While various sensors have been deployed to monitor vehicular flows, sensing pedestrian movement is still nascent. Yet walking is a significant mode of travel in many cities, especially those in Europe, Africa, and Asia. Understanding…

音频与语音处理 · 电气工程与系统科学 2025-08-11 Chaeyeon Han , Pavan Seshadri , Yiwei Ding , Noah Posner , Bon Woo Koo , Animesh Agrawal , Alexander Lerch , Subhrajit Guhathakurta

Language-guided active sensing is a robotics subtask where a robot with an onboard sensor interacts efficiently with the environment via object manipulation to maximize perceptual information, following given language instructions. These…

机器人学 · 计算机科学 2024-02-06 Weihan Chen , Hanwen Ren , Ahmed H. Qureshi

Our objective is to transform a video into a set of discrete audio-visual objects using self-supervised learning. To this end, we introduce a model that uses attention to localize and group sound sources, and optical flow to aggregate…

计算机视觉与模式识别 · 计算机科学 2020-08-11 Triantafyllos Afouras , Andrew Owens , Joon Son Chung , Andrew Zisserman

Crowd scene analysis receives growing attention due to its wide applications. Grasping the accurate crowd location (rather than merely crowd count) is important for spatially identifying high-risk regions in congested scenes. In this paper,…

计算机视觉与模式识别 · 计算机科学 2020-01-28 Yao Xue , Siming Liu , Yonghui Li , Xueming Qian

Human pose forecasting predicts future poses based on past observations, and has many significant applications in areas such as action recognition, autonomous driving or human-robot interaction. This paper evaluates a wide range of pose…

计算机视觉与模式识别 · 计算机科学 2026-03-05 Daniel Bermuth , Alexander Poeppel , Wolfgang Reif

This paper studies audio-visual noise suppression for egocentric videos -- where the speaker is not captured in the video. Instead, potential noise sources are visible on screen with the camera emulating the off-screen speaker's view of the…

声音 · 计算机科学 2023-05-04 Roshan Sharma , Weipeng He , Ju Lin , Egor Lakomkin , Yang Liu , Kaustubh Kalgaonkar

Full body trackers are utilized for surveillance and security purposes, such as person-tracking robots. In the Middle East, uniform crowd environments are the norm which challenges state-of-the-art trackers. Despite tremendous improvements…

计算机视觉与模式识别 · 计算机科学 2022-09-07 Zhibo Zhang , Omar Alremeithi , Maryam Almheiri , Marwa Albeshr , Xiaoxiong Zhang , Sajid Javed , Naoufel Werghi

Surveillance cameras are widely applied for indoor occupancy measurement and human movement perception, which benefit for building energy management and social security. To address the challenges of limited view angle of single camera as…

计算机视觉与模式识别 · 计算机科学 2021-04-13 Ping Zhang , Zhenxiang Tao , Wenjie Yang , Minze Chen , Shan Ding , Xiaodong Liu , Rui Yang , Hui Zhang

Currently, the safety of people has become a very important problem in different places including subway station, universities, colleges, airport, shopping mall and square, city squares. Therefore, considering intelligence event detection…

计算机视觉与模式识别 · 计算机科学 2020-08-11 Constantinou Miti , Demetriou Zatte , Siraj Sajid Gondal

Human beings rely heavily on estimation of poses in order to access their body movements. Human pose estimation methods take advantage of computer vision advances in order to track human body movements in real life applications. This comes…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Arindam Chaudhuri

Speaker diarization consists of assigning speech signals to people engaged in a dialogue. An audio-visual spatiotemporal diarization model is proposed. The model is well suited for challenging scenarios that consist of several participants…

计算机视觉与模式识别 · 计算机科学 2018-10-15 Israel D. Gebru , Silèye Ba , Xiaofei Li , Radu Horaud

In this paper, we describe our study on how humans allocate their attention during visual crowd counting. Using an eye tracker, we collect gaze behavior of human participants who are tasked with counting the number of people in crowd…

计算机视觉与模式识别 · 计算机科学 2020-09-29 Raji Annadi , Yupei Chen , Viresh Ranjan , Dimitris Samaras , Gregory Zelinsky , Minh Hoai

This article proposes a novel attention-based body pose encoding for human activity recognition that presents a enriched representation of body-pose that is learned. The enriched data complements the 3D body joint position data and improves…

计算机视觉与模式识别 · 计算机科学 2020-10-05 B Debnath , M O'brien , S Kumar , A Behera

Audio grounding, or speech-driven open-set object detection, aims to localize and identify objects directly from speech, enabling generalization beyond predefined categories. This task is crucial for applications like human-robot…

声音 · 计算机科学 2025-09-23 Wenhuan Lu , Xinyue Song , Wenjun Ke , Zhizhi Yu , Wenhao Yang , Jianguo Wei

Most sound event detection (SED) systems perform well on clean datasets but degrade significantly in noisy environments. Language-queried audio source separation (LASS) models show promise for robust SED by separating target events;…

声音 · 计算机科学 2025-08-12 Yuanjian Chen , Yang Xiao , Han Yin , Yadong Guan , Xubo Liu

Interpersonal spoken communication is central to human interaction and the exchange of information. Such interactive processes involve not only speech and spoken language but also non-verbal cues such as hand gestures, facial expressions,…

声音 · 计算机科学 2022-12-20 Tiantian Feng , Shrikanth Narayanan

Automatic detection of emergent leaders in small groups from nonverbal behaviour is a growing research topic in social signal processing but existing methods were evaluated on single datasets -- an unrealistic assumption for real-world…

人机交互 · 计算机科学 2019-05-07 Philipp Müller , Andreas Bulling

Speaker clustering is an essential step in conventional speaker diarization systems and is typically addressed as an audio-only speech processing task. The language used by the participants in a conversation, however, carries additional…

音频与语音处理 · 电气工程与系统科学 2022-07-12 Nikolaos Flemotomos , Shrikanth Narayanan