English
Related papers

Related papers: Audio-visual Representation Learning for Anomaly E…

200 papers

Understanding crowd behaviors in a large social event is crucial for event management. Passive WiFi sensing, by collecting WiFi probe requests sent from mobile devices, provides a better way to monitor crowds compared with people counters…

Social and Information Networks · Computer Science 2020-02-12 Yuren Zhou , Billy Pik Lik Lau , Zann Koh , Chau Yuen , Benny Kai Kiat Ng

Most existing video anomaly detectors rely solely on RGB frames, which lack the temporal resolution needed to capture abrupt or transient motion cues, key indicators of anomalous events. To address this limitation, we propose Image-Event…

Computer Vision and Pattern Recognition · Computer Science 2025-05-09 Sungheon Jeong , Jihong Park , Mohsen Imani

A framework is proposed to detect anomalies in multi-modal data. A deep neural network-based object detector is employed to extract counts of objects and sub-events from the data. A cyclostationary model is proposed to model regular…

Signal Processing · Electrical Eng. & Systems 2018-07-19 Taposh Banerjee , Gene Whipps , Prudhvi Gurram , Vahid Tarokh

Most existing deep learning-based acoustic scene classification (ASC) approaches directly utilize representations extracted from spectrograms to identify target scenes. However, these approaches pay little attention to the audio events…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-03 Yuanbo Hou , Siyang Song , Chuang Yu , Yuxin Song , Wenwu Wang , Dick Botteldooren

Anomaly detection is concerned with identifying data patterns that deviate remarkably from the expected behaviour. This is an important research problem, due to its broad set of application domains, from data analysis to e-health,…

Machine Learning · Computer Science 2021-08-23 L. Erhan , M. Ndubuaku , M. Di Mauro , W. Song , M. Chen , G. Fortino , O. Bagdasar , A. Liotta

Large language models reveal deep comprehension and fluent generation in the field of multi-modality. Although significant advancements have been achieved in audio multi-modality, existing methods are rarely leverage language model for…

Sound · Computer Science 2024-08-06 Hualei Wang , Jianguo Mao , Zhifang Guo , Jiarui Wan , Hong Liu , Xiangdong Wang

We learn rich natural sound representations by capitalizing on large amounts of unlabeled sound data collected in the wild. We leverage the natural synchronization between vision and sound to learn an acoustic representation using…

Computer Vision and Pattern Recognition · Computer Science 2016-10-31 Yusuf Aytar , Carl Vondrick , Antonio Torralba

Self-supervised learning allows for better utilization of unlabelled data. The feature representation obtained by self-supervision can be used in downstream tasks such as classification, object detection, segmentation, and anomaly…

Computer Vision and Pattern Recognition · Computer Science 2020-06-18 Rabia Ali , Muhammad Umar Karim Khan , Chong Min Kyung

Audio content analysis in terms of sound events is an important research problem for a variety of applications. Recently, the development of weak labeling approaches for audio or sound event detection (AED) and availability of large scale…

Sound · Computer Science 2018-04-26 Ankit Shah , Anurag Kumar , Alexander G. Hauptmann , Bhiksha Raj

The pervasiveness and availability of mobile phone data offer the opportunity of discovering usable knowledge about crowd behaviors in urban environments. Cities can leverage such knowledge in order to provide better services (e.g., public…

Social and Information Networks · Computer Science 2015-04-15 Yuxiao Dong , Fabio Pinelli , Yiannis Gkoufas , Zubair Nabi , Francesco Calabrese , Nitesh V. Chawla

The Dynamic Saliency Prediction (DSP) task simulates the human selective attention mechanism to perceive the dynamic scene, which is significant and imperative in many vision tasks. Most of existing methods only consider visual cues, while…

Computer Vision and Pattern Recognition · Computer Science 2022-05-03 Hailong Ning , Bin Zhao , Zhanxuan Hu , Lang He , Ercheng Pei

Crowd scene analysis receives growing attention due to its wide applications. Grasping the accurate crowd location (rather than merely crowd count) is important for spatially identifying high-risk regions in congested scenes. In this paper,…

Computer Vision and Pattern Recognition · Computer Science 2020-01-28 Yao Xue , Siming Liu , Yonghui Li , Xueming Qian

Video anomaly detection is an essential but challenging task. The prevalent methods mainly investigate the reconstruction difference between normal and abnormal patterns but ignore the semantics consistency between appearance and motion…

Computer Vision and Pattern Recognition · Computer Science 2023-03-10 Xiangyu Huang , Caidan Zhao , Zhiqiang Wu

Personalized advertisement is a crucial task for many of the online businesses and video broadcasters. Many of today's broadcasters use the same commercial for all customers, but as one can imagine different viewers have different interests…

Computer Vision and Pattern Recognition · Computer Science 2018-06-25 Shervin Minaee , Imed Bouazizi , Prakash Kolan , Hossein Najafzadeh

Anomaly detection in surveillance videos remains a challenging task due to the diversity of abnormal events, class imbalance, and scene-dependent visual clutter. To address these issues, we propose a robust deep learning framework that…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Mohammad Ali Etemadi Naeen , Hoda Mohammadzade , Saeed Bagheri Shouraki

Pattern recognition from audio signals is an active research topic encompassing audio tagging, acoustic scene classification, music classification, and other areas. Spectrogram and mel-frequency cepstral coefficients (MFCC) are among the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-18 Md. Istiaq Ansari , Taufiq Hasan

Deep learning occupies an undisputed dominance in crowd counting. In this paper, we propose a novel convolutional neural network (CNN) architecture called SegCrowdNet. Despite the complex background in crowd scenes, the proposeSegCrowdNet…

Computer Vision and Pattern Recognition · Computer Science 2022-04-18 Jiwei Chen , Zengfu Wang

In this paper we advance the state-of-the-art for crowd counting in high density scenes by further exploring the idea of a fully convolutional crowd counting model introduced by (Zhang et al., 2016). Producing an accurate and robust crowd…

Computer Vision and Pattern Recognition · Computer Science 2017-01-18 Mark Marsden , Kevin McGuinness , Suzanne Little , Noel E. O'Connor

Multimodal acoustic event classification plays a key role in audio-visual systems. Although combining audio and visual signals improves recognition, it is still difficult to align them over time and to reduce the effect of noise across…

Sound · Computer Science 2025-09-19 Yuanjian Chen , Yang Xiao , Jinjie Huang

Sound event detection is a core module for acoustic environmental analysis. Semi-supervised learning technique allows to largely scale up the dataset without increasing the annotation budget, and recently attracts lots of research…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-02 Xiaofei Li
‹ Prev 1 8 9 10 Next ›