English
Related papers

Related papers: VicKAM: Visual Conceptual Knowledge Guided Action …

200 papers

Weakly supervised visual grounding aims to predict the region in an image that corresponds to a specific linguistic query, where the mapping between the target object and query is unknown in the training stage. The state-of-the-art method…

Computer Vision and Pattern Recognition · Computer Science 2023-02-23 Viet-Quoc Pham , Nao Mishima

A popular approach to semantic image understanding is to manually tag images with keywords and then learn a mapping from vi- sual features to keywords. Manually tagging images is a subjective pro- cess and the same or very similar visual…

Computer Vision and Pattern Recognition · Computer Science 2016-09-08 Ke Sun , Xianxu Hou , Qian Zhang , Guoping Qiu

Inspired by recent advances in neural machine translation, that jointly align and translate using encoder-decoder networks equipped with attention, we propose an attentionbased LSTM model for human activity recognition. Our model jointly…

Computer Vision and Pattern Recognition · Computer Science 2017-09-01 Atousa Torabi , Leonid Sigal

The recent surge in interest in autonomous driving stems from its rapidly developing capacity to enhance safety, efficiency, and convenience. A pivotal aspect of autonomous driving technology is its perceptual systems, where core algorithms…

Computer Vision and Pattern Recognition · Computer Science 2023-11-02 Qi Zhang , Siyuan Gou , Wenbin Li

Vision-based human activity recognition (HAR) has made substantial progress in recognizing predefined gestures but lacks adaptability for emerging activities. This paper introduces a paradigm shift by harnessing generative modeling and…

Human-Computer Interaction · Computer Science 2023-12-13 Nikhil Kashyap , Manas Satish Bedmutha , Prerit Chaudhary , Brian Wood , Wanda Pratt , Janice Sabin , Andrea Hartzler , Nadir Weibel

Weakly supervised object detection aims at learning precise object detectors, given image category labels. In recent prevailing works, this problem is generally formulated as a multiple instance learning module guided by an image…

Computer Vision and Pattern Recognition · Computer Science 2019-04-02 Xiaoyan Li , Meina Kan , Shiguang Shan , Xilin Chen

Weakly supervised video anomaly detection (WS-VAD) involves identifying the temporal intervals that contain anomalous events in untrimmed videos, where only video-level annotations are provided as supervisory signals. However, a key…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Yu Wang , Shengjie Zhao

Long-term memory is increasingly important for personalized AI agents, yet existing benchmarks and methods remain largely text-centric. Even when images are included, the user-specific information needed for later questions is typically…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Viet Nguyen , Thao Nguyen , Vishal M. Patel , Yuheng Li

We present a weakly-supervised approach to segmenting proposed drivable paths in images with the goal of autonomous driving in complex urban environments. Using recorded routes from a data collection vehicle, our proposed method generates…

Robotics · Computer Science 2017-11-20 Dan Barnes , Will Maddern , Ingmar Posner

Classification networks can be used to localize and segment objects in images by means of class activation maps (CAMs). However, without pixel-level annotations, classification networks are known to (1) mainly focus on discriminative…

Computer Vision and Pattern Recognition · Computer Science 2023-04-05 Arvi Jonnarth , Michael Felsberg

Video-based computer vision tasks can benefit from estimation of the salient regions and interactions between those regions. Traditionally, this has been done by identifying the object regions in the images by utilizing pre-trained models…

Computer Vision and Pattern Recognition · Computer Science 2022-08-03 Arulkumar Subramaniam , Jayesh Vaidya , Muhammed Abdul Majeed Ameen , Athira Nambiar , Anurag Mittal

Human action recognition often struggles with deep semantic understanding, complex contextual information, and fine-grained distinction, limitations that traditional methods frequently encounter when dealing with diverse video data.…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Jingwei Peng , Zhixuan Qiu , Boyu Jin , Surasakdi Siripong

Weakly-Supervised Video Anomaly Detection aims to identify anomalous events using only video-level labels, balancing annotation efficiency with practical applicability. However, existing methods often oversimplify the anomaly space by…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Junhee Lee , ChaeBeen Bang , MyoungChul Kim , MyeongAh Cho

Action in video usually involves the interaction of human with objects. Action labels are typically composed of various combinations of verbs and nouns, but we may not have training data for all possible combinations. In this paper, we aim…

Computer Vision and Pattern Recognition · Computer Science 2022-07-06 Zhekun Luo , Shalini Ghosh , Devin Guillory , Keizo Kato , Trevor Darrell , Huijuan Xu

Group activity detection in multi-person scenes is challenging due to complex human interactions, occlusions, and variations in appearance over time. This work presents a computer vision based framework for group activity recognition and…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Narthana Sivalingam , Santhirarajah Sivasthigan , Thamayanthi Mahendranathan , G. M. R. I. Godaliyadda , M. P. B. Ekanayake , H. M. V. R. Herath

3D Convolutional Neural Network (3D CNN) captures spatial and temporal information on 3D data such as video sequences. However, due to the convolution and pooling mechanism, the information loss seems unavoidable. To improve the visual…

Computer Vision and Pattern Recognition · Computer Science 2022-08-17 Novanto Yudistira , Muthu Subash Kavitha , Takio Kurita

Weakly supervised temporal action localization is a challenging vision task due to the absence of ground-truth temporal locations of actions in the training videos. With only video-level supervision during training, most existing methods…

Computer Vision and Pattern Recognition · Computer Science 2021-03-26 Ashraful Islam , Chengjiang Long , Richard Radke

Kinematic trajectories recorded from surgical robots contain information about surgical gestures and potentially encode cues about surgeon's skill levels. Automatic segmentation of these trajectories into meaningful action units could help…

Robotics · Computer Science 2019-07-26 Beatrice van Amsterdam , Hirenkumar Nakawala , Elena De Momi , Danail Stoyanov

Collaborative perception is essential for networks of agents with limited sensing capabilities, enabling them to work together by exchanging information to achieve a robust and comprehensive understanding of their environment. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Jiuwu Hao , Liguo Sun , Ti Xiang , Yuting Wan , Haolin Song , Pin Lv

Interpretation and understanding of video presents a challenging computer vision task in numerous fields - e.g. autonomous driving and sports analytics. Existing approaches to interpreting the actions taking place within a video clip are…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Salman Khan , Izzeddin Teeti , Andrew Bradley , Mohamed Elhoseiny , Fabio Cuzzolin