English
Related papers

Related papers: PO-GUISE+: Pose and object guided transformer toke…

200 papers

Attention is sparse in vision transformers. We observe the final prediction in vision transformers is only based on a subset of most informative tokens, which is sufficient for accurate image recognition. Based on this observation, we…

Computer Vision and Pattern Recognition · Computer Science 2021-10-27 Yongming Rao , Wenliang Zhao , Benlin Liu , Jiwen Lu , Jie Zhou , Cho-Jui Hsieh

The capability to automatically detect human stress can benefit artificial intelligent agents involved in affective computing and human-computer interaction. Stress and emotion are both human affective states, and stress has proven to have…

Computation and Language · Computer Science 2021-05-19 Yiqun Yao , Michalis Papakostas , Mihai Burzo , Mohamed Abouelenien , Rada Mihalcea

In this paper, we propose a novel token selective attention approach, ToSA, which can identify tokens that need to be attended as well as those that can skip a transformer layer. More specifically, a token selector parses the current…

Computer Vision and Pattern Recognition · Computer Science 2024-06-14 Manish Kumar Singh , Rajeev Yasarla , Hong Cai , Mingu Lee , Fatih Porikli

Visual Saliency refers to the innate human mechanism of focusing on and extracting important features from the observed environment. Recently, there has been a notable surge of interest in the field of automotive research regarding the…

Computer Vision and Pattern Recognition · Computer Science 2023-08-09 Francesco Rundo , Michael Sebastian Rundo , Concetto Spampinato

Recent advances in deep learning have enabled the generation of videos from textual descriptions as well as the prediction of future sequences from input videos. Similarly, in human motion modeling, motions can be generated from text or…

Computer Vision and Pattern Recognition · Computer Science 2026-04-27 Masato Soga , Ryuki Takebayashi

Object discovery -- separating objects from the background without manual labels -- is a fundamental open challenge in computer vision. Previous methods struggle to go beyond clustering of low-level cues, whether handcrafted (e.g., color,…

Computer Vision and Pattern Recognition · Computer Science 2023-03-29 Zhipeng Bao , Pavel Tokmakov , Yu-Xiong Wang , Adrien Gaidon , Martial Hebert

Accurate understanding and prediction of human behaviors are critical prerequisites for autonomous vehicles, especially in highly dynamic and interactive scenarios such as intersections in dense urban areas. In this work, we aim at…

Computer Vision and Pattern Recognition · Computer Science 2023-06-05 Jiachen Li , Xinwei Shi , Feiyu Chen , Jonathan Stroud , Zhishuai Zhang , Tian Lan , Junhua Mao , Jeonhyung Kang , Khaled S. Refaat , Weilong Yang , Eugene Ie , Congcong Li

Human motion prediction combines the tasks of trajectory forecasting and human pose prediction. For each of the two tasks, specialized models have been developed. Combining these models for holistic human motion prediction is non-trivial,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Aadya Agrawal , Alexander Schwing

Existing methods in video action recognition mostly do not distinguish human body from the environment and easily overfit the scenes and objects. In this work, we present a conceptually simple, general and high-performance framework for…

Computer Vision and Pattern Recognition · Computer Science 2018-12-18 Jiagang Zhu , Wei Zou , Liang Xu , Yiming Hu , Zheng Zhu , Manyu Chang , Junjie Huang , Guan Huang , Dalong Du

Collecting multi-view driving scenario videos to enhance the performance of 3D visual perception tasks presents significant challenges and incurs substantial costs, making generative models for realistic data an appealing alternative. Yet,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Junpeng Jiang , Gangyi Hong , Miao Zhang , Hengtong Hu , Kun Zhan , Rui Shao , Liqiang Nie

3D hand pose is an underexplored modality for action recognition. Poses are compact yet informative and can greatly benefit applications with limited compute budgets. However, poses alone offer an incomplete understanding of actions, as…

Computer Vision and Pattern Recognition · Computer Science 2024-08-15 Md Salman Shamil , Dibyadip Chatterjee , Fadime Sener , Shugao Ma , Angela Yao

Recently, the scientific progress of Advanced Driver Assistance System solutions (ADAS) has played a key role in enhancing the overall safety of driving. ADAS technology enables active control of vehicles to prevent potentially risky…

Signal Processing · Electrical Eng. & Systems 2023-08-07 Francesco Rundo , Concetto Spampinato , Michael Rundo

In video understanding tasks, particularly those involving human motion, synthetic data generation often suffers from uncanny features, diminishing its effectiveness for training. Tasks such as sign language translation, gesture…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Vaclav Knapp , Matyas Bohacek

The study and modeling of driver's gaze dynamics is important because, if and how the driver is monitoring the driving environment is vital for driver assistance in manual mode, for take-over requests in highly automated mode and for…

Computer Vision and Pattern Recognition · Computer Science 2018-02-02 Sujitha Martin , Sourabh Vora , Kevan Yuen , Mohan M. Trivedi

Transformer has demonstrated promising performance in many 2D vision tasks. However, it is cumbersome to compute the self-attention on large-scale point cloud data because point cloud is a long sequence and unevenly distributed in 3D space.…

Computer Vision and Pattern Recognition · Computer Science 2022-03-22 Chenhang He , Ruihuang Li , Shuai Li , Lei Zhang

Recent transformer based approaches have demonstrated impressive performance in solving real-world 3D human pose estimation problems. Albeit these approaches achieve fruitful results on benchmark datasets, they tend to fall short of sports…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Zhuoer Yin , Calvin Yeung , Tomohiro Suzuki , Ryota Tanaka , Keisuke Fujii

We present WidthFormer, a novel transformer-based module to compute Bird's-Eye-View (BEV) representations from multi-view cameras for real-time autonomous-driving applications. WidthFormer is computationally efficient, robust and does not…

Computer Vision and Pattern Recognition · Computer Science 2024-07-31 Chenhongyi Yang , Tianwei Lin , Lichao Huang , Elliot J. Crowley

Identifying unusual driving behaviors exhibited by drivers during driving is essential for understanding driver behavior and the underlying causes of crashes. Previous studies have primarily approached this problem as a classification task,…

Computer Vision and Pattern Recognition · Computer Science 2023-04-18 Armstrong Aboah , Ulas Bagci , Abdul Rashid Mussah , Neema Jakisa Owor , Yaw Adu-Gyamfi

Driving in a state of drowsiness is a major cause of road accidents, resulting in tremendous damage to life and property. Developing robust, automatic, real-time systems that can infer drowsiness states of drivers has the potential of…

Computer Vision and Pattern Recognition · Computer Science 2020-10-22 Ajjen Joshi , Survi Kyal , Sandipan Banerjee , Taniya Mishra

Holistic methods based on dense trajectories are currently the de facto standard for recognition of human activities in video. Whether holistic representations will sustain or will be superseded by higher level video encoding in terms of…

Computer Vision and Pattern Recognition · Computer Science 2014-07-29 Leonid Pishchulin , Mykhaylo Andriluka , Bernt Schiele
‹ Prev 1 4 5 6 7 8 10 Next ›