English
Related papers

Related papers: Is end-to-end learning enough for fitness activity…

200 papers

Recent advances in deep learning and computer vision offer an excellent opportunity to investigate high-level visual analysis tasks such as human localization and human pose estimation. Although the performance of human localization and…

Computer Vision and Pattern Recognition · Computer Science 2022-04-21 Somnuk Phon-Amnuaisuk , Ken T. Murata , La-Or Kovavisaruch , Tiong-Hoo Lim , Praphan Pavarangkoon , Takamichi Mizuhara

Action recognition and human pose estimation are closely related but both problems are generally handled as distinct tasks in the literature. In this work, we propose a multitask framework for jointly 2D and 3D pose estimation from still…

Computer Vision and Pattern Recognition · Computer Science 2018-03-22 Diogo C. Luvizon , David Picard , Hedi Tabia

Previous part-based attribute recognition approaches perform part detection and attribute recognition in separate steps. The parts are not optimized for attribute recognition and therefore could be sub-optimal. We present an end-to-end deep…

Computer Vision and Pattern Recognition · Computer Science 2016-07-20 Luwei Yang , Ligen Zhu , Yichen Wei , Shuang Liang , Ping Tan

Recent advances in incorporating neural networks into particle filters provide the desired flexibility to apply particle filters in large-scale real-world applications. The dynamic and measurement models in this framework are learnable…

Machine Learning · Computer Science 2021-03-30 Hao Wen , Xiongjie Chen , Georgios Papagiannis , Conghui Hu , Yunpeng Li

We present a unified framework for understanding human social behaviors in raw image sequences. Our model jointly detects multiple individuals, infers their social actions, and estimates the collective actions with a single feed-forward…

Computer Vision and Pattern Recognition · Computer Science 2016-11-29 Timur Bagautdinov , Alexandre Alahi , François Fleuret , Pascal Fua , Silvio Savarese

Learning approaches have shown great success in the task of super-resolving an image given a low resolution input. Video super-resolution aims for exploiting additionally the information from multiple images. Typically, the images are…

Computer Vision and Pattern Recognition · Computer Science 2017-07-04 Osama Makansi , Eddy Ilg , Thomas Brox

One of the core components of conventional (i.e., non-learned) video codecs consists of predicting a frame from a previously-decoded frame, by leveraging temporal correlations. In this paper, we propose an end-to-end learned system for…

Image and Video Processing · Electrical Eng. & Systems 2020-04-22 Nannan Zou , Honglei Zhang , Francesco Cricri , Hamed R. Tavakoli , Jani Lainema , Emre Aksu , Miska Hannuksela , Esa Rahtu

While learning based compression techniques for images have outperformed traditional methods, they have not been widely adopted in machine learning pipelines. This is largely due to lack of standardization and lack of retention of salient…

Image and Video Processing · Electrical Eng. & Systems 2024-10-01 Kartik Gupta , Kimberley Faria , Vikas Mehta

Existing multi-person video pose estimation methods typically adopt a two-stage pipeline: detecting individuals in each frame, followed by temporal modeling for single person pose estimation. This design relies on heuristic operations such…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Yonghui Yu , Jiahang Cai , Xun Wang , Wenwu Yang

In this work we introduce a fully end-to-end approach for action detection in videos that learns to directly predict the temporal bounds of actions. Our intuition is that the process of detecting actions is naturally one of observation and…

Computer Vision and Pattern Recognition · Computer Science 2017-03-14 Serena Yeung , Olga Russakovsky , Greg Mori , Li Fei-Fei

The most performant spatio-temporal action localisation models use external person proposals and complex external memory banks. We propose a fully end-to-end, purely-transformer based model that directly ingests an input video, and outputs…

Computer Vision and Pattern Recognition · Computer Science 2023-04-25 Alexey Gritsenko , Xuehan Xiong , Josip Djolonga , Mostafa Dehghani , Chen Sun , Mario Lučić , Cordelia Schmid , Anurag Arnab

Holistic methods based on dense trajectories are currently the de facto standard for recognition of human activities in video. Whether holistic representations will sustain or will be superseded by higher level video encoding in terms of…

Computer Vision and Pattern Recognition · Computer Science 2014-07-29 Leonid Pishchulin , Mykhaylo Andriluka , Bernt Schiele

Plenty of effective methods have been proposed for face recognition during the past decade. Although these methods differ essentially in many aspects, a common practice of them is to specifically align the facial area based on the prior…

Computer Vision and Pattern Recognition · Computer Science 2017-08-02 Yuanyi Zhong , Jiansheng Chen , Bo Huang

Person re-identification across disjoint camera views has been widely applied in video surveillance yet it is still a challenging problem. One of the major challenges lies in the lack of spatial and temporal cues, which makes it difficult…

Computer Vision and Pattern Recognition · Computer Science 2017-06-28 Hao Liu , Jiashi Feng , Meibin Qi , Jianguo Jiang , Shuicheng Yan

In recent years, considerable progress has been made towards a vehicle's ability to operate autonomously. An end-to-end approach attempts to achieve autonomous driving using a single, comprehensive software component. Recent breakthroughs…

Robotics · Computer Science 2019-05-17 Hege Haavaldsen , Max Aasboe , Frank Lindseth

The need for automated real-time visual systems in applications such as smart camera surveillance, smart environments, and drones necessitates the improvement of methods for visual active monitoring and control. Traditionally, the active…

Computer Vision and Pattern Recognition · Computer Science 2021-07-29 Christos Kyrkou

Event cameras are vision sensors that record asynchronous streams of per-pixel brightness changes, referred to as "events". They have appealing advantages over frame-based cameras for computer vision, including high temporal resolution,…

Computer Vision and Pattern Recognition · Computer Science 2019-08-21 Daniel Gehrig , Antonio Loquercio , Konstantinos G. Derpanis , Davide Scaramuzza

We introduce associative embedding, a novel method for supervising convolutional neural networks for the task of detection and grouping. A number of computer vision problems can be framed in this manner including multi-person pose…

Computer Vision and Pattern Recognition · Computer Science 2017-06-12 Alejandro Newell , Zhiao Huang , Jia Deng

We have seen significant leapfrog advancement in machine learning in recent decades. The central idea of machine learnability lies on constructing learning algorithms that learn from good data. The availability of more data being made…

Computer Vision and Pattern Recognition · Computer Science 2020-08-07 Ng Hui Xian Lynnette , Henry Ng Siong Hock , Nguwi Yok Yen

Predicting high-fidelity future human poses, from a historically observed sequence, is decisive for intelligent robots to interact with humans. Deep end-to-end learning approaches, which typically train a generic pre-trained model on…

Computer Vision and Pattern Recognition · Computer Science 2023-04-14 Qiongjie Cui , Huaijiang Sun , Jianfeng Lu , Bin Li , Weiqing Li