English
Related papers

Related papers: Video2Action: Reducing Human Interactions in Actio…

200 papers

We aim to tackle the interesting yet challenging problem of generating videos of diverse and natural human motions from prescribed action categories. The key issue lies in the ability to synthesize multiple distinct motion sequences that…

Computer Vision and Pattern Recognition · Computer Science 2021-12-21 Chuan Guo , Xinxin Zuo , Sen Wang , Xinshuang Liu , Shihao Zou , Minglun Gong , Li Cheng

Annotating object ground truth in videos is vital for several downstream tasks in robot perception and machine learning, such as for evaluating the performance of an object tracker or training an image-based object detector. The accuracy of…

Computer Vision and Pattern Recognition · Computer Science 2023-12-21 Eric Price , Aamir Ahmad

Locating specific segments within an instructional video is an efficient way to acquire guiding knowledge. Generally, the task of obtaining video segments for both verbal explanations and visual demonstrations is known as visual answer…

Computer Vision and Pattern Recognition · Computer Science 2025-04-24 Chang Zong , Bin Li , Shoujun Zhou , Jian Wan , Lei Zhang

We aim to automatically identify human action reasons in online videos. We focus on the widespread genre of lifestyle vlogs, in which people perform actions while verbally describing them. We introduce and make publicly available the WhyAct…

Computer Vision and Pattern Recognition · Computer Science 2021-09-10 Oana Ignat , Santiago Castro , Hanwen Miao , Weiji Li , Rada Mihalcea

Autonomous driving systems require huge amounts of data to train. Manual annotation of this data is time-consuming and prohibitively expensive since it involves human resources. Therefore, active learning emerged as an alternative to ease…

Computer Vision and Pattern Recognition · Computer Science 2019-09-02 Javad Zolfaghari Bengar , Abel Gonzalez-Garcia , Gabriel Villalonga , Bogdan Raducanu , Hamed H. Aghdam , Mikhail Mozerov , Antonio M. Lopez , Joost van de Weijer

Distilling knowledge from human demonstrations is a promising way for robots to learn and act. Existing methods, which often rely on coarsely-aligned video pairs, are typically constrained to learning global or task-level features. As a…

Robotics · Computer Science 2025-11-18 Sicheng Xie , Haidong Cao , Zejia Weng , Zhen Xing , Haoran Chen , Shiwei Shen , Jiaqi Leng , Zuxuan Wu , Yu-Gang Jiang

Video captioning has shown impressive progress in recent years. One key reason of the performance improvements made by existing methods lie in massive paired video-sentence data, but collecting such strong annotation, i.e., high-quality…

Computer Vision and Pattern Recognition · Computer Science 2020-09-03 Jingyi Hou , Yunde Jia , Xinxiao wu , Yayun Qi

Recent advances in multimodal large language models have driven growing interest in graphical user interface (GUI) agents, yet their generalization remains constrained by the scarcity of large-scale training data spanning diverse real-world…

Computation and Language · Computer Science 2026-05-15 Weimin Xiong , Shuhao Gu , Bowen Ye , Zihao Yue , Lei Li , Feifan Song , Sujian Li , Hao Tian

In this paper we consider the task of recognizing human actions in realistic video where human actions are dominated by irrelevant factors. We first study the benefits of removing non-action video segments, which are the ones that do not…

Computer Vision and Pattern Recognition · Computer Science 2016-04-25 Yang Wang , Minh Hoai

Human annotation is always considered as ground truth in video object tracking tasks. It is used in both training and evaluation purposes. Thus, ensuring its high quality is an important task for the success of trackers and evaluations…

Computer Vision and Pattern Recognition · Computer Science 2019-11-11 Yu Pang , Xinyi Li , Lin Yuan , Haibin Ling

We introduce a framework that predicts the goals behind observable human action in video. Motivated by evidence in developmental psychology, we leverage video of unintentional action to learn video representations of goals without direct…

Computer Vision and Pattern Recognition · Computer Science 2020-12-17 Dave Epstein , Carl Vondrick

With advances in data-driven machine learning research, a wide variety of prediction models have been proposed to capture spatio-temporal features for the analysis of video streams. Recognising actions and detecting action transitions…

Computer Vision and Pattern Recognition · Computer Science 2024-03-06 Harshala Gammulle , David Ahmedt-Aristizabal , Simon Denman , Lachlan Tychsen-Smith , Lars Petersson , Clinton Fookes

Retrieving target videos based on text descriptions is a task of great practical value and has received increasing attention over the past few years. Despite recent progress, imperfect annotations in existing video retrieval datasets have…

Computer Vision and Pattern Recognition · Computer Science 2022-07-22 Zeyu Wang , Yu Wu , Karthik Narasimhan , Olga Russakovsky

Instructional videos are a common source for learning text-video or even multimodal representations by leveraging subtitles extracted with automatic speech recognition systems (ASR) from the audio signal in the videos. However, in contrast…

Computer Vision and Pattern Recognition · Computer Science 2024-09-10 Nina Shvetsova , Anna Kukleva , Xudong Hong , Christian Rupprecht , Bernt Schiele , Hilde Kuehne

Users often take notes for instructional videos to access key knowledge later without revisiting long videos. Automated note generation tools enable users to obtain informative notes efficiently. However, notes generated by existing…

Human-Computer Interaction · Computer Science 2025-08-21 Running Zhao , Zhihan Jiang , Xinchen Zhang , Chirui Chang , Handi Chen , Weipeng Deng , Luyao Jin , Xiaojuan Qi , Xun Qian , Edith C. H. Ngai

Understanding human actions in visual data is tied to advances in complementary research areas including object recognition, human dynamics, domain adaptation and semantic segmentation. Over the last decade, human action analysis evolved…

Computer Vision and Pattern Recognition · Computer Science 2017-02-02 Samitha Herath , Mehrtash Harandi , Fatih Porikli

Anticipating future actions in videos is challenging, as the observed frames provide only evidence of past activities, requiring the inference of latent intentions to predict upcoming actions. Existing transformer-based approaches, which…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Tsung-Ming Tai , Sofia Casarin , Andrea Pilzer , Werner Nutt , Oswald Lanz

We propose technology to enable a new medium of expression, where video elements can be looped, merged, and triggered, interactively. Like audio, video is easy to sample from the real world but hard to segment into clean reusable elements.…

Human-Computer Interaction · Computer Science 2017-05-23 Corneliu Ilisescu , Halil Aytac Kanaci , Matteo Romagnoli , Neill D. F. Campbell , Gabriel J. Brostow

We consider the task of temporal human action localization in lifestyle vlogs. We introduce a novel dataset consisting of manual annotations of temporal localization for 13,000 narrated actions in 1,200 video clips. We present an extensive…

Computer Vision and Pattern Recognition · Computer Science 2022-02-22 Oana Ignat , Santiago Castro , Yuhang Zhou , Jiajun Bao , Dandan Shan , Rada Mihalcea

Action recognition is an important task for video understanding with broad applications. However, developing an effective action recognition solution often requires extensive engineering efforts in building and testing different…

Computer Vision and Pattern Recognition · Computer Science 2022-07-19 Daochen Zha , Zaid Pervaiz Bhat , Yi-Wei Chen , Yicheng Wang , Sirui Ding , Jiaben Chen , Kwei-Herng Lai , Mohammad Qazim Bhat , Anmoll Kumar Jain , Alfredo Costilla Reyes , Na Zou , Xia Hu
‹ Prev 1 4 5 6 7 8 10 Next ›