中文
相关论文

相关论文: ActionCOMET: A Zero-shot Approach to Learn Image-s…

200 篇论文

Human activity recognition is typically addressed by detecting key concepts like global and local motion, features related to object classes present in the scene, as well as features related to the global context. The next open challenges…

计算机视觉与模式识别 · 计算机科学 2018-09-21 Fabien Baradel , Natalia Neverova , Christian Wolf , Julien Mille , Greg Mori

Understanding human actions is a key problem in computer vision. However, recognizing actions is only the first step of understanding what a person is doing. In this paper, we introduce the problem of predicting why a person has performed…

计算机视觉与模式识别 · 计算机科学 2016-12-01 Carl Vondrick , Deniz Oktay , Hamed Pirsiavash , Antonio Torralba

Neural networks trained on datasets such as ImageNet have led to major advances in visual object classification. One obstacle that prevents networks from reasoning more deeply about complex scenes and situations, and from integrating visual…

Current captioning datasets focus on object-centric captions, describing the visible objects in the image, e.g. "people eating food in a park". Although these datasets are useful to evaluate the ability of Vision & Language models to…

计算与语言 · 计算机科学 2023-09-26 Michele Cafagna , Kees van Deemter , Albert Gatt

Spatio-temporal action detection encompasses the tasks of localizing and classifying individual actions within a video. Recent works aim to enhance this process by incorporating interaction modeling, which captures the relationship between…

计算机视觉与模式识别 · 计算机科学 2024-12-06 Wei-Jhe Huang , Min-Hung Chen , Shang-Hong Lai

This work introduces a model that can recognize objects in images even if no training data is available for the objects. The only necessary knowledge about the unseen categories comes from unsupervised large text corpora. In our zero-shot…

计算机视觉与模式识别 · 计算机科学 2013-03-21 Richard Socher , Milind Ganjoo , Hamsa Sridhar , Osbert Bastani , Christopher D. Manning , Andrew Y. Ng

An image editing model should be able to perform diverse edits, ranging from object replacement, changing attributes or style, to performing actions or movement, which require many forms of reasoning. Current general instruction-guided…

计算机视觉与模式识别 · 计算机科学 2024-10-18 Benno Krojer , Dheeraj Vattikonda , Luis Lara , Varun Jampani , Eva Portelance , Christopher Pal , Siva Reddy

Procedural video understanding is gaining attention in the vision and language community. Deep learning-based video analysis requires extensive data. Consequently, existing works often use web videos as training resources, making it…

计算机视觉与模式识别 · 计算机科学 2024-08-06 Koki Maeda , Tosho Hirasawa , Atsushi Hashimoto , Jun Harashima , Leszek Rybicki , Yusuke Fukasawa , Yoshitaka Ushiku

Most human behaviors consist of multiple parts, steps, or subtasks. These structures guide our action planning and execution, but when we observe others, the latent structure of their actions is typically unobservable, and must be inferred…

人工智能 · 计算机科学 2018-09-28 Ryo Nakahashi , Chris L. Baker , Joshua B. Tenenbaum

Multimodal systems have great potential to assist humans in procedural activities, where people follow instructions to achieve their goals. Despite diverse application scenarios, systems are typically evaluated on traditional classification…

To support humans in their daily lives, robots are required to autonomously learn, adapt to objects and environments, and perform the appropriate actions. We tackled on the task of cooking scrambled eggs using real ingredients, in which the…

机器人学 · 计算机科学 2025-09-18 Namiko Saito , Mayu Tatsumi , Ayuna Kubo , Kanata Suzuki , Hiroshi Ito , Shigeki Sugano , Tetsuya Ogata

Computer vision algorithms performance are near or superior to humans in the visual problems including object recognition (especially those of fine-grained categories), segmentation, and 3D object reconstruction from 2D views. Humans are,…

计算机视觉与模式识别 · 计算机科学 2020-11-13 Stuart Synakowski , Qianli Feng , Aleix Martinez

Humans routinely infer taste, smell, texture, and even sound from food images a phenomenon well studied in cognitive science. However, prior vision language research on food has focused primarily on recognition tasks such as meal…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Sabab Ishraq , Aarushi Aarushi , Juncai Jiang , Chen Chen

Unlike most reinforcement learning agents which require an unrealistic amount of environment interactions to learn a new behaviour, humans excel at learning quickly by merely observing and imitating others. This ability highly depends on…

机器学习 · 计算机科学 2023-12-05 Xingyuan Zhang , Philip Becker-Ehmck , Patrick van der Smagt , Maximilian Karl

We propose a method for human action recognition, one that can localize the spatiotemporal regions that `define' the actions. This is a challenging task due to the subtlety of human actions in video and the co-occurrence of contextual…

计算机视觉与模式识别 · 计算机科学 2019-04-12 Yang Wang , Vinh Tran , Gedas Bertasius , Lorenzo Torresani , Minh Hoai

Articulation modeling enables robots to learn joint parameters of articulated objects for effective manipulation which can then be used downstream for skill learning or planning. Existing approaches often rely on prior knowledge about the…

机器人学 · 计算机科学 2026-02-04 Anmol Gupta , Weiwei Gu , Omkar Patil , Jun Ki Lee , Nakul Gopalan

A major challenge for modern AI is to learn to understand the world and learn to act largely by observation. This paper explores a self-supervised approach that combines internet-scale video data with a small amount of interaction data…

We introduce the novel concept of visually Connecting Actions and Their Effects (CATE) in video understanding. CATE can have applications in areas like task planning and learning from demonstration. We identify and explore two different…

计算机视觉与模式识别 · 计算机科学 2024-07-29 Paritosh Parmar , Eric Peh , Basura Fernando

We introduce the Action Transformer model for recognizing and localizing human actions in video clips. We repurpose a Transformer-style architecture to aggregate features from the spatiotemporal context around the person whose actions we…

计算机视觉与模式识别 · 计算机科学 2019-05-20 Rohit Girdhar , João Carreira , Carl Doersch , Andrew Zisserman

Human action recognition refers to automatic recognizing human actions from a video clip. In reality, there often exist multiple human actions in a video stream. Such a video stream is often weakly-annotated with a set of relevant human…

计算机视觉与模式识别 · 计算机科学 2019-02-07 Qian Wang , Ke Chen