English
Related papers

Related papers: Intention Action Anticipation Model with Guide-Fee…

200 papers

Monocular 3D human performance capture is indispensable for many applications in computer graphics and vision for enabling immersive experiences. However, detailed capture of humans requires tracking of multiple aspects, including the…

Computer Vision and Pattern Recognition · Computer Science 2022-10-12 Yue Jiang , Marc Habermann , Vladislav Golyanik , Christian Theobalt

Generative AI has significantly changed industries by enabling text-driven image generation, yet challenges remain in achieving high-resolution outputs that align with fine-grained user preferences. Consequently, multi-round interactions…

Computer Vision and Pattern Recognition · Computer Science 2025-04-28 Kun Li , Jianhui Wang , Yangfan He , Xinyuan Song , Ruoyu Wang , Hongyang He , Wenxin Zhang , Jiaqi Chen , Keqin Li , Sida Li , Miao Zhang , Tianyu Shi , Xueqian Wang

Generative models have shown great promise in collaborative filtering by capturing the underlying distribution of user interests and preferences. However, existing approaches struggle with inaccurate posterior approximations and…

Information Retrieval · Computer Science 2025-09-08 Chengkai Liu , Yangtian Zhang , Jianling Wang , Rex Ying , James Caverlee

Complex physical tasks entail a sequence of object interactions, each with its own preconditions -- which can be difficult for robotic agents to learn efficiently solely through their own experience. We introduce an approach to discover…

Computer Vision and Pattern Recognition · Computer Science 2021-10-18 Tushar Nagarajan , Kristen Grauman

We propose GHR-VQA, Graph-guided Hierarchical Relational Reasoning for Video Question Answering (Video QA), a novel human-centric framework that incorporates scene graphs to capture intricate human-object interactions within video…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Dionysia Danai Brilli , Dimitrios Mallis , Vassilis Pitsikalis , Petros Maragos

Click-Through Rate (CTR) prediction has long been dominated by discriminative paradigms that optimize local decision boundaries within candidate-specific subspaces. However, these models often fail to capture the global joint distribution…

Information Retrieval · Computer Science 2026-04-15 Chen Gao , Zixin Zhao , Lv Shao , Tong Liu

We present an approach to learn an object-centric forward model, and show that this allows us to plan for sequences of actions to achieve distant desired goals. We propose to model a scene as a collection of objects, each with an explicit…

Computer Vision and Pattern Recognition · Computer Science 2019-10-09 Yufei Ye , Dhiraj Gandhi , Abhinav Gupta , Shubham Tulsiani

As autonomous machines such as robots and vehicles start performing tasks involving human users, ensuring a safe interaction between them becomes an important issue. Translating methods from human-robot interaction (HRI) studies to the…

Robotics · Computer Science 2021-06-04 Erwin Jose Lopez Pulgarin , Guido Herrmann , Ute Leonards

Generative AI shifts interaction toward intent-based outcome specification, despite user intents being inherently vague, fluid, and evolving. While a growing body of HCI research has proposed diverse interaction techniques to support this…

Human-Computer Interaction · Computer Science 2026-01-29 Yoonsu Kim , Kihoon Son , Seoyoung Kim , Brandon Chin , Juho Kim

Medical Vision-Language Models (Med-VLMs) have achieved success across various tasks, yet most existing methods overlook the modality misalignment issue that can lead to untrustworthy responses in clinical settings. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Songtao Jiang , Yan Zhang , Yeying Jin , Zhihang Tang , Yangyang Wu , Yang Feng , Jian Wu , Zuozhu Liu

We propose Anticipative Video Transformer (AVT), an end-to-end attention-based video modeling architecture that attends to the previously observed video in order to anticipate future actions. We train the model jointly to predict the next…

Computer Vision and Pattern Recognition · Computer Science 2021-09-23 Rohit Girdhar , Kristen Grauman

Hindsight rationality is an approach to playing general-sum games that prescribes no-regret learning dynamics for individual agents with respect to a set of deviations, and further describes jointly rational behavior among multiple agents…

Computer Science and Game Theory · Computer Science 2022-06-24 Dustin Morrill , Ryan D'Orazio , Marc Lanctot , James R. Wright , Michael Bowling , Amy Greenwald

Group Activity Scene Graph (GASG) generation is a challenging task in computer vision, aiming to anticipate and describe relationships between subjects and objects in video sequences. Traditional Video Scene Graph Generation (VidSGG)…

Computer Vision and Pattern Recognition · Computer Science 2024-10-23 Naga VS Raviteja Chappa , Pha Nguyen , Thi Hoang Ngan Le , Khoa Luu

Technological progress increasingly envisions the use of robots interacting with people in everyday life. Human-robot collaboration (HRC) is the approach that explores the interaction between a human and a robot, during the completion of a…

Robotics · Computer Science 2022-07-12 Francesco Semeraro , Alexander Griffiths , Angelo Cangelosi

Being able to predict human gaze behavior has obvious importance for behavioral vision and for computer vision applications. Most models have mainly focused on predicting free-viewing behavior using saliency maps, but these predictions do…

Computer Vision and Pattern Recognition · Computer Science 2020-06-26 Zhibo Yang , Lihan Huang , Yupei Chen , Zijun Wei , Seoyoung Ahn , Gregory Zelinsky , Dimitris Samaras , Minh Hoai

Early accident anticipation from dashcam videos is a highly desirable yet challenging task for improving the safety of intelligent vehicles. Existing advanced accident anticipation approaches commonly model the interaction among traffic…

Computer Vision and Pattern Recognition · Computer Science 2025-02-27 Hongpu Huang , Wei Zhou , Chen Wang

Recognizing human activities in videos is challenging due to the spatio-temporal complexity and context-dependence of human interactions. Prior studies often rely on single input modalities, such as RGB or skeletal data, limiting their…

Computer Vision and Pattern Recognition · Computer Science 2024-09-05 Tuyen Tran , Thao Minh Le , Hung Tran , Truyen Tran

Advancements in egocentric video datasets like Ego4D, EPIC-Kitchens, and Ego-Exo4D have enriched the study of first-person human interactions, which is crucial for applications in augmented reality and assisted living. Despite these…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Joungbin An , Yunsu Park , Hyolim Kang , Seon Joo Kim

This work studies the problem of predicting the sequence of future actions for surround vehicles in real-world driving scenarios. To this aim, we make three main contributions. The first contribution is an automatic method to convert the…

Computer Vision and Pattern Recognition · Computer Science 2020-04-30 Jan-Nico Zaech , Dengxin Dai , Alexander Liniger , Luc Van Gool

Video-based person re-identification (Re-ID) aims to automatically retrieve video sequences of the same person under non-overlapping cameras. To achieve this goal, it is the key to fully utilize abundant spatial and temporal cues in videos.…

Computer Vision and Pattern Recognition · Computer Science 2021-04-02 Xuehu Liu , Pingping Zhang , Chenyang Yu , Huchuan Lu , Xiaoyun Yang
‹ Prev 1 8 9 10 Next ›