English
Related papers

Related papers: Entity-Centric World Models: Interaction-Aware Mas…

200 papers

Genomic Foundation Models (GFMs) typically rely on Masked Language Modeling (MLM) or Next-Token Prediction (NTP) to learn the "Laws of Nature". While effective at capturing local syntax, these generative paradigms prioritize token-level…

We propose JFAA, a JEPA-based Future Action Anticipation method for the EPIC-KITCHENS-100 (EK-100) Action Anticipation task. Inspired by the representation learning and future prediction ability of V-JEPA 2.1, JFAA uses a frozen encoder and…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Qiaohui Chu , Haoyu Zhang , Yisen Feng , Meng Liu , Weili Guan , Dongmei Jiang , Liqiang Nie

A long-standing challenge in scene analysis is the recovery of scene arrangements under moderate to heavy occlusion, directly from monocular video. While the problem remains a subject of active research, concurrent advances have been made…

Graphics · Computer Science 2019-07-19 Aron Monszpart , Paul Guerrero , Duygu Ceylan , Ersin Yumer , Niloy J. Mitra

Masked Video Autoencoder (MVA) approaches have demonstrated their potential by significantly outperforming previous video representation learning methods. However, they waste an excessive amount of computations and memory in predicting…

Computer Vision and Pattern Recognition · Computer Science 2024-06-21 Sunil Hwang , Jaehong Yoon , Youngwan Lee , Sung Ju Hwang

Current multimodal learning strategies primarily optimize in the original token space. Such a framework is easy to incorporate with the backbone of pretrained language model, but might result in modality collapse. To alleviate such issues,…

Machine Learning · Computer Science 2025-06-19 Hongyang Lei , Xiaolong Cheng , Qi Qin , Dan Wang , Kun Fan , Huazhen Huang , Qingqing Gu , Yetao Wu , Zhonglin Jiang , Yong Chen , Luo Ji

Self-Supervised Learning (SSL) has shifted from pixel-level reconstruction to latent space prediction, spearheaded by the Joint Embedding Predictive Architecture (JEPA). While effective, standard JEPA models typically rely on a…

Machine Learning · Computer Science 2026-03-03 Yongchao Huang

Pretrained video diffusion models provide powerful spatiotemporal generative priors, making them a natural foundation for robotic world models. While recent world-action models jointly optimize future videos and actions, they predominantly…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Zhaoyang Yang , Yurun Jin , Lizhe Qi , Cong Huang , Kai Chen

Eye movements have long been studied as a window into the attentional mechanisms of the human brain and made accessible as novelty style human-machine interfaces. However, not everything that we gaze upon, is something we want to interact…

Computer Vision and Pattern Recognition · Computer Science 2022-01-25 Paul Festor , Ali Shafti , Alex Harston , Michey Li , Pavel Orlov , A. Aldo Faisal

We are interested in learning scalable agents for reinforcement learning that can learn from large-scale, diverse sequential data similar to current large vision and language models. To this end, this paper presents masked decision…

Machine Learning · Computer Science 2023-05-30 Fangchen Liu , Hao Liu , Aditya Grover , Pieter Abbeel

Manipulation tasks require robots to reason about cause and effect when interacting with objects. Yet, many data-driven approaches lack causal semantics and thus only consider correlations. We introduce COBRA-PPM, a novel causal Bayesian…

Robotics · Computer Science 2025-09-01 Ricardo Cannizzaro , Michael Groom , Jonathan Routley , Robert Osazuwa Ness , Lars Kunze

Joint visual attention (JVA) provides informative cues on human behavior during social interactions. The ubiquity of egocentric eye-trackers and large-scale datasets on everyday interactions offer research opportunities in identifying JVA…

Human-Computer Interaction · Computer Science 2025-09-17 Kumushini Thennakoon , Yasasi Abeysinghe , Bhanuka Mahanama , Vikas Ashok , Sampath Jayarathna

Trajectory prediction has been a crucial task in building a reliable autonomous driving system by anticipating possible dangers. One key issue is to generate consistent trajectory predictions without colliding. To overcome the challenge, we…

Computer Vision and Pattern Recognition · Computer Science 2023-03-14 Hao Chen , Jiaze Wang , Kun Shao , Furui Liu , Jianye Hao , Chenyong Guan , Guangyong Chen , Pheng-Ann Heng

Expressing and identifying emotions through facial and physical expressions is a significant part of social interaction. Emotion recognition is an essential task in computer vision due to its various applications and mainly for allowing a…

Computer Vision and Pattern Recognition · Computer Science 2021-12-15 Willams Costa , David Macêdo , Cleber Zanchettin , Lucas S. Figueiredo , Veronica Teichrieb

The interactions between human and objects are important for recognizing object-centric actions. Existing methods usually adopt a two-stage pipeline, where object proposals are first detected using a pretrained detector, and then are fed to…

Computer Vision and Pattern Recognition · Computer Science 2024-04-19 Xunsong Li , Pengzhan Sun , Yangcen Liu , Lixin Duan , Wen Li

For efficient human-agent interaction, an agent should proactively recognize their target user and prepare for upcoming interactions. We formulate this challenging problem as the novel task of jointly forecasting a person's intent to…

Computer Vision and Pattern Recognition · Computer Science 2025-05-09 Tongfei Bian , Yiming Ma , Mathieu Chollet , Victor Sanchez , Tanaya Guha

Knowledge bases, and their representations in the form of knowledge graphs (KGs), are naturally incomplete. Since scientific and industrial applications have extensively adopted them, there is a high demand for solutions that complete their…

Artificial Intelligence · Computer Science 2025-07-30 Vítor Lourenço , Aline Paes

Masked Autoencoder (MAE) is a self-supervised approach for representation learning, widely applicable to a variety of downstream tasks in computer vision. In spite of its success, it is still not fully uncovered what and how MAE exactly…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Jeongwoo Shin , Inseo Lee , Junho Lee , Joonseok Lee

Reliable perception is essential for robots that interact with the world. But sensors alone are often insufficient to provide this capability, and they are prone to errors due to various conditions in the environment. Furthermore, there is…

Robotics · Computer Science 2021-07-08 Ying Siu Liang , Dongkyu Choi , Kenneth Kwok

Medical AI systems face catastrophic forgetting when deployed in clinical settings, where models must learn new imaging protocols while retaining prior diagnostic capabilities. This challenge is particularly acute for medical…

Multimedia · Computer Science 2025-12-23 Ziyuan Gao , Philippe Morel

To apply eyeshadow without a brush, should I use a cotton swab or a toothpick? Questions requiring this kind of physical commonsense pose a challenge to today's natural language understanding systems. While recent pretrained models (such as…

Computation and Language · Computer Science 2019-11-27 Yonatan Bisk , Rowan Zellers , Ronan Le Bras , Jianfeng Gao , Yejin Choi