English
Related papers

Related papers: Why and How Auxiliary Tasks Improve JEPA Represent…

200 papers

World models require robust relational understanding to support prediction, reasoning, and control. While object-centric representations provide a useful abstraction, they are not sufficient to capture interaction-dependent dynamics. We…

Artificial Intelligence · Computer Science 2026-05-29 Heejeong Nam , Quentin Le Lidec , Lucas Maes , Yann LeCun , Randall Balestriero

We introduce VL-JEPA, a vision-language model built on a Joint Embedding Predictive Architecture (JEPA). Instead of autoregressively generating tokens as in classical VLMs, VL-JEPA predicts continuous embeddings of the target texts. By…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Delong Chen , Mustafa Shukor , Theo Moutakanni , Willy Chung , Jade Yu , Tejaswi Kasarla , Yejin Bang , Allen Bolourchi , Yann LeCun , Pascale Fung

Self-supervision is often used for pre-training to foster performance on a downstream task by constructing meaningful representations of samples. Self-supervised learning (SSL) generally involves generating different views of the same…

Machine Learning · Computer Science 2025-05-06 Hugo Thimonier , José Lucas De Melo Costa , Fabrice Popineau , Arpad Rimmel , Bich-Liên Doan

World models compress rich sensory streams into compact latent codes that anticipate future observations. We let separate agents acquire such models from distinct viewpoints of the same environment without any parameter sharing or…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Haoran Zhang , Youjin Wang , Yi Duan , Rong Fu , Dianyu Zhao , Sicheng Fan , Shuaishuai Cao , Wentao Guo , Xiao Zhou

Recent vision-language-action (VLA) models built upon pretrained vision-language models (VLMs) have achieved significant improvements in robotic manipulation. However, current VLAs still suffer from low sample efficiency and limited…

Computer Vision and Pattern Recognition · Computer Science 2026-02-13 Shangchen Miao , Ningya Feng , Jialong Wu , Ye Lin , Xu He , Dong Li , Mingsheng Long

Inspired by the success of generative pretraining in natural language, we ask whether the same principles can yield strong self-supervised visual learners. Instead of training models to output features for downstream use, we train them to…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Sihan Xu , Ziqiao Ma , Wenhao Chai , Xuweiyi Chen , Weiyang Jin , Joyce Chai , Saining Xie , Stella X. Yu

Joint-Embedding Predictive Architectures (JEPA) learn view-invariant representations and admit projection-based distribution matching for collapse prevention. Existing approaches regularize representations towards isotropic Gaussian…

Machine Learning · Computer Science 2026-05-29 Yilun Kuang , Yash Dagade , Tim G. J. Rudner , Randall Balestriero , Yann LeCun

We evaluate JEPA-style predictive representation learning versus reconstruction-based autoencoders on a controlled "TV-series" linear dynamical system with known latent state and a single noise parameter. While an initial comparison…

Machine Learning · Computer Science 2026-03-17 Alexey Potapov , Oleg Shcherbakov , Ivan Kravchenko

The cornerstone of cognitive intelligence lies in extracting hidden patterns from observations and leveraging these principles to systematically predict future outcomes. However, current image tokenization methods demonstrate significant…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Junyeob Baek , Hosung Lee , Christopher Hoang , Mengye Ren , Sungjin Ahn

This paper explores the automated process of determining stem compatibility by identifying audio recordings of single instruments that blend well with a given musical context. To tackle this challenge, we present Stem-JEPA, a novel…

Sound · Computer Science 2024-08-06 Alain Riou , Stefan Lattner , Gaëtan Hadjeres , Michael Anslow , Geoffroy Peeters

Video Joint Embedding Predictive Architectures (V-JEPA) learn generalizable off-the-shelf video representation by predicting masked regions in latent space with an exponential moving average (EMA)-updated teacher. While EMA prevents…

Machine Learning · Computer Science 2025-09-30 Xianhang Li , Chen Huang , Chun-Liang Li , Eran Malach , Josh Susskind , Vimal Thilak , Etai Littwin

Large Language Model (LLM) pretraining, finetuning, and evaluation rely on input-space reconstruction and generative capabilities. Yet, it has been observed in vision that embedding-space training objectives, e.g., with Joint Embedding…

Computation and Language · Computer Science 2025-10-08 Hai Huang , Yann LeCun , Randall Balestriero

The joint-embedding predictive architecture (JEPA) recently has shown impressive results in extracting visual representations from unlabeled imagery under a masking strategy. However, we reveal its disadvantages, notably its insufficient…

Computer Vision and Pattern Recognition · Computer Science 2024-05-29 Shentong Mo , Sukmin Yun

We present a transformer architecture-based foundation model for tasks at high-energy particle colliders such as the Large Hadron Collider. We train the model to classify jets using a self-supervised strategy inspired by the Joint Embedding…

Machine Learning · Computer Science 2025-02-07 Jai Bardhan , Radhikesh Agrawal , Abhiram Tilak , Cyrin Neeraj , Subhadip Mitra

Reservoir simulation workflows face a fundamental data asymmetry: input parameter fields (geostatistical permeability realizations, porosity distributions) are free to generate in arbitrary quantities, yet existing neural operator…

Machine Learning · Computer Science 2026-04-10 Brandon Yee , Pairie Koh

Recent advancements in self-supervised learning in the point cloud domain have demonstrated significant potential. However, these methods often suffer from drawbacks, including lengthy pre-training time, the necessity of reconstruction in…

Computer Vision and Pattern Recognition · Computer Science 2025-02-11 Ayumu Saito , Prachi Kudeshia , Jiju Poovvancheri

We introduce Brain-JEPA, a brain dynamics foundation model with the Joint-Embedding Predictive Architecture (JEPA). This pioneering model achieves state-of-the-art performance in demographic prediction, disease diagnosis/prognosis, and…

Joint-Embedding Predictive Architectures (JEPAs) aim to learn representations by predicting target embeddings from context embeddings, inducing a scalar compatibility energy in a latent space. In contrast, Quasimetric Reinforcement Learning…

Machine Learning · Computer Science 2026-02-13 Anthony Kobanda , Waris Radji

Exploring spatial-temporal dependencies from observed motions is one of the core challenges of human motion prediction. Previous methods mainly focus on dedicated network structures to model the spatial and temporal dependencies. This paper…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Chenxin Xu , Robby T. Tan , Yuhong Tan , Siheng Chen , Xinchao Wang , Yanfeng Wang

With the advent of Joint Embedding Predictive Architectures (JEPAs), which appear to be more capable than reconstruction-based methods, this paper introduces a novel technique for creating world models using continuous-time dynamic systems…

Machine Learning · Computer Science 2025-08-15 Jonas Ulmen , Ganesh Sundaram , Daniel Görges