English
Related papers

Related papers: VJEPA: Variational Joint Embedding Predictive Arch…

200 papers

This paper presents that the masked-modeling principle driving the success of large foundational vision models can be effectively applied to audio by making predictions in a latent space. We introduce Audio-based Joint-Embedding Predictive…

Sound · Computer Science 2024-01-12 Zhengcong Fei , Mingyuan Fan , Junshi Huang

Single-cell foundation models learn by reconstructing masked gene expression, implicitly treating technical noise as signal. With dropout rates exceeding 90%, reconstruction objectives encourage models to encode measurement artifacts rather…

Computational Engineering, Finance, and Science · Computer Science 2026-02-03 Ali ElSheikh , Rui-Xi Wang , Weimin Wu , Yibo Wen , Payam Dibaeinia , Jennifer Yuntong Zhang , Jerry Yao-Chieh Hu , Mei Knudson , Sudarshan Babu , Shao-Hua Sun , Aly A. Khan , Han Liu

Two competing paradigms exist for self-supervised learning of data representations. Joint Embedding Predictive Architecture (JEPA) is a class of architectures in which semantically similar inputs are encoded into representations that are…

Machine Learning · Computer Science 2024-07-08 Etai Littwin , Omid Saremi , Madhu Advani , Vimal Thilak , Preetum Nakkiran , Chen Huang , Joshua Susskind

Learning audio representations from raw waveforms overcomes key limitations of spectrogram-based audio representation learning, such as the long latency of spectrogram computation and the loss of phase information. Yet, while…

Accurately modeling and controlling vehicle exhaust emissions during transient events, such as rapid acceleration, is critical for meeting environmental regulations and optimizing powertrains. Conventional data-driven methods, such as…

Systems and Control · Electrical Eng. & Systems 2026-01-28 Ganesh Sundaram , Tobias Gehra , Jonas Ulmen , Mirjan Heubaum , Daniel Görges , Michael Günthner

Ultrasound (US) imaging poses unique challenges for representation learning due to its inherently noisy acquisition process. The low signal-to-noise ratio and stochastic speckle patterns hinder standard self-supervised learning methods…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Ashwath Radhachandran , Vedrana Ivezić , Shreeram Athreya , Ronit Anilkumar , Corey W. Arnold , William Speier

Image-based Joint-Embedding Predictive Architecture (IJEPA) offers an attractive alternative to Masked Autoencoder (MAE) for representation learning using the Masked Image Modeling framework. IJEPA drives representations to capture useful…

Machine Learning · Computer Science 2024-10-15 Etai Littwin , Vimal Thilak , Anand Gopalakrishnan

Pretraining Vision-Language-Action (VLA) policies on internet-scale video is appealing, yet current latent-action objectives often learn the wrong thing: they remain anchored to pixel variation rather than action-relevant state transitions,…

Robotics · Computer Science 2026-02-17 Jingwen Sun , Wenyao Zhang , Zekun Qi , Shaojie Ren , Zezhi Liu , Hanxin Zhu , Guangzhong Sun , Xin Jin , Zhibo Chen

Recent progress in latent world models (e.g., V-JEPA2) has shown promising capability in forecasting future world states from video observations. Nevertheless, dense prediction from a short observation window limits temporal context and can…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Haichao Zhang , Yijiang Li , Shwai He , Tushar Nagarajan , Mingfei Chen , Jianglin Lu , Ang Li , Yun Fu

Joint-Embedding Predictive Architecture (JEPA) is increasingly used for visual representation learning and as a component in model-based RL, but its behavior remains poorly understood. We provide a theoretical characterization of a simple,…

Machine Learning · Computer Science 2025-10-21 Jiacan Yu , Siyi Chen , Mingrui Liu , Nono Horiuchi , Vladimir Braverman , Zicheng Xu , Dan Haramati , Randall Balestriero

Recent vision-language-action (VLA) models built upon pretrained vision-language models (VLMs) have achieved significant improvements in robotic manipulation. However, current VLAs still suffer from low sample efficiency and limited…

Computer Vision and Pattern Recognition · Computer Science 2026-02-13 Shangchen Miao , Ningya Feng , Jialong Wu , Ye Lin , Xu He , Dong Li , Mingsheng Long

Language representation learning has emerged as a promising approach for sequential recommendation, thanks to its ability to learn generalizable representations. However, despite its advantages, this approach still struggles with data…

Information Retrieval · Computer Science 2025-08-08 Minh-Anh Nguyen , Dung D. Le

World models compress rich sensory streams into compact latent codes that anticipate future observations. We let separate agents acquire such models from distinct viewpoints of the same environment without any parameter sharing or…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Haoran Zhang , Youjin Wang , Yi Duan , Rong Fu , Dianyu Zhao , Sicheng Fan , Shuaishuai Cao , Wentao Guo , Xiao Zhou

Joint Embedding Predictive Architectures (JEPAs) offer a compelling framework for learning world models in compact latent spaces, yet existing methods remain fragile, relying on complex multi-term losses, exponential moving averages,…

Machine Learning · Computer Science 2026-03-26 Lucas Maes , Quentin Le Lidec , Damien Scieur , Yann LeCun , Randall Balestriero

This paper addresses the problem of self-supervised general-purpose audio representation learning. We explore the use of Joint-Embedding Predictive Architectures (JEPA) for this task, which consists of splitting an input mel-spectrogram…

Sound · Computer Science 2024-05-15 Alain Riou , Stefan Lattner , Gaëtan Hadjeres , Geoffroy Peeters

We introduce Variational Joint Embedding (VJE), a reconstruction-free latent-variable framework for non-contrastive self-supervised learning in representation space. VJE maximizes a symmetric conditional evidence lower bound (ELBO) on…

Machine Learning · Computer Science 2026-04-27 Amin Oji , Paul Fieguth

We present V-JEPA 2.1, a family of self-supervised models that learn dense, high-quality visual representations for both images and videos while retaining strong global scene understanding. The approach combines four key components. First,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Lorenzo Mur-Labadia , Matthew Muckley , Amir Bar , Mido Assran , Koustuv Sinha , Mike Rabbat , Yann LeCun , Nicolas Ballas , Adrien Bardes

In wireless networked control systems, ensuring timely and reliable state updates from distributed devices to remote controllers is essential for robust control performance. However, when multiple devices transmit high-dimensional states…

Systems and Control · Electrical Eng. & Systems 2026-02-10 Abanoub M. Girgis , Ibtissam Labriji , Mehdi Bennis

Building deep learning models that can reason about their environment requires capturing its underlying dynamics. Joint-Embedded Predictive Architectures (JEPA) provide a promising framework to model such dynamics by learning…

Machine Learning · Computer Science 2026-01-06 Matthieu Destrade , Oumayma Bounou , Quentin Le Lidec , Jean Ponce , Yann LeCun

Joint-embedding self-supervised learning (SSL) commonly relies on transformations such as data augmentation and masking to learn visual representations, a task achieved by enforcing invariance or equivariance with respect to these…

Computer Vision and Pattern Recognition · Computer Science 2026-01-12 Hafez Ghaemi , Eilif Muller , Shahab Bakhtiari