English
Related papers

Related papers: Audio-JEPA: Joint-Embedding Predictive Architectur…

200 papers

This paper focuses on multimodal alignment within the realm of Artificial Intelligence, particularly in text and image modalities. The semantic gap between the textual and visual modality poses a discrepancy problem towards the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Khang H. N. Vo , Duc P. T. Nguyen , Thong Nguyen , Tho T. Quan

Joint-Embedding Predictive Architecture (JEPA) is increasingly used for visual representation learning and as a component in model-based RL, but its behavior remains poorly understood. We provide a theoretical characterization of a simple,…

Machine Learning · Computer Science 2025-10-21 Jiacan Yu , Siyi Chen , Mingrui Liu , Nono Horiuchi , Vladimir Braverman , Zicheng Xu , Dan Haramati , Randall Balestriero

The representation of urban trajectory data plays a critical role in effectively analyzing spatial movement patterns. Despite considerable progress, the challenge of designing trajectory representations that can capture diverse and…

Machine Learning · Computer Science 2025-07-02 Lihuan Li , Hao Xue , Shuang Ao , Yang Song , Flora Salim

Learning predictive world models from unlabelled video is a foundational challenge in artificial intelligence. While Joint Embedding Predictive Architectures (JEPA) have set new benchmarks in semantic classification, they often remain…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Santosh Kumar Paidi

We present Clin-JEPA, a multi-phase co-training framework for joint-embedding predictive (JEPA) pretraining on EHR patient trajectories. JEPA architectures have enabled latent-space planning in robotics and high-quality representation…

Self-supervised learning of visual representations has been focusing on learning content features, which do not capture object motion or location, and focus on identifying and differentiating objects in images and videos. On the other hand,…

Computer Vision and Pattern Recognition · Computer Science 2023-07-25 Adrien Bardes , Jean Ponce , Yann LeCun

Robotic imitation learning is often treated as reproducing demonstrated actions, but actions are inherently embodiment-specific. When demonstrations come from humans or robots with different morphology, kinematics, or action spaces, this…

Robotics · Computer Science 2026-05-21 Jingyang He , Guangrun Li , Jieyu Zhang , Chengkai Hou , Zhengping Che , Shanghang Zhang

Building deep learning models that can reason about their environment requires capturing its underlying dynamics. Joint-Embedded Predictive Architectures (JEPA) provide a promising framework to model such dynamics by learning…

Machine Learning · Computer Science 2026-01-06 Matthieu Destrade , Oumayma Bounou , Quentin Le Lidec , Jean Ponce , Yann LeCun

Recent general-purpose audio representations show state-of-the-art performance on various audio tasks. These representations are pre-trained by self-supervised learning methods that create training signals from the input. For example,…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-09 Daisuke Niizumi , Daiki Takeuchi , Yasunori Ohishi , Noboru Harada , Kunio Kashino

Video Joint Embedding Predictive Architectures (V-JEPA) learn generalizable off-the-shelf video representation by predicting masked regions in latent space with an exponential moving average (EMA)-updated teacher. While EMA prevents…

Machine Learning · Computer Science 2025-09-30 Xianhang Li , Chen Huang , Chun-Liang Li , Eran Malach , Josh Susskind , Vimal Thilak , Etai Littwin

Joint-embedding predictive architectures (JEPAs) propose that a model should learn more useful abstractions when trained to predict latent representations rather than observed outputs. For autoregressive language-model fine-tuning the…

Machine Learning · Computer Science 2026-05-18 Biswa Sengupta

The rapid expansion of remote sensing image archives demands the development of strong and efficient techniques for content-based image retrieval (RS-CBIR). This paper presents REJEPA (Retrieval with Joint-Embedding Predictive…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Shabnam Choudhury , Yash Salunkhe , Sarthak Mehrotra , Biplab Banerjee

Semi-supervised learning has emerged as a powerful paradigm for leveraging large amounts of unlabeled data to improve the performance of machine learning models when labeled data are scarce. Among existing approaches, methods derived from…

Machine Learning · Computer Science 2026-04-29 Ali Aghababaei-Harandi , Aude Sportisse , Massih-Reza Amini

Video world models trained with Joint Embedding Predictive Architectures (JEPA) acquire rich spatiotemporal representations by predicting masked regions in latent space rather than reconstructing pixels. This removes the visual verification…

Machine Learning · Computer Science 2026-03-24 Liu hung ming

Modern Text-to-Image (T2I) generation increasingly relies on token-centric architectures that are trained with self-supervision, yet effectively fusing text with visual tokens remains a challenge. We propose \textbf{JEPA-T}, a unified…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Siheng Wan , Zhengtao Yao , Zhengdao Li , Junhao Dong , Yanshu Li , Yikai Li , Linshan Li , Haoyan Xu , Yijiang Li , Zhikang Dong , Huacan Wang , Jifeng Shen

Aerodynamic surrogate models are increasingly used to replace repeated high-fidelity CFD evaluations in many-query design settings, but current approaches still face two important limitations: they often scale poorly to the very large…

Joint-Embedding Predictive Architectures (JEPAs) provide a simpleframework for learning world models by predicting future latent representations.However, JEPA training is subject to a bias-variance tradeoff.Without sufficient structural…

Machine Learning · Computer Science 2026-05-12 Kai Zhao , Dongliang Nie , Yuchen Lin , Zhehan Luo , Yixiao Gu , Deng-Ping Fan , Dan Zeng

Language-audio joint representation learning frameworks typically depend on deterministic embeddings, assuming a one-to-one correspondence between audio and text. In real-world settings, however, the language-audio relationship is…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-22 Toranosuke Manabe , Yuchi Ishikawa , Hokuto Munakata , Tatsuya Komatsu

The development of multimodal models for pulmonary nodule diagnosis is limited by the scarcity of labeled data and the tendency for these models to overfit on the training distribution. In this work, we leverage self-supervised learning…

Computer Vision and Pattern Recognition · Computer Science 2025-09-22 Thomas Z. Li , Aravind R. Krishnan , Lianrui Zuo , John M. Still , Kim L. Sandler , Fabien Maldonado , Thomas A. Lasko , Bennett A. Landman

We evaluate JEPA-style predictive representation learning versus reconstruction-based autoencoders on a controlled "TV-series" linear dynamical system with known latent state and a single noise parameter. While an initial comparison…

Machine Learning · Computer Science 2026-03-17 Alexey Potapov , Oleg Shcherbakov , Ivan Kravchenko