English
Related papers

Related papers: DMT-JEPA: Discriminative Masked Targets for Joint-…

200 papers

Graph Neural Networks (GNNs) have shown promise in learning dynamic functional connectivity for distinguishing phenotypes from human brain networks. However, obtaining extensive labeled clinical data for training is often…

Machine Learning · Computer Science 2025-05-06 Jungwon Choi , Hyungi Lee , Byung-Hoon Kim , Juho Lee

Current multimodal learning strategies primarily optimize in the original token space. Such a framework is easy to incorporate with the backbone of pretrained language model, but might result in modality collapse. To alleviate such issues,…

Machine Learning · Computer Science 2025-06-19 Hongyang Lei , Xiaolong Cheng , Qi Qin , Dan Wang , Kun Fan , Huazhen Huang , Qingqing Gu , Yetao Wu , Zhonglin Jiang , Yong Chen , Luo Ji

This paper introduces a novel application of Video Joint-Embedding Predictive Architectures (V-JEPAs) for Facial Expression Recognition (FER). Departing from conventional pre-training methods for video understanding that rely on pixel-level…

Computer Vision and Pattern Recognition · Computer Science 2026-01-15 Lennart Eing , Cristina Luna-Jiménez , Silvan Mertes , Elisabeth André

We introduce a two-stage self-supervised framework that combines the Joint-Embedding Predictive Architecture (JEPA) with a Density Adaptive Attention Mechanism (DAAM) for learning robust speech representations. Stage~1 uses JEPA with DAAM…

This paper focuses on multimodal alignment within the realm of Artificial Intelligence, particularly in text and image modalities. The semantic gap between the textual and visual modality poses a discrepancy problem towards the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Khang H. N. Vo , Duc P. T. Nguyen , Thong Nguyen , Tho T. Quan

Joint-Embedding Predictive Architectures (JEPA) have recently become popular as promising architectures for self-supervised learning. Vision transformers have been trained using JEPA to produce embeddings from images and videos, which have…

Computer Vision and Pattern Recognition · Computer Science 2025-04-24 Tristan Kenneweg , Philip Kenneweg , Barbara Hammer

Joint-embedding self-supervised learning (SSL) commonly relies on transformations such as data augmentation and masking to learn visual representations, a task achieved by enforcing invariance or equivariance with respect to these…

Computer Vision and Pattern Recognition · Computer Science 2026-01-12 Hafez Ghaemi , Eilif Muller , Shahab Bakhtiari

Inspired by the success of generative pretraining in natural language, we ask whether the same principles can yield strong self-supervised visual learners. Instead of training models to output features for downstream use, we train them to…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Sihan Xu , Ziqiao Ma , Wenhao Chai , Xuweiyi Chen , Weiyang Jin , Joyce Chai , Saining Xie , Stella X. Yu

Latent prediction--where agents learn by predicting their own latents--has emerged as a powerful paradigm for training general representations in machine learning. In reinforcement learning (RL), this approach has been explored to define…

Machine Learning · Computer Science 2025-10-02 Marco Bagatella , Matteo Pirotta , Ahmed Touati , Alessandro Lazaric , Andrea Tirinzoni

Self-supervised learning (SSL) has become an important approach in pretraining large neural networks, enabling unprecedented scaling of model and dataset sizes. While recent advances like I-JEPA have shown promising results for Vision…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 András Kalapos , Bálint Gyires-Tóth

Joint-embedding predictive architectures (JEPAs) propose that a model should learn more useful abstractions when trained to predict latent representations rather than observed outputs. For autoregressive language-model fine-tuning the…

Machine Learning · Computer Science 2026-05-18 Biswa Sengupta

Joint-Embedding Predictive Architectures (JEPAs) aim to learn representations by predicting target embeddings from context embeddings, inducing a scalar compatibility energy in a latent space. In contrast, Quasimetric Reinforcement Learning…

Machine Learning · Computer Science 2026-02-13 Anthony Kobanda , Waris Radji

The development of multimodal models for pulmonary nodule diagnosis is limited by the scarcity of labeled data and the tendency for these models to overfit on the training distribution. In this work, we leverage self-supervised learning…

Computer Vision and Pattern Recognition · Computer Science 2025-09-22 Thomas Z. Li , Aravind R. Krishnan , Lianrui Zuo , John M. Still , Kim L. Sandler , Fabien Maldonado , Thomas A. Lasko , Bennett A. Landman

Joint-Embedding Predictive Architecture (JEPA) is increasingly used for visual representation learning and as a component in model-based RL, but its behavior remains poorly understood. We provide a theoretical characterization of a simple,…

Machine Learning · Computer Science 2025-10-21 Jiacan Yu , Siyi Chen , Mingrui Liu , Nono Horiuchi , Vladimir Braverman , Zicheng Xu , Dan Haramati , Randall Balestriero

Self-supervised learning has emerged as a powerful paradigm for learning visual representations without manual annotations, yet most methods still operate on a single modality and therefore miss the complementary structure available from…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Ciem Cornelissen , Sam Leroux , Pieter Simoens

Geospatial foundation models provide precomputed embeddings that serve as compact feature vectors for large-scale satellite remote sensing data. While these embeddings can reduce data-transfer bottlenecks and computational costs, Earth…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Erik Scheurer , Rocco Sedona , Stefan Kesselheim , Gabriele Cavallaro

In remote control systems, transmitting large data volumes (e.g., images, video frames) from wireless sensors to remote controllers is challenging when uplink capacity is limited (e.g., RedCap devices or massive wireless sensor networks).…

Information Theory · Computer Science 2025-07-03 Abanoub M. Girgis , Alvaro Valcarce , Mehdi Bennis

Video Joint Embedding Predictive Architectures (V-JEPA) learn generalizable off-the-shelf video representation by predicting masked regions in latent space with an exponential moving average (EMA)-updated teacher. While EMA prevents…

Machine Learning · Computer Science 2025-09-30 Xianhang Li , Chen Huang , Chun-Liang Li , Eran Malach , Josh Susskind , Vimal Thilak , Etai Littwin

In wireless networked control systems, ensuring timely and reliable state updates from distributed devices to remote controllers is essential for robust control performance. However, when multiple devices transmit high-dimensional states…

Systems and Control · Electrical Eng. & Systems 2026-02-10 Abanoub M. Girgis , Ibtissam Labriji , Mehdi Bennis

While nowadays deep neural networks achieve impressive performances on semantic segmentation tasks, they are usually trained by optimizing pixel-wise losses such as cross-entropy. As a result, the predictions outputted by such networks…

Computer Vision and Pattern Recognition · Computer Science 2019-05-07 Yifu Chen , Arnaud Dapogny , Matthieu Cord