中文
相关论文

相关论文: Self-Supervised Pre-Training with Joint-Embedding …

200 篇论文

Deep learning models need a sufficient amount of data in order to be able to find the hidden patterns in it. It is the purpose of generative modeling to learn the data distribution, thus allowing us to sample more data and augment the…

机器学习 · 计算机科学 2024-11-28 José Fernando Núñez , Jamie Arjona , Javier Béjar

This paper introduces a novel application of Video Joint-Embedding Predictive Architectures (V-JEPAs) for Facial Expression Recognition (FER). Departing from conventional pre-training methods for video understanding that rely on pixel-level…

计算机视觉与模式识别 · 计算机科学 2026-01-15 Lennart Eing , Cristina Luna-Jiménez , Silvan Mertes , Elisabeth André

In wireless networked control systems, ensuring timely and reliable state updates from distributed devices to remote controllers is essential for robust control performance. However, when multiple devices transmit high-dimensional states…

系统与控制 · 电气工程与系统科学 2026-02-10 Abanoub M. Girgis , Ibtissam Labriji , Mehdi Bennis

Visual Speech Recognition (VSR) tasks are generally recognized to have a lower theoretical performance ceiling than Automatic Speech Recognition (ASR), owing to the inherent limitations of conveying semantic information visually. To…

计算机视觉与模式识别 · 计算机科学 2024-03-29 Chang Sun , Hong Yang , Bo Qin

We evaluate JEPA-style predictive representation learning versus reconstruction-based autoencoders on a controlled "TV-series" linear dynamical system with known latent state and a single noise parameter. While an initial comparison…

机器学习 · 计算机科学 2026-03-17 Alexey Potapov , Oleg Shcherbakov , Ivan Kravchenko

Language representation learning has emerged as a promising approach for sequential recommendation, thanks to its ability to learn generalizable representations. However, despite its advantages, this approach still struggles with data…

信息检索 · 计算机科学 2025-08-08 Minh-Anh Nguyen , Dung D. Le

Task-specific pre-training is essential when task representations diverge from generic pre-training features. Existing task-general pre-training EEG models struggle with complex tasks like emotion recognition due to mismatches between…

机器学习 · 计算机科学 2025-10-28 Qingzhu Zhang , Jiani Zhong , Zongsheng Li , Xinke Shen , Quanying Liu

Heart rate (HR) estimation from photoplethysmography (PPG) signals is a key feature of modern wearable devices for health and wellness monitoring. While deep learning models show promise, their performance relies on the availability of…

Autonomous driving, as an agent operating in the physical world, requires the fundamental capability to build \textit{world models} that capture how the environment evolves spatiotemporally in order to support long-term planning. At the…

计算机视觉与模式识别 · 计算机科学 2026-02-16 Haoran Zhu , Anna Choromanska

Electrocardiogram (ECG) is a simple non-invasive measure to identify heart-related issues such as irregular heartbeats known as arrhythmias. While artificial intelligence and machine learning is being utilized in a wide range of healthcare…

机器学习 · 计算机科学 2022-07-11 Minh Cao , Tianqi Zhao , Yanxun Li , Wenhao Zhang , Peyman Benharash , Ramin Ramezani

Joint-Embedding Predictive Architectures (JEPAs) provide a simpleframework for learning world models by predicting future latent representations.However, JEPA training is subject to a bias-variance tradeoff.Without sufficient structural…

机器学习 · 计算机科学 2026-05-12 Kai Zhao , Dongliang Nie , Yuchen Lin , Zhehan Luo , Yixiao Gu , Deng-Ping Fan , Dan Zeng

Cardiovascular diseases are a leading cause of death and disability worldwide. Electrocardiogram (ECG) is critical for diagnosing and monitoring cardiac health, but obtaining large-scale annotated ECG datasets is labor-intensive and…

信号处理 · 电气工程与系统科学 2025-09-22 Mingsheng Cai , Jiuming Jiang , Wenhao Huang , Che Liu , Rossella Arcucci

Building deep learning models that can reason about their environment requires capturing its underlying dynamics. Joint-Embedded Predictive Architectures (JEPA) provide a promising framework to model such dynamics by learning…

机器学习 · 计算机科学 2026-01-06 Matthieu Destrade , Oumayma Bounou , Quentin Le Lidec , Jean Ponce , Yann LeCun

Electrocardiogram (ECG) arrhythmia classification remains challenging due to signal variability, noise, limited labeled data, and the difficulty in achieving both accuracy and efficiency in models. While self-supervised learning reduces…

机器学习 · 计算机科学 2026-05-14 Mahsa Gazeran , Sayvan Soleymanbaigi , Fatemeh Daneshfar , Amjad Seyedi , Fardin Akhlaghian Tab

Myocardial infarction is a major cause of death globally, and accurate early diagnosis from electrocardiograms (ECGs) remains a clinical priority. Deep learning models have shown promise for automated ECG interpretation, but require large…

图像与视频处理 · 电气工程与系统科学 2025-07-01 Lachin Naghashyar

Electrocardiogram (ECG) signal is one of the most effective sources of information mainly employed for the diagnosis and prediction of cardiovascular diseases (CVDs) connected with the abnormalities in heart rhythm. Clearly, single modality…

信号处理 · 电气工程与系统科学 2022-10-13 Thinh Phan , Duc Le , Patel Brijesh , Donald Adjeroh , Jingxian Wu , Morten Olgaard Jensen , Ngan Le

Electrocardiogram (ECG) diagnosis remains challenging due to limited labeled data and the need to capture subtle yet clinically meaningful variations in rhythm and morphology. We present CREMA (Contrastive Regularized Masked Autoencoder), a…

机器学习 · 计算机科学 2025-08-22 Junho Song , Jong-Hwan Jang , DongGyun Hong , Joon-myoung Kwon , Yong-Yeon Jo

Current multimodal learning strategies primarily optimize in the original token space. Such a framework is easy to incorporate with the backbone of pretrained language model, but might result in modality collapse. To alleviate such issues,…

机器学习 · 计算机科学 2025-06-19 Hongyang Lei , Xiaolong Cheng , Qi Qin , Dan Wang , Kun Fan , Huazhen Huang , Qingqing Gu , Yetao Wu , Zhonglin Jiang , Yong Chen , Luo Ji

A major challenge for modern AI is to learn to understand the world and learn to act largely by observation. This paper explores a self-supervised approach that combines internet-scale video data with a small amount of interaction data…

Self-supervised learning has produced impressive results in multimedia domains of audio, vision and speech. This paradigm is equally, if not more, relevant for the domain of biosignals, owing to the scarcity of labelled data in such…

机器学习 · 计算机科学 2024-03-07 Aditya Kommineni , Kleanthis Avramidis , Richard Leahy , Shrikanth Narayanan