中文
相关论文

相关论文: TI-JEPA: An Innovative Energy-based Joint Embeddin…

200 篇论文

Modern Text-to-Image (T2I) generation increasingly relies on token-centric architectures that are trained with self-supervision, yet effectively fusing text with visual tokens remains a challenge. We propose \textbf{JEPA-T}, a unified…

计算机视觉与模式识别 · 计算机科学 2025-10-02 Siheng Wan , Zhengtao Yao , Zhengdao Li , Junhao Dong , Yanshu Li , Yikai Li , Linshan Li , Haoyan Xu , Yijiang Li , Zhikang Dong , Huacan Wang , Jifeng Shen

Image-based Joint-Embedding Predictive Architecture (I-JEPA) offers a promising approach to visual self-supervised learning through masked feature prediction. However with the inherent visual uncertainty at masked positions, feature…

机器学习 · 计算机科学 2026-05-06 Chen Huang , Xianhang Li , Vimal Thilak , Etai Littwin , Josh Susskind

This paper demonstrates an approach for learning highly semantic image representations without relying on hand-crafted data-augmentations. We introduce the Image-based Joint-Embedding Predictive Architecture (I-JEPA), a non-generative…

计算机视觉与模式识别 · 计算机科学 2023-04-14 Mahmoud Assran , Quentin Duval , Ishan Misra , Piotr Bojanowski , Pascal Vincent , Michael Rabbat , Yann LeCun , Nicolas Ballas

Current multimodal learning strategies primarily optimize in the original token space. Such a framework is easy to incorporate with the backbone of pretrained language model, but might result in modality collapse. To alleviate such issues,…

机器学习 · 计算机科学 2025-06-19 Hongyang Lei , Xiaolong Cheng , Qi Qin , Dan Wang , Kun Fan , Huazhen Huang , Qingqing Gu , Yetao Wu , Zhonglin Jiang , Yong Chen , Luo Ji

Self-supervised learning has emerged as a powerful paradigm for learning visual representations without manual annotations, yet most methods still operate on a single modality and therefore miss the complementary structure available from…

计算机视觉与模式识别 · 计算机科学 2026-04-02 Ciem Cornelissen , Sam Leroux , Pieter Simoens

Image-text matching is a key multimodal task that aims to model the semantic association between images and text as a matching relationship. With the advent of the multimedia information age, image, and text data show explosive growth, and…

机器学习 · 计算机科学 2024-06-24 Jinyin Wang , Haijing Zhang , Yihao Zhong , Yingbin Liang , Rongwei Ji , Yiru Cang

Self-supervised learning has seen great success recently in unsupervised representation learning, enabling breakthroughs in natural language and image processing. However, these methods often rely on autoregressive and masked modeling,…

机器学习 · 计算机科学 2025-10-01 Sofiane Ennadir , Siavash Golkar , Leopoldo Sarra

We present EB-JEPA, an open-source library for learning representations and world models using Joint-Embedding Predictive Architectures (JEPAs). JEPAs learn to predict in representation space rather than pixel space, avoiding the pitfalls…

Image-to-point cross-modal learning has emerged to address the scarcity of large-scale 3D datasets in 3D representation learning. However, current methods that leverage 2D data often result in large, slow-to-train models, making them…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Avishka Perera , Kumal Hewagamage , Saeedha Nazar , Kavishka Abeywardana , Hasitha Gallella , Ranga Rodrigo , Mohamed Afham

Future wireless systems increasingly require predictive and transferable representations that can support multiple physical-layer (PHY) tasks under dynamic environments. However, most existing supervised learning-based methods are designed…

信号处理 · 电气工程与系统科学 2026-04-01 Can Zheng , Jiguang He , Guofa Cai , Nannan Li , Mehdi Bennis , Henk Wymeersch , Merouane Debbah

This work introduces JEMA (Joint Embedding with Multimodal Alignment), a novel co-learning framework tailored for laser metal deposition (LMD), a pivotal process in metal additive manufacturing. As Industry 5.0 gains traction in industrial…

计算机视觉与模式识别 · 计算机科学 2024-11-01 Joao Sousa , Roya Darabi , Armando Sousa , Frank Brueckner , Luís Paulo Reis , Ana Reis

Large Language Model (LLM) pretraining, finetuning, and evaluation rely on input-space reconstruction and generative capabilities. Yet, it has been observed in vision that embedding-space training objectives, e.g., with Joint Embedding…

计算与语言 · 计算机科学 2025-10-08 Hai Huang , Yann LeCun , Randall Balestriero

We introduce VL-JEPA, a vision-language model built on a Joint Embedding Predictive Architecture (JEPA). Instead of autoregressively generating tokens as in classical VLMs, VL-JEPA predicts continuous embeddings of the target texts. By…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Delong Chen , Mustafa Shukor , Theo Moutakanni , Willy Chung , Jade Yu , Tejaswi Kasarla , Yejin Bang , Allen Bolourchi , Yann LeCun , Pascale Fung

This paper presents that the masked-modeling principle driving the success of large foundational vision models can be effectively applied to audio by making predictions in a latent space. We introduce Audio-based Joint-Embedding Predictive…

声音 · 计算机科学 2024-01-12 Zhengcong Fei , Mingyuan Fan , Junshi Huang

Image-based Joint-Embedding Predictive Architecture (IJEPA) offers an attractive alternative to Masked Autoencoder (MAE) for representation learning using the Masked Image Modeling framework. IJEPA drives representations to capture useful…

机器学习 · 计算机科学 2024-10-15 Etai Littwin , Vimal Thilak , Anand Gopalakrishnan

Multi-modality image fusion is a technique that combines information from different sensors or modalities, enabling the fused image to retain complementary features from each modality, such as functional highlights and texture details.…

计算机视觉与模式识别 · 计算机科学 2024-04-17 Zixiang Zhao , Haowen Bai , Jiangshe Zhang , Yulun Zhang , Kai Zhang , Shuang Xu , Dongdong Chen , Radu Timofte , Luc Van Gool

In recent years, advances in the large-scale pretraining of language and text-to-image models have revolutionized the field of machine learning. Yet, integrating these two modalities into a single, robust model capable of generating…

机器学习 · 计算机科学 2023-09-29 Emanuele Aiello , Lili Yu , Yixin Nie , Armen Aghajanyan , Barlas Oguz

This paper introduces a two-phase deep feature engineering framework for efficient learning of semantics enhanced joint embedding, which clearly separates the deep feature engineering in data preprocessing from training the text-image joint…

计算机视觉与模式识别 · 计算机科学 2021-10-25 Zhongwei Xie , Ling Liu , Yanzhao Wu , Luo Zhong , Lin Li

While MLLMs perform well on perceptual tasks, they lack precise multimodal alignment, limiting performance. To address this challenge, we propose Vision Dynamic Embedding-Guided Pretraining (VDEP), a hybrid autoregressive training paradigm…

计算机视觉与模式识别 · 计算机科学 2025-02-14 Mingxiao Li , Fang Qu , Zhanpeng Chen , Na Su , Zhizhou Zhong , Ziyang Chen , Nan Du , Xiaolong Li

Recent advances in self-supervised visual representation learning have demonstrated the effectiveness of predictive latent-space objectives for learning transferable features. In particular, Image-based Joint-Embedding Predictive…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Xiangteng He , Shunsuke Sakai , Shivam Chandhok , Sara Beery , Kun Yuan , Nicolas Padoy , Tatsuhito Hasegawa , Leonid Sigal
‹ 上一页 1 2 3 10 下一页 ›