中文
相关论文

相关论文: HanoiWorld : A Joint Embedding Predictive Architec…

200 篇论文

We evaluate JEPA-style predictive representation learning versus reconstruction-based autoencoders on a controlled "TV-series" linear dynamical system with known latent state and a single noise parameter. While an initial comparison…

机器学习 · 计算机科学 2026-03-17 Alexey Potapov , Oleg Shcherbakov , Ivan Kravchenko

Vision-Language-Action (VLA) models have recently achieved notable progress in end-to-end autonomous driving by integrating perception, reasoning, and control within a unified multimodal framework. However, they often lack explicit modeling…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Guoqing Wang , Pin Tang , Xiangxuan Ren , Guodongfang Zhao , Bailan Feng , Chao Ma

End-to-end autonomous driving provides a feasible way to automatically maximize overall driving system performance by directly mapping the raw pixels from a front-facing camera to control signals. Recent advanced methods construct a latent…

机器学习 · 计算机科学 2024-05-21 Zeyu Gao , Yao Mu , Chen Chen , Jingliang Duan , Shengbo Eben Li , Ping Luo , Yanfeng Lu

Robotic imitation learning is often treated as reproducing demonstrated actions, but actions are inherently embodiment-specific. When demonstrations come from humans or robots with different morphology, kinematics, or action spaces, this…

机器人学 · 计算机科学 2026-05-21 Jingyang He , Guangrun Li , Jieyu Zhang , Chengkai Hou , Zhengping Che , Shanghang Zhang

It is expected that many human drivers will still prefer to drive themselves even if the self-driving technologies are ready. Therefore, human-driven vehicles and autonomous vehicles (AVs) will coexist in a mixed traffic for a long time. To…

机器人学 · 计算机科学 2019-10-14 Dong Chen , Longsheng Jiang , Yue Wang , Zhaojian Li

This paper presents that the masked-modeling principle driving the success of large foundational vision models can be effectively applied to audio by making predictions in a latent space. We introduce Audio-based Joint-Embedding Predictive…

声音 · 计算机科学 2024-01-12 Zhengcong Fei , Mingyuan Fan , Junshi Huang

Building on the Joint-Embedding Predictive Architecture (JEPA) paradigm, a recent self-supervised learning framework that predicts latent representations of masked regions in high-level feature spaces, we propose Audio-JEPA (Audio…

声音 · 计算机科学 2025-07-08 Ludovic Tuncay , Etienne Labbé , Emmanouil Benetos , Thomas Pellegrini

World models enable model-based planning through learned latent dynamics, but imagined rollouts become unstable as the planning horizon grows or the dynamics distribution shifts. We argue that this instability reflects two missing…

人工智能 · 计算机科学 2026-05-08 Haoyun Tang , Haodong Cui , Keyao Xu , Kun Wang , Zhandong Mei

Adapting to unforeseen novelties in open-world environments remains a major challenge for autonomous systems. While hybrid planning and reinforcement learning (RL) approaches show promise, they often suffer from sample inefficiency, slow…

机器人学 · 计算机科学 2026-01-27 Pierrick Lorang

In the area of learning-driven artificial intelligence advancement, the integration of machine learning (ML) into self-driving (SD) technology stands as an impressive engineering feat. Yet, in real-world applications outside the confines of…

机器人学 · 计算机科学 2023-09-06 Haozhe Lei , Quanyan Zhu

End-to-end autonomous driving has emerged as a compelling alternative to traditional modular pipelines by directly mapping raw sensor data to driving actions. While recent approaches achieve strong performance on single-domain datasets,…

机器人学 · 计算机科学 2026-05-20 Hoonhee Cho , Giwon Lee , Jae-Young Kang , Hyemin Yang , Heejun Park , Kuk-Jin Yoon

Collision avoidance systems can play a vital role in reducing the number of accidents and saving human lives. In this paper, we introduce and validate a novel method for vehicles reactive collision avoidance using evolutionary neural…

神经与进化计算 · 计算机科学 2016-09-28 Hesham Eraqi , Youssef EmadEldin , Mohamed Moustafa

Navigating to a visually specified goal given natural language instructions remains a fundamental challenge in embodied AI. Existing approaches either rely on reactive policies that struggle with long-horizon planning, or employ world…

机器人学 · 计算机科学 2026-03-30 Amirhosein Chahe , Lifeng Zhou

Learning predictive world models from unlabelled video is a foundational challenge in artificial intelligence. While Joint Embedding Predictive Architectures (JEPA) have set new benchmarks in semantic classification, they often remain…

计算机视觉与模式识别 · 计算机科学 2026-05-18 Santosh Kumar Paidi

Recent advances in self-supervised visual representation learning have demonstrated the effectiveness of predictive latent-space objectives for learning transferable features. In particular, Image-based Joint-Embedding Predictive…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Xiangteng He , Shunsuke Sakai , Shivam Chandhok , Sara Beery , Kun Yuan , Nicolas Padoy , Tatsuhito Hasegawa , Leonid Sigal

To safely navigate intricate real-world scenarios, autonomous vehicles must be able to adapt to diverse road conditions and anticipate future events. World model (WM) based reinforcement learning (RL) has emerged as a promising approach by…

机器人学 · 计算机科学 2024-07-29 Dechen Gao , Shuangyu Cai , Hanchu Zhou , Hang Wang , Iman Soltani , Junshan Zhang

We present a transformer architecture-based foundation model for tasks at high-energy particle colliders such as the Large Hadron Collider. We train the model to classify jets using a self-supervised strategy inspired by the Joint Embedding…

机器学习 · 计算机科学 2025-02-07 Jai Bardhan , Radhikesh Agrawal , Abhiram Tilak , Cyrin Neeraj , Subhadip Mitra

Vision-centric autonomous driving has recently raised wide attention due to its lower cost. Pre-training is essential for extracting a universal representation. However, current vision-centric pre-training typically relies on either 2D or…

计算机视觉与模式识别 · 计算机科学 2024-05-08 Chen Min , Dawei Zhao , Liang Xiao , Jian Zhao , Xinli Xu , Zheng Zhu , Lei Jin , Jianshu Li , Yulan Guo , Junliang Xing , Liping Jing , Yiming Nie , Bin Dai

Autonomous driving has attracted great interest due to its potential capability in full-unsupervised driving. Model-based and learning-based methods are widely used in autonomous driving. Model-based methods rely on pre-defined models of…

Modern self-supervised predictive architectures excel at capturing complex statistical correlations from high-dimensional data but lack mechanisms to internalize verifiable human logic, leaving them susceptible to spurious correlations and…

机器学习 · 计算机科学 2026-03-17 Yongchao Huang , Hassan Raza