English
Related papers

Related papers: LeWorldModel: Stable End-to-End Joint-Embedding Pr…

200 papers

While video-generation-based embodied world models have gained increasing attention, their reliance on large-scale embodied interaction data remains a key bottleneck. The scarcity, difficulty of collection, and high dimensionality of…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Hao Li , Qiao Sun

With generative models producing high quality images that are indistinguishable from real ones, there is growing concern regarding the malicious usage of AI-generated images. Imperceptible image watermarking is one viable solution towards…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Ahmad Rezaei , Mohammad Akbari , Saeed Ranjbar Alvar , Arezou Fatemi , Yong Zhang

Self-supervised learning in healthcare has largely relied on invariance-based objectives, which maximize similarity between different views of the same patient. While effective for static anatomy, this paradigm is fundamentally misaligned…

Machine Learning · Computer Science 2026-04-27 Jose Geraldo Fernandes , Luiz Facury , Pedro Robles Dutenhefner , Wagner Meira

Video Joint Embedding Predictive Architectures (V-JEPA) learn generalizable off-the-shelf video representation by predicting masked regions in latent space with an exponential moving average (EMA)-updated teacher. While EMA prevents…

Machine Learning · Computer Science 2025-09-30 Xianhang Li , Chen Huang , Chun-Liang Li , Eran Malach , Josh Susskind , Vimal Thilak , Etai Littwin

We show how perceptual embeddings of the visual system can be constructed at inference-time with no training data or deep neural network features. Our perceptual embeddings are solutions to a weighted least squares (WLS) problem, defined at…

Computer Vision and Pattern Recognition · Computer Science 2023-10-11 Daniel Severo , Lucas Theis , Johannes Ballé

Learning audio representations from raw waveforms overcomes key limitations of spectrogram-based audio representation learning, such as the long latency of spectrogram computation and the loss of phase information. Yet, while…

Modern vision-based world models can represent observations as compact yet expressive latent manifolds, but fast goal-oriented planning in these spaces remains challenging. This raises a central question: when does a learned representation…

Robotics · Computer Science 2026-05-12 Hoang Nguyen , Xiaohao Xu , Xiaonan Huang

A promising approach to preserving model performance in linearized transformers is to employ position-based re-weighting functions. However, state-of-the-art re-weighting functions rely heavily on target sequence lengths, making it…

Computation and Language · Computer Science 2024-05-24 Victor Agostinelli , Sanghyun Hong , Lizhong Chen

Various world model frameworks are being developed today based on autoregressive frameworks that rely on discrete representations of actions and observations, and these frameworks are succeeding in constructing interactive generative models…

Machine Learning · Computer Science 2025-03-14 Kohei Hayashi , Masanori Koyama , Julian Jorge Andrade Guerreiro

Artificial intelligence systems in critical fields like autonomous driving and medical imaging analysis often continually learn new tasks using a shared stream of input data. For instance, after learning to detect traffic signs, a model may…

Machine Learning · Computer Science 2025-11-18 Hanchen David Wang , Siwoo Bae , Zirong Chen , Meiyi Ma

A major challenge in deploying world models is the trade-off between size and performance. Large world models can capture rich physical dynamics but require massive computing resources, making them impractical for edge devices. Small world…

Artificial Intelligence · Computer Science 2025-09-17 Dingrui Wang , Zhexiao Sun , Zhouheng Li , Cheng Wang , Youlun Peng , Hongyuan Ye , Baha Zarrouki , Wei Li , Mattia Piccinini , Lei Xie , Johannes Betz

This paper proposes a Learnable Multiplicative absolute position Embedding based Conformer (LMEC). It contains a kernelized linear attention (LA) module called LMLA to solve the time-consuming problem for long sequence speech recognition as…

Audio and Speech Processing · Electrical Eng. & Systems 2022-12-06 Yuguang Yang , Yu Pan , Jingjing Yin , Heng Lu

Recent advancements in self-supervised learning in the point cloud domain have demonstrated significant potential. However, these methods often suffer from drawbacks, including lengthy pre-training time, the necessity of reconstruction in…

Computer Vision and Pattern Recognition · Computer Science 2025-02-11 Ayumu Saito , Prachi Kudeshia , Jiju Poovvancheri

Deep Learning based Weather Prediction (DLWP) models have been improving rapidly over the last few years, surpassing state of the art numerical weather forecasts by significant margins. While much of the optimization effort is focused on…

Atmospheric and Oceanic Physics · Physics 2024-08-15 Haoyu Qin , Yungang Chen , Qianchuan Jiang , Pengchao Sun , Xiancai Ye , Chao Lin

Legged locomotion over various terrains is challenging and requires precise perception of the robot and its surroundings from both proprioception and vision. However, learning directly from high-dimensional visual input is often…

Robotics · Computer Science 2024-09-26 Hang Lai , Jiahang Cao , Jiafeng Xu , Hongtao Wu , Yunfeng Lin , Tao Kong , Yong Yu , Weinan Zhang

We propose a new architecture for the learning of predictive spatio-temporal motion models from data alone. Our approach, dubbed the Dropout Autoencoder LSTM, is capable of synthesizing natural looking motion sequences over long time…

Computer Vision and Pattern Recognition · Computer Science 2017-12-05 Partha Ghosh , Jie Song , Emre Aksan , Otmar Hilliges

Learning latent representations that capture both semantic and spatial information is central to efficient spatio-semantic reasoning. However, many existing approaches rely on implicit latent structures combined with dense feature maps or…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 SeongMin Jin , Doo Seok Jeong

Accurate and automatic detection of mood serves as a building block for use cases like user profiling which in turn power applications such as advertising, recommendation systems, and many more. One primary source indicative of an…

Computation and Language · Computer Science 2022-02-09 Harichandana B S S , Sumit Kumar

World models are widely explored in embodied intelligence, yet they typically predict distinct evolutions of the world and the ego within a single stream, where the world captures persistent instruction-agnostic scene regularities and the…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Zuyao Lin , Jianhui Zhang , Peidong Jia , Xiaoguang Zhao , Shanghang Zhang , Xingyu Chen

Neural networks are very effective when trained on large datasets for a large number of iterations. However, when they are trained on non-stationary streams of data and in an online fashion, their performance is reduced (1) by the online…

Machine Learning · Computer Science 2023-07-04 Albin Soutif--Cormerais , Antonio Carta , Joost Van de Weijer