English
Related papers

Related papers: Neural Implicit Action Fields: From Discrete Waypo…

200 papers

This paper presents the Never Ending Open Learning Adaptive Framework (NEOLAF), an integrated neural-symbolic cognitive architecture that models and constructs intelligent agents. The NEOLAF framework is a superior approach to constructing…

Our environment is filled with rich and dynamic acoustic information. When we walk into a cathedral, the reverberations as much as appearance inform us of the sanctuary's wide open space. Similarly, as an object moves around us, we expect…

Sound · Computer Science 2023-01-18 Andrew Luo , Yilun Du , Michael J. Tarr , Joshua B. Tenenbaum , Antonio Torralba , Chuang Gan

Medical image segmentation is often considered as the task of labelling each pixel or voxel as being inside or outside a given anatomy. Processing the images at their original size and resolution often result in insuperable memory…

Image and Video Processing · Electrical Eng. & Systems 2025-04-28 Kristine Sørensen , Oscar Camara , Ole de Backer , Klaus Kofoed , Rasmus Paulsen

Flow-based vision-language-action (VLA) models excel in embodied control but suffer from intractable likelihoods during multi-step sampling, hindering online reinforcement learning. We propose \textbf{\textit{$\boldsymbol{\pi}$-StepNFT}}…

Robotics · Computer Science 2026-03-10 Siting Wang , Xiaofeng Wang , Zheng Zhu , Minnan Pei , Xinyu Cui , Cheng Deng , Jian Zhao , Guan Huang , Haifeng Zhang , Jun Wang

Recent advances in Vision-Language-Action (VLA) models have established a two-component architecture, where a pre-trained Vision-Language Model (VLM) encodes visual observations and task descriptions, and an action decoder maps these…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Denis Tarasov , Alexander Nikulin , Ilya Zisman , Albina Klepach , Nikita Lyubaykin , Andrei Polubarov , Alexander Derevyagin , Vladislav Kurenkov

One of the central challenges preventing robots from acquiring complex manipulation skills is the prohibitive cost of collecting large-scale robot demonstrations. In contrast, humans are able to learn efficiently by watching others interact…

Robotics · Computer Science 2025-11-13 Changhe Chen , Quantao Yang , Xiaohao Xu , Nima Fazeli , Olov Andersson

Vision-language-action (VLA) models have recently emerged as a promising paradigm for robotic control, enabling end-to-end policies that ground natural language instructions into visuomotor actions. However, current VLAs often struggle to…

Robotics · Computer Science 2025-09-18 Zijian An , Ran Yang , Yiming Feng , Lifeng Zhou

We present Neural Articulated Radiance Field (NARF), a novel deformable 3D representation for articulated objects learned from images. While recent advances in 3D implicit representation have made it possible to learn models of complex…

Computer Vision and Pattern Recognition · Computer Science 2021-08-19 Atsuhiro Noguchi , Xiao Sun , Stephen Lin , Tatsuya Harada

Hybrid neural-physics modeling frameworks through differentiable programming have emerged as powerful tools in scientific machine learning, enabling the integration of known physics with data-driven learning to improve prediction accuracy…

Machine Learning · Computer Science 2025-04-04 Deepak Akhare , Pan Du , Tengfei Luo , Jian-Xun Wang

We propose a new continuous video modeling framework based on implicit neural representations (INRs) called ActINR. At the core of our approach is the observation that INRs can be considered as a learnable dictionary, with the shapes of the…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Alper Kayabasi , Anil Kumar Vadathya , Guha Balakrishnan , Vishwanath Saragadam

World models predict future transitions from observations and actions. Existing works predominantly focus on image generation only. Visual feature-based world models, on the other hand, predict future visual features instead of raw video…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Xinyu Zhang , Zhengtong Xu , Yutian Tao , Yeping Wang , Yu She , Abdeslam Boularias

Deep Implicit Functions (DIFs) represent 3D geometry with continuous signed distance functions learned through deep neural nets. Recently DIFs-based methods have been proposed to handle shape reconstruction and dense point correspondences…

Computer Vision and Pattern Recognition · Computer Science 2022-03-23 Shanlin Sun , Kun Han , Deying Kong , Hao Tang , Xiangyi Yan , Xiaohui Xie

Feedback Alignment (FA) methods are biologically inspired local learning rules for training neural networks with reduced communication between layers. While FA has potential applications in distributed and privacy-aware ML, limitations in…

Machine Learning · Computer Science 2024-06-05 Zachary Robertson , Oluwasanmi Koyejo

Vision-Language-Action (VLA) models are increasingly expected to not only complete robot tasks, but also follow human instructions about how those tasks should be executed. However, existing robot datasets usually pair trajectories with…

This paper tackles spatial perception and manipulation challenges in Vision-Language-Action (VLA) models. To address depth ambiguity from monocular input, we leverage a pre-trained multi-view diffusion model to synthesize latent novel views…

We propose Avi, a novel 3D Vision-Language-Action (VLA) architecture that reframes robotic action generation as a problem of 3D perception and spatial reasoning, rather than low-level policy learning. While existing VLA models primarily…

Robotics · Computer Science 2025-10-28 Harris Song , Long Le

Video action localization aims to find the timings of specific actions from a long video. Although existing learning-based approaches have been successful, they require annotating videos, which comes with a considerable labor cost. This…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Naoki Wake , Atsushi Kanehira , Kazuhiro Sasabuchi , Jun Takamatsu , Katsushi Ikeuchi

Normalizing flows and autoregressive models have been successfully combined to produce state-of-the-art results in density estimation, via Masked Autoregressive Flows (MAF), and to accelerate state-of-the-art WaveNet-based speech synthesis…

Machine Learning · Computer Science 2018-04-04 Chin-Wei Huang , David Krueger , Alexandre Lacoste , Aaron Courville

Learning latent actions from large-scale videos is crucial for the pre-training of scalable embodied foundation models, yet existing methods often struggle with action-irrelevant distractors. Although incorporating action supervision can…

Robotics · Computer Science 2026-03-24 Xizhou Bu , Jiexi Lyu , Fulei Sun , Ruichen Yang , Zhiqiang Ma , Wei Li

The ability to accurately model interatomic interactions in large-scale systems is fundamental to understanding a wide range of physical and chemical phenomena, from drug-protein binding to the behavior of next-generation materials. While…

Materials Science · Physics 2025-05-26 Taskin Mehereen , Sourav Saha , Intesar Jawad Jaigirdar , Chanwook Park
‹ Prev 1 4 5 6 7 8 10 Next ›