English
Related papers

Related papers: Any3D-VLA: Enhancing VLA Robustness via Diverse Po…

200 papers

Successful generalist Vision-Language-Action (VLA) models rely on effective training across diverse robotic platforms with large-scale, cross-embodiment, heterogeneous datasets. To facilitate and leverage the heterogeneity in rich, diverse…

The rapid advancement of generative AI and multi-modal foundation models has shown significant potential in advancing robotic manipulation. Vision-language-action (VLA) models, in particular, have emerged as a promising approach for…

Software Engineering · Computer Science 2025-05-13 Zhijie Wang , Zhehua Zhou , Jiayang Song , Yuheng Huang , Zhan Shu , Lei Ma

This paper tackles spatial perception and manipulation challenges in Vision-Language-Action (VLA) models. To address depth ambiguity from monocular input, we leverage a pre-trained multi-view diffusion model to synthesize latent novel views…

In the last few years, deep neural networks opened the doors for big advances in novel view synthesis. Many of these approaches are based on a (coarse) proxy geometry obtained by structure from motion algorithms. Small deficiencies in this…

Computer Vision and Pattern Recognition · Computer Science 2023-11-09 Linus Franke , Darius Rückert , Laura Fink , Matthias Innmann , Marc Stamminger

We present a new deep point cloud rendering pipeline through multi-plane projections. The input to the network is the raw point cloud of a scene and the output are image or image sequences from a novel view or along a novel camera…

Computer Vision and Pattern Recognition · Computer Science 2020-06-26 Peng Dai , Yinda Zhang , Zhuwen Li , Shuaicheng Liu , Bing Zeng

Point clouds captured by different sensors such as RGB-D cameras and LiDAR possess non-negligible domain gaps. Most existing methods design different network architectures and train separately on point clouds from various sensors.…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Shengjun Zhang , Xin Fei , Yueqi Duan

Recent Vision-Language-Action (VLA) models have made impressive progress toward general-purpose robotic manipulation by post-training large Vision-Language Models (VLMs) for action prediction. Yet most VLAs entangle perception and control…

Vision-language-action (VLA) models perform well on training-seen robotic tasks but struggle to generalize to unseen scenes and objects. A key limitation lies in their implicit visual representations, which entangle object appearance,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Hanyu Zhou , Chuanhao Ma , Gim Hee Lee

Applying pre-trained models to assist point cloud understanding has recently become a mainstream paradigm in 3D perception. However, existing application strategies are straightforward, utilizing only the final output of the pre-trained…

Computer Vision and Pattern Recognition · Computer Science 2025-05-28 Yaohua Zha , Yanzi Wang , Hang Guo , Jinpeng Wang , Tao Dai , Bin Chen , Zhihao Ouyang , Xue Yuerong , Ke Chen , Shu-Tao Xia

Generating realistic and diverse LiDAR point clouds is crucial for autonomous driving simulation. Although previous methods achieve LiDAR point cloud generation from user inputs, they struggle to attain high-quality results while enabling…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Haiyun Wei , Fan Lu , Yunwei Zhu , Zehan Zheng , Weiyi Xue , Lin Shao , Xudong Zhang , Ya Wu , Rong Fu , Guang Chen

Robotic laboratories play a critical role in autonomous scientific discovery by enabling scalable, continuous experimental execution. Recent vision-language-action (VLA) models offer a promising foundation for robotic laboratories. However,…

Robotics · Computer Science 2026-02-11 Yiwen Pang , Bo Zhou , Changjin Li , Xuanhao Wang , Shengxiang Xu , Deng-Bao Wang , Min-Ling Zhang , Shimin Di

The 3D visual grounding task has been explored with visual and language streams comprehending referential language to identify target objects in 3D scenes. However, most existing methods devote the visual stream to capturing the 3D visual…

Computer Vision and Pattern Recognition · Computer Science 2022-11-28 Eslam Mohamed Bakr , Yasmeen Alsaedy , Mohamed Elhoseiny

In contrast to extensive studies on general vision, pre-training for scalable visual autonomous driving remains seldom explored. Visual autonomous driving applications require features encompassing semantics, 3D geometry, and temporal…

Computer Vision and Pattern Recognition · Computer Science 2024-01-01 Zetong Yang , Li Chen , Yanan Sun , Hongyang Li

Modern image encoders achieve high generalization by decoupling semantic meaning from resolution, an ability yet to be fully realized in the 3D domain. We investigate the failure of 3D point cloud encoders to achieve similar generalization…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Chun-Peng Chang , Shaoxiang Wang , Alain Pagani , Dariu Gavrila , Holger Caesar

Domain Adaptation (DA) approaches achieved significant improvements in a wide range of machine learning and computer vision tasks (i.e., classification, detection, and segmentation). However, as far as we are aware, there are few methods…

Computer Vision and Pattern Recognition · Computer Science 2019-11-26 Can Qin , Haoxuan You , Lichen Wang , C. -C. Jay Kuo , Yun Fu

Vision-Language-Action (VLA) models have shown remarkable achievements, driven by the rich implicit knowledge of their vision-language components. However, achieving generalist robotic agents demands precise grounding into physical…

Robotics · Computer Science 2025-07-15 Jialei Huang , Shuo Wang , Fanqi Lin , Yihang Hu , Chuan Wen , Yang Gao

Vision-Language-Action (VLA) models demonstrate remarkable potential for generalizable robotic manipulation. The performance of VLA models can be improved by integrating with action chunking, a critical technique for effective control.…

The development of practical applications, such as autonomous driving and robotics, has brought increasing attention to 3D point cloud understanding. While deep learning has achieved remarkable success on image-based tasks, there are many…

Computer Vision and Pattern Recognition · Computer Science 2021-05-25 Haoming Lu , Humphrey Shi

In this paper, we propose LiOn-XA, an unsupervised domain adaptation (UDA) approach that combines LiDAR-Only Cross-Modal (X) learning with Adversarial training for 3D LiDAR point cloud semantic segmentation to bridge the domain gap arising…

Computer Vision and Pattern Recognition · Computer Science 2024-10-22 Thomas Kreutz , Jens Lemke , Max Mühlhäuser , Alejandro Sanchez Guinea

Vision-Language-Action (VLA) models are emerging as a next-generation paradigm for robotics. We introduce dVLA, a diffusion-based VLA that leverages a multimodal chain-of-thought to unify visual perception, language reasoning, and robotic…

Robotics · Computer Science 2025-10-01 Junjie Wen , Minjie Zhu , Jiaming Liu , Zhiyuan Liu , Yicun Yang , Linfeng Zhang , Shanghang Zhang , Yichen Zhu , Yi Xu