English
Related papers

Related papers: 3D Move to See: Multi-perspective visual servoing …

200 papers

Multimodal Large Language Models (MLLMs) exhibit impressive capabilities across a variety of tasks, especially when equipped with carefully designed visual prompts. However, existing studies primarily focus on logical reasoning and visual…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Dingning Liu , Cheng Wang , Peng Gao , Renrui Zhang , Xinzhu Ma , Yuan Meng , Zhihui Wang

In autonomous driving, 3D object detection is essential for accurately identifying and tracking objects. Despite the continuous development of various technologies for this task, a significant drawback is observed in most of them-they…

Computer Vision and Pattern Recognition · Computer Science 2025-02-05 Hsin-Cheng Lu , Chung-Yi Lin , Winston H. Hsu

Combining 3D vision with tactile sensing could unlock a greater level of dexterity for robots and improve several manipulation tasks. However, obtaining a close-up 3D view of the location where manipulation contacts occur can be…

Robotics · Computer Science 2023-03-14 Etienne Roberge , Guillaume Fornes , Jean-Philippe Roberge

Multi-view photometric stereo (MVPS) is a preferred method for detailed and precise 3D acquisition of an object from images. Although popular methods for MVPS can provide outstanding results, they are often complex to execute and limited to…

Computer Vision and Pattern Recognition · Computer Science 2022-10-17 Berk Kaya , Suryansh Kumar , Carlos Oliveira , Vittorio Ferrari , Luc Van Gool

Understanding semantics and dynamics has been crucial for embodied agents in various tasks. Both tasks have much more data redundancy than the static scene understanding task. We formulate the view selection problem as an active learning…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Yiqian Li , Wen Jiang , Kostas Daniilidis

Accurate localization and 3D maps are increasingly needed for various artificial intelligence based IoT applications such as augmented reality, intelligent transportation, crowd monitoring, robotics, etc. This article proposes a novel…

Robotics · Computer Science 2021-03-23 Max Jwo Lem Lee , Li-Ta Hsu

What is a good visual representation for autonomous agents? We address this question in the context of semantic visual navigation, which is the problem of a robot finding its way through a complex environment to a target object, e.g. go to…

Computer Vision and Pattern Recognition · Computer Science 2019-07-04 Arsalan Mousavian , Alexander Toshev , Marek Fiser , Jana Kosecka , Ayzaan Wahid , James Davidson

The recent advances in 3D Gaussian Splatting (3DGS) show promising results on the novel view synthesis (NVS) task. With its superior rendering performance and high-fidelity rendering quality, 3DGS is excelling at its previous NeRF…

Computer Vision and Pattern Recognition · Computer Science 2024-10-30 Yu Chen , Gim Hee Lee

The motivation of this paper is to develop a smart system using multi-modal vision for next-generation mechanical assembly. It includes two phases where in the first phase human beings teach the assembly structure to a robot and in the…

Robotics · Computer Science 2016-01-27 Weiwei Wan , Feng Lu , Zepei Wu , Kensuke Harada

This paper proposes a new approach to achieve direct visual servoing (DVS) based on discrete orthogonal moments (DOMs). DVS is performed in such a way that the extraction of geometric primitives, matching, and tracking steps in the…

Robotics · Computer Science 2024-02-14 Yuhan Chen , Max Q. -H. Meng , Li Liu

In this work, we present a method for tracking and learning the dynamics of all objects in a large scale robot environment. A mobile robot patrols the environment and visits the different locations one by one. Movable objects are discovered…

Robotics · Computer Science 2018-01-30 Nils Bore , Patric Jensfelt , John Folkesson

This paper presents a strategy to guide a mobile ground robot equipped with a camera or depth sensor, in order to autonomously map the visible part of a bounded three-dimensional structure. We describe motion planning algorithms that…

Robotics · Computer Science 2017-11-15 Manikandasriram Srinivasan Ramanagopal , André Phu-Van Nguyen , Jerome Le Ny

We introduce a test-time framework for multiview Transformers (MVTs) that incorporates priors (e.g., camera poses, intrinsics, and depth) to improve 3D tasks without retraining or modifying pre-trained image-only networks. Rather than…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Lei Zhou , Haoyu Wu , Akshat Dave , Dimitris Samaras

Text-to-3D is an emerging task that allows users to create 3D content with infinite possibilities. Existing works tackle the problem by optimizing a 3D representation with guidance from pre-trained diffusion models. An apparent drawback is…

Computer Vision and Pattern Recognition · Computer Science 2023-06-06 Yiji Cheng , Fei Yin , Xiaoke Huang , Xintong Yu , Jiaxiang Liu , Shikun Feng , Yujiu Yang , Yansong Tang

We study how the choice of visual perspective affects learning and generalization in the context of physical manipulation from raw sensor observations. Compared with the more commonly used global third-person perspective, a hand-centric…

Robotics · Computer Science 2022-03-25 Kyle Hsu , Moo Jin Kim , Rafael Rafailov , Jiajun Wu , Chelsea Finn

Human decision-making often relies on visual information from multiple perspectives or views. In contrast, machine learning-based object recognition utilizes information from a single image of the object. However, the information conveyed…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Mona Alzahrani , Muhammad Usman , Salma Kammoun , Saeed Anwar , Tarek Helmy

This paper presents a method for constrained motion planning from vision, which enables a robot to move its end-effector over an observed surface, given start and destination points. The robot has no prior knowledge of the surface shape,…

Robotics · Computer Science 2020-06-23 T. Pardi , V. Ortenzi , C. Fairbairn , T. Pipe , A. M. Ghalamzan E. , R. Stolkin

This paper investigates coverage control for visual sensor networks based on gradient descent techniques on matrix manifolds. We consider the scenario that networked vision sensors with controllable orientations are distributed over 3-D…

Systems and Control · Computer Science 2013-09-24 Takeshi Hatanaka , Riku Funada , Masayuki Fujita

Nowadays robots play an increasingly important role in our daily life. In human-centered environments, robots often encounter piles of objects, packed items, or isolated objects. Therefore, a robot must be able to grasp and manipulate…

Robotics · Computer Science 2022-10-06 Hamidreza Kasaei , Mohammadreza Kasaei

Recent camera-based 3D object detection methods have introduced sequential frames to improve the detection performance hoping that multiple frames would mitigate the large depth estimation error. Despite improved detection performance,…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Sanmin Kim , Youngseok Kim , In-Jae Lee , Dongsuk Kum