中文
相关论文

相关论文: Text2Pos: Text-to-Point-Cloud Cross-Modal Localiza…

200 篇论文

Navigating drones through natural language commands remains challenging due to the dearth of accessible multi-modal datasets and the stringent precision requirements for aligning visual and textual data. To address this pressing need, we…

计算机视觉与模式识别 · 计算机科学 2024-08-01 Meng Chu , Zhedong Zheng , Wei Ji , Tingyu Wang , Tat-Seng Chua

This paper proposes a voxel-based approach for creating a digital twin of an urban environment that is capable of efficiently managing smart spaces. The paper explains the registration and localization procedure of the point cloud dataset,…

机器人学 · 计算机科学 2024-06-24 F. S. Mortazavi , O. Shkedova , U. Feuerhake , C. Brenner , M. Sester

Object transportation in cluttered environments is a fundamental task in various domains, including domestic service and warehouse logistics. In cooperative object transport, multiple robots must coordinate to move objects that are too…

机器人学 · 计算机科学 2025-10-13 Noah Steinkrüger , Nisarga Nilavadi , Wolfram Burgard , Tanja Katharina Kaiser

Robust localization in a given map is a crucial component of most autonomous robots. In this paper, we address the problem of localizing in an indoor environment that changes and where prominent structures have no correspondence in the map…

机器人学 · 计算机科学 2024-10-28 Nicky Zimmerman , Louis Wiesmann , Tiziano Guadagnino , Thomas Läbe , Jens Behley , Cyrill Stachniss

Image to point cloud global localization is crucial for robot navigation in GNSS-denied environments and has become increasingly important for multi-robot map fusion and urban asset management. The modality gap between images and point…

计算机视觉与模式识别 · 计算机科学 2024-12-23 Yuhao Li , Jianping Li , Zhen Dong , Yuan Wang , Bisheng Yang

Speech-to-text capabilities on mobile devices have proven helpful for hearing and speech accessibility, language translation, note-taking, and meeting transcripts. However, our foundational large-scale survey (n=263) shows that the…

人机交互 · 计算机科学 2025-03-06 Artem Dementyev , Dimitri Kanevsky , Samuel J. Yang , Mathieu Parvaix , Chiong Lai , Alex Olwal

Research connecting text and images has recently seen several breakthroughs, with models like CLIP, DALL-E 2, and Stable Diffusion. However, the connection between text and other visual modalities, such as lidar data, has received less…

计算机视觉与模式识别 · 计算机科学 2023-05-03 Georg Hess , Adam Tonderski , Christoffer Petersson , Kalle Åström , Lennart Svensson

In this research, we present an end-to-end data-driven pipeline for determining the long-term stability status of objects within a given environment, specifically distinguishing between static and dynamic objects. Understanding object…

计算机视觉与模式识别 · 计算机科学 2023-06-13 Ibrahim Hroob , Sergi Molina , Riccardo Polvara , Grzegorz Cielniak , Marc Hanheide

Text-to-point-cloud localization enables robots to understand spatial positions through natural language descriptions, which is crucial for human-robot collaboration in applications such as autonomous driving and last-mile delivery.…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Tianyi Shang , Zhenyu Li

Reorienting objects by using supports is a practical yet challenging manipulation task. Owing to the intricate geometry of objects and the constrained feasible motions of the robot, multiple manipulation steps are required for object…

机器人学 · 计算机科学 2023-08-30 Peng Xu , Hu Cheng , Jiankun Wang , Max Q. -H. Meng

We present Lang2Motion, a framework for language-guided point trajectory generation by aligning motion manifolds with joint embedding spaces. Unlike prior work focusing on human motion or video synthesis, we generate explicit trajectories…

计算机视觉与模式识别 · 计算机科学 2025-12-12 Bishoy Galoaa , Xiangyu Bai , Sarah Ostadabbas

Localizing objects in 3D scenes based on natural language requires understanding and reasoning about spatial relations. In particular, it is often crucial to distinguish similar objects referred by the text, such as "the left most chair"…

计算机视觉与模式识别 · 计算机科学 2022-11-18 Shizhe Chen , Pierre-Louis Guhur , Makarand Tapaswi , Cordelia Schmid , Ivan Laptev

Text-to-image (T2I) diffusion models excel at generating photorealistic images but often fail to render accurate spatial relationships. We identify two core issues underlying this common failure: 1) the ambiguous nature of data concerning…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Gaoyang Zhang , Bingtao Fu , Qingnan Fan , Qi Zhang , Runxing Liu , Hong Gu , Huaqi Zhang , Xinguo Liu

Text spotting in natural scene images is of great importance for many image understanding tasks. It includes two sub-tasks: text detection and recognition. In this work, we propose a unified network that simultaneously localizes and…

计算机视觉与模式识别 · 计算机科学 2021-06-29 Peng Wang , Hui Li , Chunhua Shen

Localization has been a challenging task for autonomous navigation. A loop detection algorithm must overcome environmental changes for the place recognition and re-localization of robots. Therefore, deep learning has been extensively…

机器人学 · 计算机科学 2023-04-19 Alex Junho Lee , Seungwon Song , Hyungtae Lim , Woojoo Lee , Hyun Myung

End-to-end speech translation (ST) is the task of translating speech signals in the source language into text in the target language. As a cross-modal task, end-to-end ST is difficult to train with limited data. Existing methods often try…

计算与语言 · 计算机科学 2023-05-26 Yan Zhou , Qingkai Fang , Yang Feng

We introduce Talk2Move, a reinforcement learning (RL) based diffusion framework for text-instructed spatial transformation of objects within scenes. Spatially manipulating objects in a scene through natural language poses a challenge for…

计算机视觉与模式识别 · 计算机科学 2026-01-09 Jing Tan , Zhaoyang Zhang , Yantao Shen , Jiarui Cai , Shuo Yang , Jiajun Wu , Wei Xia , Zhuowen Tu , Stefano Soatto

Place Recognition enables the estimation of a globally consistent map and trajectory by providing non-local constraints in Simultaneous Localisation and Mapping (SLAM). This paper presents Locus, a novel place recognition method using 3D…

机器人学 · 计算机科学 2022-09-27 Kavisha Vidanapathirana , Peyman Moghadam , Ben Harwood , Muming Zhao , Sridha Sridharan , Clinton Fookes

Large-scale Text-to-Image (T2I) diffusion models demonstrate significant generation capabilities based on textual prompts. Based on the T2I diffusion models, text-guided image editing research aims to empower users to manipulate generated…

计算机视觉与模式识别 · 计算机科学 2024-05-03 Chuanming Tang , Kai Wang , Fei Yang , Joost van de Weijer

Affordance detection and pose estimation are of great importance in many robotic applications. Their combination helps the robot gain an enhanced manipulation capability, in which the generated pose can facilitate the corresponding…

机器人学 · 计算机科学 2023-09-21 Toan Nguyen , Minh Nhat Vu , Baoru Huang , Tuan Van Vo , Vy Truong , Ngan Le , Thieu Vo , Bac Le , Anh Nguyen