English
Related papers

Related papers: NanoMVG: USV-Centric Low-Power Multi-Task Visual G…

200 papers

Visual context provides grounding information for multimodal machine translation (MMT). However, previous MMT models and probing studies on visual features suggest that visual information is less explored in MMT as it is often redundant to…

Computer Vision and Pattern Recognition · Computer Science 2021-01-14 Dexin Wang , Deyi Xiong

Underwater docking is critical to enable the persistent operation of Autonomous Underwater Vehicles (AUVs). For this, the AUV must be capable of detecting and localizing the docking station, which is complex due to the highly dynamic…

Robotics · Computer Science 2024-10-28 Jalil Chavez-Galaviz , Jianwen Li , Matthew Bergman , Miras Mengdibayev , Nina Mahmoudian

We present Fast-Slow Transformer for Visually Grounding Speech, or FaST-VGS. FaST-VGS is a Transformer-based model for learning the associations between raw speech waveforms and visual images. The model unifies dual-encoder and…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-03 Puyuan Peng , David Harwath

Vision-Language Models (VLMs) enable multimodal reasoning for robotic perception and interaction, but their deployment in real-world systems remains constrained by latency, limited onboard resources, and privacy risks of cloud offloading.…

Robotics · Computer Science 2026-01-22 Sarat Ahmad , Maryam Hafeez , Syed Ali Raza Zaidi

Temporal Video Grounding (TVG) aims to localize a moment from an untrimmed video given the language description. Since the annotation of TVG is labor-intensive, TVG under limited supervision has accepted attention in recent years. The great…

Computer Vision and Pattern Recognition · Computer Science 2024-06-12 Xing Zhang , Jiaxi Gu , Haoyu Zhao , Shicong Wang , Hang Xu , Renjing Pei , Songcen Xu , Zuxuan Wu , Yu-Gang Jiang

Accurate terrain perception is essential for terrain-following flight of agricultural unmanned aerial vehicles (UAVs), yet remains challenging in real-world farmland due to occlusions, complex terrain geometry, and environmental…

Robotics · Computer Science 2026-05-05 Zhihao Zhan , Le Tao , Shaobin Li , Chenxin Fang , Xingrui Yang , Liang Li , Rui Fan , Yuhang Ming

Recently worldwide interest is growing toward commercial, military or scientific Unmanned Surface Vehicle (USV) and hence there is required to develop their guidance, navigation, and control (GNC) systems. Real USVs are a relatively new…

Robotics · Computer Science 2020-09-04 Pouyan Asgharian , Zati Hakim Azizul

We present GLM-5V-Turbo, a step toward native foundation models for multimodal agents. As foundation models are increasingly deployed in real environments, agentic capability depends not only on language reasoning, but also on the ability…

Unmanned ground vehicles have a huge development potential in both civilian and military fields, and have become the focus of research in various countries. In addition, high-precision, high-reliability sensors are significant for UGVs'…

Computer Vision and Pattern Recognition · Computer Science 2020-07-14 Qi Liu , Shihua Yuan , Zirui Li

Intelligent detection and tracking of the vessels on the sea play a significant role in conducting traffic avoidance in unmanned surface vessels(USV). Current traffic avoidance software relies mainly on Automated Identification System (AIS)…

Artificial Intelligence · Computer Science 2024-05-21 Srikanth Vemula , Eulises Franco , Michael Frye

In this paper, we explore a novel task named visual Relation Grounding in Videos (vRGV). The task aims at spatio-temporally localizing the given relations in the form of subject-predicate-object in the videos, so as to provide supportive…

Computer Vision and Pattern Recognition · Computer Science 2020-07-22 Junbin Xiao , Xindi Shang , Xun Yang , Sheng Tang , Tat-Seng Chua

Visual grounding aims to predict the locations of target objects specified by textual descriptions. For this task with linguistic and visual modalities, there is a latest research line that focuses on only selecting the linguistic-relevant…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Jingchao Wang , Wenlong Zhang , Dingjiang Huang , Hong Wang , Yefeng Zheng

Feed-forward reconstruction has been progressed rapidly, with the Visual Geometry Grounded Transformer (VGGT) being a notable baseline. However, directly applying VGGT to autonomous driving (AD) fails to capture three domain-specific…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Xiaosong Jia , Yanhao Liu , Yu Hong , Renqiu Xia , Junqi You , Bin Sun , Zhihui Hao , Junchi Yan

Navigation applications relying on the Global Navigation Satellite System (GNSS) are limited in indoor environments and GNSS-denied outdoor terrains such as dense urban or forests. In this paper, we present a novel accurate, robust and…

Robotics · Computer Science 2019-12-04 Qin Shi , Xiaowei Cui , Wei Li , Yu Xia , Mingquan Lu

Indoor navigation is challenging due to the absence of satellite positioning. This challenge is manifold greater for Visually Impaired People (VIPs) who lack the ability to get information from wayfinding signage. Other sensor signals…

Computer Vision and Pattern Recognition · Computer Science 2024-10-25 Jun Yu , Yifan Zhang , Badrinadh Aila , Vinod Namboodiri

Vision--language models (VLMs) achieve strong performance on many multimodal benchmarks but remain brittle on spatial reasoning tasks that require aligning abstract overhead representations with egocentric views. We introduce m2sv, a…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 Yosub Shin , Michael Buriek , Igor Molybog

We present a photo-realistic training and evaluation simulator (Sim4CV) with extensive applications across various fields of computer vision. Built on top of the Unreal Engine, the simulator integrates full featured physics based cars,…

Computer Vision and Pattern Recognition · Computer Science 2018-03-28 Matthias Müller , Vincent Casser , Jean Lahoud , Neil Smith , Bernard Ghanem

Different from Object Detection, Visual Grounding deals with detecting a bounding box for each text-image pair. This one box for each text-image data provides sparse supervision signals. Although previous works achieve impressive results,…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Weitai Kang , Gaowen Liu , Mubarak Shah , Yan Yan

While Unmanned Aerial Vehicles (UAVs) are increasingly deployed in several missions, their inability of reliable and consistent autonomous landing poses a major setback for deploying such systems truly autonomously. In this paper we present…

Robotics · Computer Science 2022-10-18 Michalis Piponidis , Panayiotis Aristodemou , Theocharis Theocharides

In recent years, photogrammetry has been widely used in many areas to create photorealistic 3D virtual data representing the physical environment. The innovation of small unmanned aerial vehicles (sUAVs) has provided additional…

Computer Vision and Pattern Recognition · Computer Science 2022-10-18 Meida Chen , Andrew Feng , Yu Hou , Kyle McCullough , Pratusha Bhuvana Prasad , Lucio Soibelman