English
Related papers

Related papers: NAUTILUS: A Large Multimodal Model for Underwater …

200 papers

Object detection models typically perform well on images captured in controlled environments with stable lighting, water clarity, and viewpoint, but their performance degrades substantially in real-world underwater settings characterized by…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Eleanor Wiesler , Trace Baxley

Unmanned Surface Vehicles (USVs) have emerged as a major platform in maritime operations, capable of supporting a wide range of applications. USVs can help reduce labor costs, increase safety, save energy, and allow for difficult unmanned…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Linh Trinh , Siegfried Mercelis , Ali Anwar

This paper introduces the first publicly accessible labeled multi-modal perception dataset for autonomous maritime navigation, focusing on in-water obstacles within the aquatic environment to enhance situational awareness for Autonomous…

Vision-language navigation (VLN) is a challenging task due to its large searching space in the environment. To address this problem, previous works have proposed some methods of fine-tuning a large model that pretrained on large-scale…

Computer Vision and Pattern Recognition · Computer Science 2022-03-09 Xiwen Liang , Fengda Zhu , Lingling Li , Hang Xu , Xiaodan Liang

The quality of supervised fine-tuning (SFT) data is crucial for the performance of large multimodal models (LMMs), yet current data enhancement methods often suffer from factual errors and hallucinations due to inadequate visual perception.…

Artificial Intelligence · Computer Science 2025-10-20 Tingqiao Xu , Ziru Zeng , Jiayu Chen

Multimodal large language models (MLLMs) deployed on devices must adapt to continuously changing visual scenarios such as variations in background and perspective, to effectively perform complex visual tasks. To investigate catastrophic…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Kai Jiang , Siqi Huang , Xiangyu Chen , Jiawei Shao , Hongyuan Zhang , Ping Luo , Xuelong Li

Autonomous and targeted underwater visual monitoring and exploration using Autonomous Underwater Vehicles (AUVs) can be a challenging task due to both online and offline constraints. The online constraints comprise limited onboard storage…

Marine biodiversity monitoring requires scalability and reliability across complex underwater environments to support conservation and invasive-species management. Yet existing detection solutions often exhibit a pronounced deployment gap,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Marco Piccolo , Qiwei Han , Astrid van Toor , Joachim Vanneste

Underwater image enhancement is such an important vision task due to its significance in marine engineering and aquatic robot. It is usually work as a pre-processing step to improve the performance of high level vision tasks such as…

Computer Vision and Pattern Recognition · Computer Science 2020-06-30 Long Chen , Lei Tong , Feixiang Zhou , Zheheng Jiang , Zhenyang Li , Jialin Lv , Junyu Dong , Huiyu Zhou

While Multimodal Large Language Models (MLLMs) have become adept at recognizing objects, they often lack the intuitive, human-like understanding of the world's underlying physical and social principles. This high-level vision-grounded…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Tianxiang Jiang , Sheng Xia , Yicheng Xu , Linquan Wu , Xiangyu Zeng , Limin Wang , Yu Qiao , Yi Wang

Construction sites are challenging environments for autonomous systems due to their unstructured nature and the presence of dynamic actors, such as workers and machinery. This work presents a comprehensive panoptic scene understanding…

Robotics · Computer Science 2024-10-08 Lorenzo Terenzi , Julian Nubert , Pol Eyschen , Pascal Roth , Simin Fei , Edo Jelavic , Marco Hutter

Autonomous Underwater Vehicles (AUVs) and Remotely Operated Vehicles (ROVs) demand robust spatial perception capabilities, including Simultaneous Localization and Mapping (SLAM), to support both remote and autonomous tasks. Vision-based…

Robotics · Computer Science 2025-06-10 Pushyami Kaveti , Ambjorn Grimsrud Waldum , Hanumant Singh , Martin Ludvigsen

Semantic scene understanding is crucial for robust and safe autonomous navigation, particularly so in off-road environments. Recent deep learning advances for 3D semantic segmentation rely heavily on large sets of training data, however…

Computer Vision and Pattern Recognition · Computer Science 2022-05-26 Peng Jiang , Philip Osteen , Maggie Wigness , Srikanth Saripalli

Visual instruction tuning has made considerable strides in enhancing the capabilities of Large Multimodal Models (LMMs). However, existing open LMMs largely focus on single-image tasks, their applications to multi-image scenarios remains…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Feng Li , Renrui Zhang , Hao Zhang , Yuanhan Zhang , Bo Li , Wei Li , Zejun Ma , Chunyuan Li

Navigation is a fundamental capability in embodied AI, representing the intelligence required to perceive and interact within physical environments following language instructions. Despite significant progress in large Vision-Language…

Large Vision-Language Models (LVLMs) have shown impressive capabilities across a range of tasks that integrate visual and textual understanding, such as image captioning and visual question answering. These models are trained on large-scale…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Xiaomei Zhang , Hanyu Zheng , Xiangyu Zhu , Jinghuan Wei , Junhong Zou , Zhen Lei , Zhaoxiang Zhang

Recent work in Machine Learning and Computer Vision has highlighted the presence of various types of systematic flaws inside ground truth object recognition benchmark datasets. Our basic tenet is that these flaws are rooted in the…

Computer Vision and Pattern Recognition · Computer Science 2023-07-27 Fausto Giunchiglia , Mayukh Bagchi , Xiaolei Diao

Visual Text-to-Speech (VTTS) aims to take the environmental image as the prompt to synthesize the reverberant speech for the spoken content. The challenge of this task lies in understanding the spatial environment from the image. Many…

Computer Vision and Pattern Recognition · Computer Science 2025-01-16 Rui Liu , Shuwei He , Yifan Hu , Haizhou Li

Metaverse has attracted great attention from industry and academia in recent years. Metaverse for the ocean (Meta-ocean) is the implementation of the Metaverse technologies in virtual emersion of the ocean which is beneficial for people…

Human-Computer Interaction · Computer Science 2023-08-14 Jinyu Li , Ping Hu , Weicheng Cui , Tianyi Huang , Shenghui Cheng

Multimodal large language models (MLLMs) have demonstrated impressive cross-domain capabilities, yet their proficiency in specialized scientific fields like marine biology remains underexplored. In this work, we systematically evaluate…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Faizan Farooq Khan , Yousef Radwan , Eslam Abdelrahman , Abdulwahab Felemban , Aymen Mir , Nico K. Michiels , Andrew J. Temple , Michael L. Berumen , Mohamed Elhoseiny
‹ Prev 1 4 5 6 7 8 10 Next ›