中文
相关论文

相关论文: NAUTILUS: A Large Multimodal Model for Underwater …

200 篇论文

Underwater image enhancement (UIE) presents a significant challenge within computer vision research. Despite the development of numerous UIE algorithms, a thorough and systematic review is still absent. To foster future advancements, we…

计算机视觉与模式识别 · 计算机科学 2026-02-26 Xiaofeng Cong , Yu Zhao , Jie Gui , Junming Hou , Dacheng Tao

Multimodal large language models (MLLMs) have achieved impressive performance across various tasks such as image captioning and visual question answer(VQA); however, they often struggle to accurately interpret depth information inherent in…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Hao Yang , Hongbo Zhang , Yanyan Zhao , Bing Qin

Depth information plays a crucial role in autonomous systems for environmental perception and robot state estimation. With the rapid development of deep neural network technology, depth estimation has been extensively studied and shown…

机器人学 · 计算机科学 2024-11-11 Quang Truong Nguyen , Thanh Nguyen Canh , Xiem HoangVan

Recently, multi-modality scene perception tasks, e.g., image fusion and scene understanding, have attracted widespread attention for intelligent vision systems. However, early efforts always consider boosting a single task unilaterally and…

计算机视觉与模式识别 · 计算机科学 2023-05-12 Zhu Liu , Jinyuan Liu , Guanyao Wu , Long Ma , Xin Fan , Risheng Liu

The powerful representation capacity of deep learning has made it inevitable for the underwater image enhancement community to employ its potential. The exploration of deep underwater image enhancement networks is increasing over time, and…

计算机视觉与模式识别 · 计算机科学 2019-07-19 Saeed Anwar , Chongyi Li

For aquaculture resource evaluation and ecological environment monitoring, automatic detection and identification of marine organisms is critical. However, due to the low quality of underwater images and the characteristics of underwater…

计算机视觉与模式识别 · 计算机科学 2022-05-23 Zheng Liu , Yaoming Zhuang , Pengrun Jia , Chengdong Wu , Hongli Xu ang Zhanlin Liu

Small-sized unmanned surface vehicles (USV) are coastal water devices with a broad range of applications such as environmental control and surveillance. A crucial capability for autonomous operation is obstacle detection for timely reaction…

计算机视觉与模式识别 · 计算机科学 2022-02-10 Borja Bovcon , Jon Muhovič , Duško Vranac , Dean Mozetič , Janez Perš , Matej Kristan

The visible-light camera, which is capable of environment perception and navigation assistance, has emerged as an essential imaging sensor for marine surface vessels in intelligent waterborne transportation systems (IWTS). However, the…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Ryan Wen Liu , Yuxu Lu , Yuan Gao , Yu Guo , Wenqi Ren , Fenghua Zhu , Fei-Yue Wang

Due to the selective absorption and scattering of light by diverse aquatic media, underwater images usually suffer from various visual degradations. Existing underwater image enhancement (UIE) approaches that combine underwater physical…

计算机视觉与模式识别 · 计算机科学 2025-08-04 Dazhao Du , Lingyu Si , Fanjiang Xu , Jianwei Niu , Fuchun Sun

Underwater image enhancement plays a crucial role in providing reliable visual information for underwater platforms, since strong absorption and scattering in water-related environments generally lead to image quality degradation. Existing…

计算机视觉与模式识别 · 计算机科学 2026-03-18 Yiqiang Zhou , Yifan Chen , Zhe Sun , Jijun Lu , Ye Zheng , Xuelong Li

Automated fetal ultrasound interpretation requires a workflow from visual perception, including plane recognition and anatomical segmentation, to clinical understanding, including biometric measurement and diagnostic reporting. However, the…

Automated waterway environment perception is crucial for enabling unmanned surface vessels (USVs) to understand their surroundings and make informed decisions. Most existing waterway perception models primarily focus on instance-level…

计算机视觉与模式识别 · 计算机科学 2025-07-02 Runwei Guan , Ningwei Ouyang , Tianhao Xu , Shaofeng Liang , Wei Dai , Yafeng Sun , Shang Gao , Songning Lai , Shanliang Yao , Xuming Hu , Ryan Wen Liu , Yutao Yue , Hui Xiong

In robotics and computer vision communities, extensive studies have been widely conducted regarding surveillance tasks, including human detection, tracking, and motion recognition with a camera. Additionally, deep learning algorithms are…

Understanding the deep semantics of images is essential in the era dominated by social media. However, current research works primarily on the superficial description of images, revealing a notable deficiency in the systematic investigation…

计算与语言 · 计算机科学 2024-06-21 Yixin Yang , Zheng Li , Qingxiu Dong , Heming Xia , Zhifang Sui

Underwater surveys provide long-term data for informing management strategies, monitoring coral reef health, and estimating blue carbon stocks. Advances in broad-scale survey methods, such as robotic underwater vehicles, have increased the…

计算机视觉与模式识别 · 计算机科学 2024-11-19 Scarlett Raine , Frederic Maire , Niko Suenderhauf , Tobias Fischer

In this paper, we present CaveSeg - the first visual learning pipeline for semantic segmentation and scene parsing for AUV navigation inside underwater caves. We address the problem of scarce annotated training data by preparing a…

机器人学 · 计算机科学 2024-05-13 A. Abdullah , T. Barua , R. Tibbetts , Z. Chen , M. J. Islam , I. Rekleitis

Due to the wavelength-dependent light attenuation, refraction and scattering, underwater images usually suffer from color distortion and blurred details. However, due to the limited number of paired underwater images with undistorted images…

图像与视频处理 · 电气工程与系统科学 2022-11-23 Qi Qi , Kunqian Li , Haiyong Zheng , Xiang Gao , Guojia Hou , Kun Sun

Underwater monocular depth estimation serves as the foundation for tasks such as 3D reconstruction of underwater scenes. However, due to the influence of light and medium, the underwater environment undergoes a distinctive imaging process,…

计算机视觉与模式识别 · 计算机科学 2024-07-26 Jian Wang , Jing Wang , Shenghui Rong , Bo He

Open-world 3D scene understanding is a critical challenge that involves recognizing and distinguishing diverse objects and categories from 3D data, such as point clouds, without relying on manual annotations. Traditional methods struggle…

计算机视觉与模式识别 · 计算机科学 2025-09-18 Yuru Wang , Pei Liu , Songtao Wang , Zehan Zhang , Xinyan Lu , Changwei Cai , Hao Li , Fu Liu , Peng Jia , Xianpeng Lang

Training multimodal large language models (MLLMs) for video understanding requires large-scale annotated data spanning diverse tasks such as object counting, question answering, and segmentation. However, collecting and annotating…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Tanzila Rahman , Renjie Liao , Leonid Sigal