中文
相关论文

相关论文: STING-BEE: Towards Vision-Language Model for Real-…

200 篇论文

Intelligent inspection robots are widely used in substation patrol inspection, which can help check potential safety hazards by patrolling the substation and sending back scene images. However, when patrolling some marginal areas with weak…

计算机视觉与模式识别 · 计算机科学 2024-04-16 Senran Fan , Haotai Liang , Chen Dong , Xiaodong Xu , Geng Liu

Bird's-Eye-View (BEV) perception has become a vital component of autonomous driving systems due to its ability to integrate multiple sensor inputs into a unified representation, enhancing performance in various downstream tasks. However,…

机器人学 · 计算机科学 2024-10-10 Yuxin Li , Yiheng Li , Xulei Yang , Mengying Yu , Zihang Huang , Xiaojun Wu , Chai Kiat Yeo

Laboratories are prone to severe injuries from minor unsafe actions, yet continuous safety monitoring -- beyond mandatory pre-lab safety training -- is limited by human availability. Vision language models (VLMs) offer promise for…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Trishna Chakraborty , Udita Ghosh , Aldair Ernesto Gongora , Ruben Glatt , Yue Dong , Jiachen Li , Amit K. Roy-Chowdhury , Chengyu Song

Dynamic scene understanding remains a persistent challenge in robotic applications. Early dynamic mapping methods focused on mitigating the negative influence of short-term dynamic objects on camera motion estimation by masking or tracking…

机器人学 · 计算机科学 2024-12-04 Chenguang Huang , Shengchao Yan , Wolfram Burgard

Existing Earth Vision datasets are either suitable for semantic segmentation or object detection. In this work, we introduce the first benchmark dataset for instance segmentation in aerial imagery that combines instance-level object…

计算机视觉与模式识别 · 计算机科学 2019-08-29 Syed Waqas Zamir , Aditya Arora , Akshita Gupta , Salman Khan , Guolei Sun , Fahad Shahbaz Khan , Fan Zhu , Ling Shao , Gui-Song Xia , Xiang Bai

We propose CARE (Collision Avoidance via Repulsive Estimation) to improve the robustness of learning-based visual navigation methods. Recently, visual navigation models, particularly foundation models, have demonstrated promising…

机器人学 · 计算机科学 2025-08-11 Joonkyung Kim , Joonyeol Sim , Woojun Kim , Katia Sycara , Changjoo Nam

Semantic Scene Completion (SSC) is essential for 3D perception in mobile robotics, as it enables holistic scene understanding by jointly estimating dense volumetric occupancy and per-voxel semantics. Although SSC has been widely studied in…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Markus Gross , Sai B. Matha , Aya Fahmy , Rui Song , Daniel Cremers , Henri Meess

In this cloud-dependent era, various security techniques, such as encryption, steganography, and hybrid approaches, have been utilized in cloud computing to enhance security, maintain enormous storage capacity, and provide ease of access.…

密码学与安全 · 计算机科学 2025-05-26 Naima Sultana Ayesha , Mehrin Anannya , Md Biplob Hosen , Rashed Mazumder

Traffic forecasting, crucial for urban planning, requires accurate predictions of spatial-temporal traffic patterns across urban areas. Existing research mainly focuses on designing complex models that capture spatial-temporal dependencies…

机器学习 · 计算机科学 2024-07-30 Jiarui Sun , Yujie Fan , Chin-Chia Michael Yeh , Wei Zhang , Girish Chowdhary

The success of Emergency Response (ER) scenarios, such as search and rescue, is often dependent upon the prompt location of a lost or injured person. With the increasing use of small Unmanned Aerial Systems (sUAS) as "eyes in the sky"…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Arturo Miguel Russell Bernal , Jane Cleland-Huang , Walter Scheirer

Accurate point tracking in surgical environments remains challenging due to complex visual conditions, including smoke occlusion, specular reflections, and tissue deformation. While existing surgical tracking datasets provide coordinate…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Rulin Zhou , Wenlong He , An Wang , Jianhang Zhang , Xuanhui Zeng , Xi Zhang , Chaowei Zhu , Haijun Hu , Hongliang Ren

Synthetic medical data generation has opened up new possibilities in the healthcare domain, offering a powerful tool for simulating clinical scenarios, enhancing diagnostic and treatment quality, gaining granular medical knowledge, and…

图像与视频处理 · 电气工程与系统科学 2024-05-01 Hyungyung Lee , Da Young Lee , Wonjae Kim , Jin-Hwa Kim , Tackeun Kim , Jihang Kim , Leonard Sunwoo , Edward Choi

In autonomous driving, addressing occlusion scenarios is crucial yet challenging. Robust surrounding perception is essential for handling occlusions and aiding motion planning. State-of-the-art models fuse Lidar and Camera data to produce…

This work targets what we consider to be the foundational step for urban airborne robots, a safe landing. Our attention is directed toward what we deem the most crucial aspect of the safe landing perception stack: segmentation. We present a…

机器人学 · 计算机科学 2024-10-16 Haechan Mark Bong , Rongge Zhang , Ricardo de Azambuja , Giovanni Beltrame

Background & Purpose: Chest X-Ray (CXR) use in pre-MRI safety screening for Lead-Less Implanted Electronic Devices (LLIEDs), easily overlooked or misidentified on a frontal view (often only acquired), is common. Although most LLIED types…

图像与视频处理 · 电气工程与系统科学 2022-04-28 Mutlu Demirer , Richard D. White , Vikash Gupta , Ronnie A. Sebro , Barbaros S. Erdal

Although recent traffic benchmarks have advanced multimodal data analysis, they generally lack systematic evaluation aligned with official safety standards. To fill this gap, we introduce RoadSafe365, a large-scale vision-language benchmark…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Xinyu Liu , Darryl C. Jacob , Yuxin Liu , Xinsong Du , Muchao Ye , Bolei Zhou , Pan He

Bird's-eye-view (BEV) is a powerful and widely adopted representation for road scenes that captures surrounding objects and their spatial locations, along with overall context in the scene. In this work, we focus on bird's eye semantic…

计算机视觉与模式识别 · 计算机科学 2020-06-24 Mong H. Ng , Kaahan Radia , Jianfei Chen , Dequan Wang , Ionel Gog , Joseph E. Gonzalez

Existing X-ray based pre-trained vision models are usually conducted on a relatively small-scale dataset (less than 500k samples) with limited resolution (e.g., 224 $\times$ 224). However, the key to the success of self-supervised…

图像与视频处理 · 电气工程与系统科学 2024-04-30 Xiao Wang , Yuehang Li , Wentao Wu , Jiandong Jin , Yao Rong , Bo Jiang , Chuanfu Li , Jin Tang

Masked token prediction has emerged as a powerful pre-training objective across language, vision, and speech, offering the potential to unify these diverse modalities through a single pre-training task. However, its application for general…

Visual Question Answering (VQA) research seeks to create AI systems to answer natural language questions in images, yet VQA methods often yield overly simplistic and short answers. This paper aims to advance the field by introducing Visual…

计算机视觉与模式识别 · 计算机科学 2024-11-13 Jialu Li , Manish Kumar Thota , Ruslan Gokhman , Radek Holik , Youshan Zhang
‹ 上一页 1 8 9 10 下一页 ›