English
Related papers

Related papers: UAVD-Mamba: Deformable Token Fusion Vision Mamba f…

200 papers

Recently, deep learning models have achieved excellent performance in hyperspectral image (HSI) classification. Among the many deep models, Transformer has gradually attracted interest for its excellence in modeling the long-range…

Computer Vision and Pattern Recognition · Computer Science 2024-08-02 Lingbo Huang , Yushi Chen , Xin He

Diffusion models currently demonstrate impressive performance over various generative tasks. Recent work on image diffusion highlights the strong capabilities of Mamba (state space models) due to its efficient handling of long-range…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Jiaxu Liu , Li Li , Hubert P. H. Shum , Toby P. Breckon

Aerial object detection using unmanned aerial vehicles (UAVs) faces critical challenges including sub-10px targets, dense occlusions, and stringent computational constraints. Existing detectors struggle to balance accuracy and efficiency…

Computer Vision and Pattern Recognition · Computer Science 2025-04-22 Liu Wenbin

Accurate detection of Unmanned Aerial Vehicles (UAVs) is critical for surveillance, security, and airspace monitoring. However, existing datasets remain limited in scale, resolution, and the ability to capture objects across extreme size…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Yu-Hsi Chen

Accurate retinal vessel segmentation provides essential structural information for ophthalmic image analysis. However, existing methods struggle with challenges such as multi-scale vessel variability, complex curvatures, and ambiguous…

Image and Video Processing · Electrical Eng. & Systems 2025-04-21 Yihao Ouyang , Xunheng Kuang , Mengjia Xiong , Zhida Wang , Yuanquan Wang

Transformer-based architectures have become the backbone of both uni-modal and multi-modal foundation models, largely due to their scalability via attention mechanisms, resulting in a rich ecosystem of publicly available pre-trained models…

Computer Vision and Pattern Recognition · Computer Science 2025-10-10 Xiuwei Chen , Wentao Hu , Xiao Dong , Sihao Lin , Zisheng Chen , Meng Cao , Yina Zhuang , Jianhua Han , Hang Xu , Xiaodan Liang

Weakly supervised semantic segmentation offers a label-efficient solution to train segmentation models for volumetric medical imaging. However, existing approaches often rely on 2D encoders that neglect the inherent volumetric nature of the…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 Yiheng Lyu , Lian Xu , Mohammed Bennamoun , Farid Boussaid , Coen Arrow , Girish Dwivedi

Inspired by the excellent performance of Mamba networks, we propose a novel Deep Mamba Multi-modal Learning (DMML). It can be used to achieve the fusion of multi-modal features. We apply DMML to the field of multimedia retrieval and propose…

Multimedia · Computer Science 2024-06-27 Jian Zhu , Xin Zou , Yu Cui , Zhangmin Huang , Chenshu Hu , Bo Lyu

Extracting robust discriminative features is a critical challenge in person re-identification (ReID). While Transformer-based methods have successfully addressed some limitations of convolutional neural networks (CNNs), such as their local…

Computer Vision and Pattern Recognition · Computer Science 2025-11-12 Hongyang Gu , Qisong Yang , Lei Pu , Siming Han , Yao Ding

Multi-view 3D object detection is a crucial component of autonomous driving systems. Contemporary query-based methods primarily depend either on dataset-specific initialization of 3D anchors, introducing bias, or utilize dense attention…

Robotics · Computer Science 2024-11-12 Michelle Adeline , Junn Yong Loo , Vishnu Monn Baskaran

Tunnel construction using the drill-and-blast method requires the 3D measurement of the excavation front to evaluate underbreak locations. Considering the inspection and measurement task's safety, cost, and efficiency, deploying lightweight…

Robotics · Computer Science 2024-01-17 Zhefan Xu , Baihan Chen , Xiaoyang Zhan , Yumeng Xiu , Christopher Suzuki , Kenji Shimada

Accurate detection of obstacles in 3D is an essential task for autonomous driving and intelligent transportation. In this work, we propose a general multimodal fusion framework FusionPainting to fuse the 2D RGB image and 3D point clouds at…

Computer Vision and Pattern Recognition · Computer Science 2021-08-11 Shaoqing Xu , Dingfu Zhou , Jin Fang , Junbo Yin , Zhou Bin , Liangjun Zhang

Unmanned aerial vehicles (UAV)-based object detection with visible (RGB) and infrared (IR) images facilitates robust around-the-clock detection, driven by advancements in deep learning techniques and the availability of high-quality…

Computer Vision and Pattern Recognition · Computer Science 2025-10-16 Chen Chen , Kangcheng Bin , Ting Hu , Jiahao Qi , Xingyue Liu , Tianpeng Liu , Zhen Liu , Yongxiang Liu , Ping Zhong

Mamba, based on state space model (SSM) with its linear complexity and great success in classification provide its superiority in 3D point cloud analysis. Prior to that, Transformer has emerged as one of the most prominent and successful…

Computer Vision and Pattern Recognition · Computer Science 2024-06-11 Jia-wei Chen , Yu-jie Xiong , Yong-bin Gao

In recent years, deep learning has shown near-expert performance in segmenting complex medical tissues and tumors. However, existing models are often task-specific, with performance varying across modalities and anatomical regions.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-30 T-Mai Bui , Fares Bougourzi , Fadi Dornaika , Vinh Truong Hoang

Accurate segmentation of coronary arteries from computed tomography angiography (CTA) images is of paramount clinical importance for the diagnosis and treatment planning of cardiovascular diseases. However, coronary artery segmentation…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Xiaochan Yuan , Pai Zeng

Achieving both high accuracy and topological continuity in road segmentation from satellite imagery is a critical goal for applications ranging from urban planning to disaster response. State-of-the-art methods often rely on Vision…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Jules Decaestecker , Nicolas Vigne

Environmental perception with the multi-modal fusion of radar and camera is crucial in autonomous driving to increase accuracy, completeness, and robustness. This paper focuses on utilizing millimeter-wave (MMW) radar and camera sensor…

Computer Vision and Pattern Recognition · Computer Science 2022-11-01 Taohua Zhou , Yining Shi , Junjie Chen , Kun Jiang , Mengmeng Yang , Diange Yang

The development of multi-modal learning for Unmanned Aerial Vehicles (UAVs) typically relies on a large amount of pixel-aligned multi-modal image data. However, existing datasets face challenges such as limited modalities, high construction…

Computer Vision and Pattern Recognition · Computer Science 2025-02-14 Liang Yao , Fan Liu , Shengxiang Xu , Chuanyi Zhang , Xing Ma , Jianyu Jiang , Zequan Wang , Shimin Di , Jun Zhou

3-D object detection based on 4-D radar-vision is an important part in Internet of Vehicles (IoV). However, there are two challenges which need to be faced. First, the 4-D radar point clouds are sparse, leading to poor 3-D representation.…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Shucong Li , Xiaoluo Zhou , Yuqian He , Zhenyu Liu
‹ Prev 1 8 9 10 Next ›