中文
相关论文

相关论文: Sector Patch Embedding: An Embedding Module Confor…

200 篇论文

3D Gaussian Splatting (3DGS) has enabled efficient 3D scene reconstruction from everyday images with real-time, high-fidelity rendering, greatly advancing VR/AR applications. Fisheye cameras, with their wider field of view (FOV), promise…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Zhengxian Yang , Fei Xie , Xutao Xue , Rui Zhang , Taicheng Huang , Yang Liu , Mengqi Ji , Tao Yu

Bird's Eye View (BEV) representations are tremendously useful for perception-related automated driving tasks. However, generating BEVs from surround-view fisheye camera images is challenging due to the strong distortions introduced by such…

计算机视觉与模式识别 · 计算机科学 2023-08-03 Ekta U. Samani , Feng Tao , Harshavardhan R. Dasari , Sihao Ding , Ashis G. Banerjee

Recent learning-based correction approaches in EPI estimate a displacement field, unwarp the reversed-PE image pair with the estimated field, and average the unwarped pair to yield a corrected image. Unsupervised learning in these…

图像与视频处理 · 电气工程与系统科学 2023-10-12 Abdallah Zaid Alkilani , Tolga Çukur , Emine Ulku Saritas

Pedestrian detection in images is a topic that has been studied extensively, but existing detectors designed for perspective images do not perform as successfully on images taken with top-view fisheye cameras, mainly due to the orientation…

计算机视觉与模式识别 · 计算机科学 2020-10-26 Sheng-Ho Chiang , Tsaipei Wang , Yi-Fu Chen

Transformers with powerful global relation modeling abilities have been introduced to fundamental computer vision tasks recently. As a typical example, the Vision Transformer (ViT) directly applies a pure transformer architecture on image…

计算机视觉与模式识别 · 计算机科学 2021-08-05 Xiaoyu Yue , Shuyang Sun , Zhanghui Kuang , Meng Wei , Philip Torr , Wayne Zhang , Dahua Lin

We introduce the notion of a Patch Sampling Schedule (PSS), that varies the number of Vision Transformer (ViT) patches used per batch during training. Since all patches are not equally important for most vision objectives (e.g.,…

计算机视觉与模式识别 · 计算机科学 2022-08-23 Bradley McDanel , Chi Phuong Huynh

Face anti-spoofing (FAS) plays a critical role in securing face recognition systems from different presentation attacks. Previous works leverage auxiliary pixel-level supervision and domain generalization approaches to address unseen spoof…

计算机视觉与模式识别 · 计算机科学 2022-03-29 Chien-Yi Wang , Yu-Ding Lu , Shang-Ta Yang , Shang-Hong Lai

With their motion-responsive nature, event-based cameras offer significant advantages over traditional cameras for optical flow estimation. While deep learning has improved upon traditional methods, current neural networks adopted for…

计算机视觉与模式识别 · 计算机科学 2025-04-16 Gokul Raju Govinda Raju , Nikola Zubić , Marco Cannici , Davide Scaramuzza

Detecting manipulated media has now become a pressing issue with the recent rise of deepfakes. Most existing approaches fail to generalize across diverse datasets and generation techniques. We thus propose a novel ensemble framework,…

计算机视觉与模式识别 · 计算机科学 2025-10-07 Vrushank Ahire , Aniruddh Muley , Shivam Zample , Siddharth Verma , Pranav Menon , Surbhi Madan , Abhinav Dhall

Line segment detection is essential for high-level tasks in computer vision and robotics. Currently, most stateof-the-art (SOTA) methods are dedicated to detecting straight line segments in undistorted pinhole images, thus distortions on…

计算机视觉与模式识别 · 计算机科学 2020-11-09 Hao Li , Huai Yu , Wen Yang , Lei Yu , Sebastian Scherer

Underwater Image Enhancement (UIE) is essential for robust visual perception in marine applications. However, existing methods predominantly rely on uniform mapping tailored to average dataset distributions, leading to over-processing…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Hang Xu , Chen Long , Bing Wang , Hao Chen , Zhen Dong

Recently, 3D Gaussian Splatting (3DGS) has garnered attention for its high fidelity and real-time rendering. However, adapting 3DGS to different camera models, particularly fisheye lenses, poses challenges due to the unique 3D to 2D…

计算机视觉与模式识别 · 计算机科学 2024-09-12 Zimu Liao , Siyan Chen , Rong Fu , Yi Wang , Zhongling Su , Hao Luo , Li Ma , Linning Xu , Bo Dai , Hengjie Li , Zhilin Pei , Xingcheng Zhang

Underwater vision suffers from severe effects due to selective attenuation and scattering when light propagates through water. Such degradation not only affects the quality of underwater images but limits the ability of vision tasks.…

计算机视觉与模式识别 · 计算机科学 2018-01-16 Chongyi Li , Jichang Guo , Chunle Guo

Image based localization is one of the important problems in computer vision due to its wide applicability in robotics, augmented reality, and autonomous systems. There is a rich set of methods described in the literature how to…

计算机视觉与模式识别 · 计算机科学 2017-12-12 Pulak Purkait , Cheng Zhao , Christopher Zach

Many telepresence robots are equipped with a forward-facing camera for video communication and a downward-facing camera for navigation. In this paper, we propose to stitch videos from the FF-camera with a wide-angle lens and the DF-camera…

机器人学 · 计算机科学 2019-03-18 Yanmei Dong , Mingtao Pei , Lijia Zhang , Bin Xu , Yuwei Wu , Yunde Jia

Current Transformer-based methods for small object detection continue emerging, yet they have still exhibited significant shortcomings. This paper introduces HeatMap Position Embedding (HMPE), a novel Transformer Optimization technique that…

计算机视觉与模式识别 · 计算机科学 2025-04-21 YangChen Zeng

Depth estimation is a critical technology in autonomous driving, and multi-camera systems are often used to achieve a 360$^\circ$ perception. These 360$^\circ$ camera sets often have limited or low-quality overlap regions, making multi-view…

计算机视觉与模式识别 · 计算机科学 2024-04-03 Jialei Xu , Wei Yin , Dong Gong , Junjun Jiang , Xianming Liu

Video semantic segmentation is active in recent years benefited from the great progress of image semantic segmentation. For such a task, the per-frame image segmentation is generally unacceptable in practice due to high computation cost. To…

计算机视觉与模式识别 · 计算机科学 2020-11-13 Jiafan Zhuang , Zilei Wang , Bingke Wang

Objective: Long-axial field-of-view (LAFOV) positron emission tomography (PET) systems allow higher sensitivity, with an increased number of detected lines of response induced by a larger angle of acceptance. However, this extended angle…

A novel Face Pyramid Vision Transformer (FPVT) is proposed to learn a discriminative multi-scale facial representations for face recognition and verification. In FPVT, Face Spatial Reduction Attention (FSRA) and Dimensionality Reduction…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Khawar Islam , Muhammad Zaigham Zaheer , Arif Mahmood