English
Related papers

Related papers: SphereUFormer: A U-Shaped Transformer for Spherica…

200 papers

The shape of objects is an important source of visual information in a wide range of applications. One of the core challenges of shape quantification is to ensure that the extracted measurements remain invariant to transformations that…

Computer Vision and Pattern Recognition · Computer Science 2025-07-02 Anna Foix Romero , Craig Russell , Alexander Krull , Virginie Uhlmann

Accurate modelling of object deformations is crucial for a wide range of robotic manipulation tasks, where interacting with soft or deformable objects is essential. Current methods struggle to generalise to unseen forces or adapt to new…

Robotics · Computer Science 2025-05-20 Sean M. V. Collins , Brendan Tidd , Mahsa Baktashmotlagh , Peyman Moghadam

In this paper, we present an InSphereNet method for the problem of 3D object classification. Unlike previous methods that use points, voxels, or multi-view images as inputs of deep neural network (DNN), the proposed method constructs a…

Computer Vision and Pattern Recognition · Computer Science 2020-01-06 Hui Cao , Haikuan Du , Siyu Zhang , Shen Cai

Recently, Transformers have shown promising performance in various vision tasks. To reduce the quadratic computation complexity caused by the global self-attention, various methods constrain the range of attention within a local region to…

Computer Vision and Pattern Recognition · Computer Science 2021-12-30 Sitong Wu , Tianyi Wu , Haoru Tan , Guodong Guo

Very recently, Window-based Transformers, which computed self-attention within non-overlapping local windows, demonstrated promising results on image classification, semantic segmentation, and object detection. However, less study has been…

Computer Vision and Pattern Recognition · Computer Science 2021-06-08 Zilong Huang , Youcheng Ben , Guozhong Luo , Pei Cheng , Gang Yu , Bin Fu

We propose a novel transformer-based framework that reconstructs two high fidelity hands from multi-view RGB images. Unlike existing hand pose estimation methods, where one typically trains a deep network to regress hand model parameters…

Computer Vision and Pattern Recognition · Computer Science 2023-08-23 Tze Ho Elden Tse , Franziska Mueller , Zhengyang Shen , Danhang Tang , Thabo Beeler , Mingsong Dou , Yinda Zhang , Sasa Petrovic , Hyung Jin Chang , Jonathan Taylor , Bardia Doosti

Underwater scene reconstruction is essential for immersive exploration of aquatic environments, yet remains challenging due to complex participating-media effects such as absorption and scattering, as well as the limited field of view (FoV)…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Jiangbei Hu , Weichao Song , Shibo Yu , Mohan Wang , Zihan Yi , Rui Wu , Mingkang Xiang , Na Lei , Shengfa Wang , Zhongxuan Luo , Ying He

In this paper, we present a hybrid X-shaped vision Transformer, named Xformer, which performs notably on image denoising tasks. We explore strengthening the global representation of tokens from different scopes. In detail, we adopt two…

Computer Vision and Pattern Recognition · Computer Science 2024-02-27 Jiale Zhang , Yulun Zhang , Jinjin Gu , Jiahua Dong , Linghe Kong , Xiaokang Yang

The generation of immersive and navigable 3D environments is increasingly prevalent with the growing adoption of virtual reality and 3D content. However, recent methods face a fundamental limitation: they cannot produce 3D worlds that…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Antoine Schnepf , Karim Kassab , Flavian Vasile , Andrew Comport

Multi-task scene understanding aims to design models that can simultaneously predict several scene understanding tasks with one versatile model. Previous studies typically process multi-task features in a more local way, and thus cannot…

Computer Vision and Pattern Recognition · Computer Science 2023-06-09 Hanrong Ye , Dan Xu

Learning representations on large graphs is a long-standing challenge due to the inter-dependence nature. Transformers recently have shown promising performance on small graphs thanks to its global attention for capturing all-pair…

Machine Learning · Computer Science 2024-09-16 Qitian Wu , Kai Yang , Hengrui Zhang , David Wipf , Junchi Yan

Background and objective: High-resolution radiographic images play a pivotal role in the early diagnosis and treatment of skeletal muscle-related diseases. It is promising to enhance image quality by introducing single-image…

Image and Video Processing · Electrical Eng. & Systems 2023-12-29 Yongsong Huang , Tomo Miyazaki , Xiaofeng Liu , Kaiyuan Jiang , Zhengmi Tang , Shinichiro Omachi

We introduce a method to convert stereo 360{\deg} (omnidirectional stereo) imagery into a layered, multi-sphere image representation for six degree-of-freedom (6DoF) rendering. Stereo 360{\deg} imagery can be captured from multi-camera…

Computer Vision and Pattern Recognition · Computer Science 2020-08-18 Benjamin Attal , Selena Ling , Aaron Gokaslan , Christian Richardt , James Tompkin

4D millimeter-wave radar has emerged as a promising sensing modality for autonomous driving due to its robustness and affordability. However, its sparse and weak geometric cues make reliable instance activation difficult, limiting the…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Xiaokai Bai , Lianqing Zheng , Si-Yuan Cao , Xiaohan Zhang , Zhe Wu , Beinan Yu , Fang Wang , Jie Bai , Hui-Liang Shen

Three-dimensional shape sensing in soft and continuum robotics is a crucial aspect for stable actuation and control in fields such as Minimally Invasive surgery, as the estimation of complex curvatures while using continuum robotic tools is…

Signal Processing · Electrical Eng. & Systems 2024-03-26 Dalia Osman , Xinli Du , Timothy Minton , Yohan Noh

Transformer models rely on self-attention to capture token dependencies but face challenges in effectively integrating positional information while allowing multi-head attention (MHA) flexibility. Prior methods often model semantic and…

Machine Learning · Computer Science 2025-05-28 Jintian Shao , Hongyi Huang , Jiayi Wu , Beiwen Zhang , ZhiYu Wu , You Shan , MingKai Zheng

Lightweight image super-resolution (SR) methods aim at increasing the resolution and restoring the details of an image using a lightweight neural network. However, current lightweight SR methods still suffer from inferior performance and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-04 Jikai Wang , Huan Zheng , Jianbing Shen

In this paper, we propose a robust end-to-end multi-modal pipeline for place recognition where the sensor systems can differ from the map building to the query. Our approach operates directly on images and LiDAR scans without requiring any…

Robotics · Computer Science 2022-01-13 Lukas Bernreiter , Lionel Ott , Juan Nieto , Roland Siegwart , Cesar Cadena

In this paper, we propose a dense depth estimation pipeline for multiview 360{\deg} images. The proposed pipeline leverages a spherical camera model that compensates for radial distortion in 360{\deg} images. The key contribution of this…

Computer Vision and Pattern Recognition · Computer Science 2022-03-11 Seongyeop Yang , Kunhee Kim , Yeejin Lee

Transformer-based models have achieved strong performance in remote sensing image captioning by capturing long-range dependencies and contextual information. However, their practical deployment is hindered by high computational costs,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Swadhin Das , Divyansh Mundra , Priyanshu Dayal , Raksha Sharma