English
Related papers

Related papers: HexFormer: Hyperbolic Vision Transformer with Expo…

200 papers

Vision transformers have recently emerged as an effective alternative to convolutional networks for action recognition. However, vision transformers still struggle with geometric variations prevalent in video data. This paper proposes a…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Jinhui Ye , Jiaming Zhou , Hui Xiong , Junwei Liang

In recent years, the task of weakly supervised audio-visual violence detection has gained considerable attention. The goal of this task is to identify violent segments within multimodal data based on video-level labels. Despite advances in…

Computer Vision and Pattern Recognition · Computer Science 2024-02-14 Xiaogang Peng , Hao Wen , Yikai Luo , Xiao Zhou , Keyang Yu , Ping Yang , Zizhao Wu

This paper presents a new Vision Transformer (ViT) architecture Multi-Scale Vision Longformer, which significantly enhances the ViT of \cite{dosovitskiy2020image} for encoding high-resolution images using two techniques. The first is the…

Computer Vision and Pattern Recognition · Computer Science 2021-05-28 Pengchuan Zhang , Xiyang Dai , Jianwei Yang , Bin Xiao , Lu Yuan , Lei Zhang , Jianfeng Gao

Medical image segmentation is a cornerstone of modern clinical diagnostics. While Vision Transformers that leverage shifted window-based self-attention have established new benchmarks in this field, they are often hampered by a critical…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Fuchen Zheng , Xinyi Chen , Weixuan Li , Quanjun Li , Junhua Zhou , Xiaojiao Guo , Xuhang Chen , Chi-Man Pun , Shoujun Zhou

Visual environments are inherently hierarchical, as a panoramic view naturally encompasses and organizes multiple perspective views within its field. Capturing this hierarchy is crucial for effective perspective-to-equirectangular (P2E)…

Computer Vision and Pattern Recognition · Computer Science 2026-03-06 Suhan Woo , Seongwon Lee , Jinwoo Jang , Euntai Kim

Vision Transformer (ViT) models have demonstrated a breakthrough in a wide range of computer vision tasks. However, compared to the Convolutional Neural Network (CNN) models, it has been observed that the ViT models struggle to capture…

Computer Vision and Pattern Recognition · Computer Science 2023-09-04 Reza Azad , Amirhossein Kazerouni , Babak Azad , Ehsan Khodapanah Aghdam , Yury Velichko , Ulas Bagci , Dorit Merhof

Slot attention has emerged as a powerful framework for unsupervised object-centric learning, decomposing visual scenes into a small set of compact vector representations called \emph{slots}, each capturing a distinct region or object.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Neelu Madan , Àlex Pujol , Andreas Møgelmose , Sergio Escalera , Kamal Nasrollahi , Graham W. Taylor , Thomas B. Moeslund

Event cameras, with microsecond temporal resolution and high dynamic range (HDR) characteristics, emit high-speed event stream for perception tasks. Despite the recent advancement in GNN-based perception methods, they are prone to use…

Computer Vision and Pattern Recognition · Computer Science 2025-08-25 Haosheng Chen , Lian Luo , Mengjingcheng Mo , Zhanjie Wu , Guobao Xiao , Ji Gan , Jiaxu Leng , Xinbo Gao

Electroencephalography (EEG)-based brain-computer interfaces facilitate direct communication with a computer, enabling promising applications in human-computer interactions. However, their utility is currently limited because EEG decoding…

Machine Learning · Computer Science 2026-02-10 Shanglin Li , Shiwen Chu , Okan Koç , Yi Ding , Qibin Zhao , Motoaki Kawanabe , Ziheng Chen

Incomplete Multi-View Clustering (IMVC) faces the challenge of learning discriminative representations from fragmentary observations while maintaining robustness against missing views. However, prevalent Euclidean-based methods suffer from…

Machine Learning · Computer Science 2026-04-21 Tianyi Chen , Haobo Wang , Kai Tang , Gengyu Lyu , Tianlei Hu , Gang Chen , Hong Ma , Meixiang Xiang

Visual Grounding aims to localize the referring object in an image given a natural language expression. Recent advancements in DETR-based visual grounding methods have attracted considerable attention, as they directly predict the…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Yabing Wang , Zhuotao Tian , Qingpei Guo , Zheng Qin , Sanping Zhou , Ming Yang , Le Wang

With the achievements of Transformer in the field of natural language processing, the encoder-decoder and the attention mechanism in Transformer have been applied to computer vision. Recently, in multiple tasks of computer vision (image…

Computer Vision and Pattern Recognition · Computer Science 2022-03-03 Rui-Yang Ju , Ting-Yu Lin , Jen-Shiun Chiang , Jia-Hao Jian , Yu-Shian Lin , Liu-Rui-Yi Huang

Referring image segmentation is a fundamental vision-language task that aims to segment out an object referred to by a natural language expression from an image. One of the key challenges behind this task is leveraging the referring…

Computer Vision and Pattern Recognition · Computer Science 2022-04-07 Zhao Yang , Jiaqi Wang , Yansong Tang , Kai Chen , Hengshuang Zhao , Philip H. S. Torr

Vision Transformers (ViTs) have achieved overwhelming success, yet they suffer from vulnerable resolution scalability, i.e., the performance drops drastically when presented with input resolutions that are unseen during training. We…

Computer Vision and Pattern Recognition · Computer Science 2023-04-04 Rui Tian , Zuxuan Wu , Qi Dai , Han Hu , Yu Qiao , Yu-Gang Jiang

Learning in hyperbolic spaces has attracted increasing attention due to its superior ability to model hierarchical structures of data. Most existing hyperbolic learning methods use fixed distance measures for all data, assuming a uniform…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Pengxiang Li , Yuwei Wu , Zhi Gao , Xiaomeng Fan , Wei Wu , Zhipeng Lu , Yunde Jia , Mehrtash Harandi

Conventional Vision Transformer simplifies visual modeling by standardizing input resolutions, often disregarding the variability of natural visual data and compromising spatial-contextual fidelity. While preliminary explorations have…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Limeng Qiao , Yiyang Gan , Bairui Wang , Jie Qin , Shuang Xu , Siqi Yang , Lin Ma

Vision-based Bird's Eye View (BEV) representation is an emerging perception formulation for autonomous driving. The core challenge is to construct BEV space with multi-camera features, which is a one-to-many ill-posed problem. Diving into…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Yiming Wu , Ruixiang Li , Zequn Qin , Xinhai Zhao , Xi Li

Due to the emergency and homogenization of Artificial Intelligence (AI) technology development, transformer-based foundation models have revolutionized scientific applications, such as drug discovery, materials research, and astronomy.…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Huiwen Wu , Shuo Zhang , Yi Liu , Hongbin Ye

This paper presents ViewFormer, a simple yet effective model for multi-view 3d shape recognition and retrieval. We systematically investigate the existing methods for aggregating multi-view information and propose a novel ``view set"…

Computer Vision and Pattern Recognition · Computer Science 2023-05-02 Hongyu Sun , Yongcai Wang , Peng Wang , Xudong Cai , Deying Li

Despite the widespread adoption of transformers in medical applications, the exploration of multi-scale learning through transformers remains limited, while hierarchical representations are considered advantageous for computer-aided medical…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Xiaoya Tang , Bodong Zhang , Man Minh Ho , Beatrice S. Knudsen , Tolga Tasdizen