English
Related papers

Related papers: Improving the Behaviour of Vision Transformers wit…

200 papers

We focus on building robustness in the convolutions of neural visual classifiers, especially against natural perturbations like elastic deformations, occlusions and Gaussian noise. Existing CNNs show outstanding performance on clean images,…

Computer Vision and Pattern Recognition · Computer Science 2021-10-25 Sadaf Gulshad , Ivan Sosnovik , Arnold Smeulders

Stochastic resonance shows that under some circumstances noise can enhance the response of a system to a periodic force. While this effect has been extensively investigated theoretically and demonstrated experimentally in classical systems,…

Quantum Physics · Physics 2009-11-06 S. F. Huelga , M. B. Plenio

Stochastic restoration algorithms allow to explore the space of solutions that correspond to the degraded input. In this paper we reveal additional fundamental advantages of stochastic methods over deterministic ones, which further motivate…

Image and Video Processing · Electrical Eng. & Systems 2024-05-21 Guy Ohayon , Theo Adrai , Michael Elad , Tomer Michaeli

Neural networks encode inputs as high-dimensional vectors, known as representations, that capture how models process data by encoding task-relevant structure and semantics. Representation alignment refers to the degree to which different…

Computational Geometry · Computer Science 2026-05-26 Xinyuan Yan , Rita Sevastjanova , Mennatallah El-Assady , Bei Wang

Effective token compression remains a critical challenge for scaling models to handle increasingly complex and diverse datasets. A novel mechanism based on contextual reinforcement is introduced, dynamically adjusting token importance…

Computation and Language · Computer Science 2025-08-11 Naderdel Piero , Zacharias Cromwell , Nathaniel Wainwright , Matthias Nethercott

We introduce Virtual Width Networks (VWN), a framework that delivers the benefits of wider representations without incurring the quadratic cost of increasing the hidden size. VWN decouples representational width from backbone width,…

Machine Learning · Computer Science 2025-11-18 Seed , Baisheng Li , Banggu Wu , Bole Ma , Bowen Xiao , Chaoyi Zhang , Cheng Li , Chengyi Wang , Chengyin Xu , Chi Zhang , Chong Hu , Daoguang Zan , Defa Zhu , Dongyu Xu , Du Li , Faming Wu , Fan Xia , Ge Zhang , Guang Shi , Haobin Chen , Hongyu Zhu , Hongzhi Huang , Huan Zhou , Huanzhang Dou , Jianhui Duan , Jianqiao Lu , Jianyu Jiang , Jiayi Xu , Jiecao Chen , Jin Chen , Jin Ma , Jing Su , Jingji Chen , Jun Wang , Jun Yuan , Juncai Liu , Jundong Zhou , Kai Hua , Kai Shen , Kai Xiang , Kaiyuan Chen , Kang Liu , Ke Shen , Liang Xiang , Lin Yan , Lishu Luo , Mengyao Zhang , Ming Ding , Mofan Zhang , Nianning Liang , Peng Li , Penghao Huang , Pengpeng Mu , Qi Huang , Qianli Ma , Qiyang Min , Qiying Yu , Renming Pang , Ru Zhang , Shen Yan , Shen Yan , Shixiong Zhao , Shuaishuai Cao , Shuang Wu , Siyan Chen , Siyu Li , Siyuan Qiao , Tao Sun , Tian Xin , Tiantian Fan , Ting Huang , Ting-Han Fan , Wei Jia , Wenqiang Zhang , Wenxuan Liu , Xiangzhong Wu , Xiaochen Zuo , Xiaoying Jia , Ximing Yang , Xin Liu , Xin Yu , Xingyan Bin , Xintong Hao , Xiongcai Luo , Xujing Li , Xun Zhou , Yanghua Peng , Yangrui Chen , Yi Lin , Yichong Leng , Yinghao Li , Yingshuan Song , Yiyuan Ma , Yong Shan , Yongan Xiang , Yonghui Wu , Yongtao Zhang , Yongzhen Yao , Yu Bao , Yuehang Yang , Yufeng Yuan , Yunshui Li , Yuqiao Xian , Yutao Zeng , Yuxuan Wang , Zehua Hong , Zehua Wang , Zengzhi Wang , Zeyu Yang , Zhengqiang Yin , Zhenyi Lu , Zhexi Zhang , Zhi Chen , Zhi Zhang , Zhiqi Lin , Zihao Huang , Zilin Xu , Ziyun Wei , Zuo Wang

Token interaction operation is one of the core modules in MLP-based models to exchange and aggregate information between different spatial locations. However, the power of token interaction on the spatial dimension is highly dependent on…

Computer Vision and Pattern Recognition · Computer Science 2023-07-24 Guiping Cao , Shengda Luo , Wenjian Huang , Xiangyuan Lan , Dongmei Jiang , Yaowei Wang , Jianguo Zhang

In this paper, we propose a simple yet effective transformer framework for self-supervised learning called DenseDINO to learn dense visual representations. To exploit the spatial information that the dense prediction tasks require but…

Computer Vision and Pattern Recognition · Computer Science 2023-06-09 Yike Yuan , Xinghe Fu , Yunlong Yu , Xi Li

Artificial neural networks can learn complex, salient data features to achieve a given task. On the opposite end of the spectrum, mathematically grounded methods such as topological data analysis allow users to design analysis pipelines…

Machine Learning · Computer Science 2022-12-29 Mattia G. Bergomi , Massimo Ferri , Alessandro Mella , Pietro Vertechi

Deep network architectures struggle to continually learn new tasks without forgetting the previous tasks. A recent trend indicates that dynamic architectures based on an expansion of the parameters can reduce catastrophic forgetting…

Computer Vision and Pattern Recognition · Computer Science 2022-08-09 Arthur Douillard , Alexandre Ramé , Guillaume Couairon , Matthieu Cord

Vision Transformers (ViTs) have demonstrated superior performance across a wide range of computer vision tasks. However, structured noise artifacts in their feature maps hinder downstream applications such as segmentation and depth…

Computer Vision and Pattern Recognition · Computer Science 2025-09-25 Sumit Mamtani

Transformers have become the predominant architecture in foundation models due to their excellent performance across various domains. However, the substantial cost of scaling these models remains a significant concern. This problem arises…

Incorporating stochasticity into the training process of deep convolutional networks is a widely used technique to reduce overfitting and improve regularization. Existing techniques often require modifying the architecture of the network by…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Evgeny Hershkovitch Neiterman , Gil Ben-Artzi

Diffusion transformers have gained substantial interest in diffusion generative modeling due to their outstanding performance. However, their computational demands, particularly the quadratic complexity of attention mechanisms and…

Machine Learning · Computer Science 2026-01-28 Jinming Lou , Wenyang Luo , Yufan Liu , Bing Li , Xinmiao Ding , Weiming Hu , Yuming Li , Chenguang Ma

Transformer-based document cross-encoder rerankers are a central component of modern information retrieval systems. Despite their success, these models suffer from high computational costs due to processing long query-document sequences at…

Information Retrieval · Computer Science 2026-05-22 Shengyao Zhuang , Zhichao Xu , Ivano Lauriola

Token compression techniques have recently emerged as powerful tools for accelerating Vision Transformer (ViT) inference in computer vision. Due to the quadratic computational complexity with respect to the token sequence length, these…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Phat Nguyen , Ngai-Man Cheung

The Vision Transformer has emerged as a powerful tool for image classification tasks, surpassing the performance of convolutional neural networks (CNNs). Recently, many researchers have attempted to understand the robustness of Transformers…

Computer Vision and Pattern Recognition · Computer Science 2023-12-18 Gihyun Kim , Juyeop Kim , Jong-Seok Lee

Production LLM systems often rely on separate models for safety and other classification-heavy steps, increasing latency, VRAM footprint, and operational complexity. We instead reuse computation already paid for by the serving LLM: we train…

Computation and Language · Computer Science 2026-04-28 Gonzalo Ariel Meyoyan , Luciano Del Corro

This paper investigates how to efficiently deploy vision transformers on edge devices for small workloads. Recent methods reduce the latency of transformer neural networks by removing or merging tokens, with small accuracy degradation.…

Machine Learning · Computer Science 2024-11-12 Nick John Eliopoulos , Purvish Jajal , James C. Davis , Gaowen Liu , George K. Thiravathukal , Yung-Hsiang Lu

Understanding the inner workings of Transformers is crucial for achieving more accurate and efficient predictions. In this work, we analyze the computation performed by Transformers in the layers after the top-1 prediction has become fixed,…

Computation and Language · Computer Science 2024-10-29 Daria Lioubashevski , Tomer Schlank , Gabriel Stanovsky , Ariel Goldstein