English
Related papers

Related papers: TREET: TRansfer Entropy Estimation via Transformer…

200 papers

Transformers demonstrate competitive performance in terms of precision on the problem of vision-based object detection. However, they require considerable computational resources due to the quadratic size of the attention weights. In this…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Giorgos Savathrakis , Antonis Argyros

The Information Plane is a conceptual framework used to analyze the flow of information in neural networks, but traditional methods based on activations may not fully capture the dynamics of information processing. This paper introduces a…

Machine Learning · Computer Science 2024-08-28 Jaouad Dabounou , Amine Baazzouz

This technical report presents an effective method for motion prediction in autonomous driving. We develop a Transformer-based method for input encoding and trajectory prediction. Besides, we propose the Temporal Flow Header to enhance the…

Computer Vision and Pattern Recognition · Computer Science 2022-07-04 Yuting Wang , Hangning Zhou , Zhigang Zhang , Chen Feng , Huadong Lin , Chaofei Gao , Yizhi Tang , Zhenting Zhao , Shiyu Zhang , Jie Guo , Xuefeng Wang , Ziyao Xu , Chi Zhang

The great success of Transformer-based models benefits from the powerful multi-head self-attention mechanism, which learns token dependencies and encodes contextual information from the input. Prior work strives to attribute model decisions…

Computation and Language · Computer Science 2021-02-26 Yaru Hao , Li Dong , Furu Wei , Ke Xu

Adaptive physical and biological systems continually process fluctuating information from their environments. When the environment is nonstationary, inference itself becomes a nonequilibrium process with thermodynamic cost. We analyse a…

Statistical Mechanics · Physics 2026-03-23 Aditya Gupta

As transfer learning techniques are increasingly used to transfer knowledge from the source model to the target task, it becomes important to quantify which source models are suitable for a given target task without performing…

Our work combines aspects of three promising paradigms in machine learning, namely, attention mechanism, energy-based models, and associative memory. Attention is the power-house driving modern deep learning successes, but it lacks clear…

Next activity prediction in predictive business process monitoring is crucial for operational efficiency and informed decision-making. While machine learning and Artificial Intelligence have achieved promising results, challenges remain in…

Machine Learning · Computer Science 2026-04-08 Hadi Zare , Mostafa Abbasi , Maryam Ahang , Homayoun Najjaran

This work explores entropy analysis as a tool for probing information distribution within Transformer-based architectures. By quantifying token-level uncertainty and examining entropy patterns across different stages of processing, we aim…

Computation and Language · Computer Science 2025-07-31 Amedeo Buonanno , Alessandro Rivetti , Francesco A. N. Palmieri , Giovanni Di Gennaro , Gianmarco Romano

The major problem in information theoretic analysis of neural responses and other biological data is the reliable estimation of entropy--like quantities from small samples. We apply a recently introduced Bayesian entropy estimator to…

Data Analysis, Statistics and Probability · Physics 2009-09-29 Ilya Nemenman , William Bialek , Rob de Ruyter van Steveninck

Entropy estimation is of practical importance in information theory and statistical science. Many existing entropy estimators suffer from fast growing estimation bias with respect to dimensionality, rendering them unsuitable for…

Information Theory · Computer Science 2023-08-22 Ziqiao Ao , Jinglai Li

Sampling-based motion planning (SBMP) algorithms are renowned for their robust global search capabilities. However, the inherent randomness in their sampling mechanisms often result in inconsistent path quality and limited search…

Robotics · Computer Science 2024-10-27 Lei Zhuang , Jingdong Zhao , Yuntao Li , Zichun Xu , Liangliang Zhao , Hong Liu

Transfer learning across heterogeneous data distributions (a.k.a. domains) and distinct tasks is a more general and challenging problem than conventional transfer learning, where either domains or tasks are assumed to be the same. While…

Machine Learning · Computer Science 2021-03-26 Yang Tan , Yang Li , Shao-Lun Huang

We conducted empirical experiments to assess the transferability of a light curve transformer to datasets with different cadences and magnitude distributions using various positional encodings (PEs). We proposed a new approach to…

Vehicle arrival time prediction has been studied widely. With the emergence of IoT devices and deep learning techniques, estimated time of arrival (ETA) has become a critical component in intelligent transportation systems. Though many…

Machine Learning · Computer Science 2022-06-20 Hieu Tran , Son Nguyen , I-Ling Yen , Farokh Bastani

Inferring models, predicting the future, and estimating the entropy rate of discrete-time, discrete-event processes is well-worn ground. However, a much broader class of discrete-event processes operates in continuous-time. Here, we provide…

Statistical Mechanics · Physics 2020-05-11 S. E. Marzen , J. P. Crutchfield

Tensor Attention extends traditional attention mechanisms by capturing high-order correlations across multiple modalities, addressing the limitations of classical matrix-based attention. Meanwhile, Rotary Position Embedding…

Machine Learning · Computer Science 2024-12-25 Xiaoyu Li , Yingyu Liang , Zhenmei Shi , Zhao Song , Mingda Wan

Fine-tuning deep neural networks pre-trained on large scale datasets is one of the most practical transfer learning paradigm given limited quantity of training samples. To obtain better generalization, using the starting point as the…

Machine Learning · Computer Science 2022-02-28 Xingjian Li , Di Hu , Xuhong Li , Haoyi Xiong , Zhi Ye , Zhipeng Wang , Chengzhong Xu , Dejing Dou

While attention is all you need may be proving true, we do not know why: attention-based transformer models such as BERT are superior but how information flows from input tokens to output predictions are unclear. We introduce influence…

Computation and Language · Computer Science 2021-12-02 Kaiji Lu , Zifan Wang , Piotr Mardziel , Anupam Datta

Transformer is a new kind of neural architecture which encodes the input data as powerful features via the attention mechanism. Basically, the visual transformers first divide the input images into several local patches and then calculate…

Computer Vision and Pattern Recognition · Computer Science 2021-10-27 Kai Han , An Xiao , Enhua Wu , Jianyuan Guo , Chunjing Xu , Yunhe Wang