中文

基于最大流的通用注意力流:Transformer模型的特征归因方法

机器学习 2025-04-29 v1 材料科学 人工智能 计算物理

摘要

本文介绍Generalized Attention Flow(GAF),一种用于Transformer模型的新型特征归因方法,以解决当前方法的局限性。通过扩展注意力流并用广义信息张量取代注意力权重,GAF将注意力权重、它们的梯度、最大流问题和屏障方法整合,以提高特征归因的性能。该方法具有关键理论属性,缓解了仅依赖简单注意力权重聚合的先前技术的局限。我们对序列分类任务进行全面基准测试,结果显示GAF的特定变体在大多数评估设置下 consistently优于最先进的特征归因方法,为Transformer模型输出提供更可靠的解释。

关键词

引用

@article{arxiv.2502.15764,
  title  = {High-Throughput Computational Screening and Interpretable Machine Learning of Metal-organic Frameworks for Iodine Capture},
  author = {Haoyi Tan and Yukun Teng and Guangcun Shan},
  journal= {arXiv preprint arXiv:2502.15764},
  year   = {2025}
}

备注

13 page,6 figures, submitted to npjCM