基于最大流的通用注意力流:Transformer模型的特征归因方法
机器学习
2025-04-29 v1 材料科学
人工智能
计算物理
摘要
本文介绍Generalized Attention Flow(GAF),一种用于Transformer模型的新型特征归因方法,以解决当前方法的局限性。通过扩展注意力流并用广义信息张量取代注意力权重,GAF将注意力权重、它们的梯度、最大流问题和屏障方法整合,以提高特征归因的性能。该方法具有关键理论属性,缓解了仅依赖简单注意力权重聚合的先前技术的局限。我们对序列分类任务进行全面基准测试,结果显示GAF的特定变体在大多数评估设置下 consistently优于最先进的特征归因方法,为Transformer模型输出提供更可靠的解释。
关键词
引用
@article{arxiv.2502.15764,
title = {High-Throughput Computational Screening and Interpretable Machine Learning of Metal-organic Frameworks for Iodine Capture},
author = {Haoyi Tan and Yukun Teng and Guangcun Shan},
journal= {arXiv preprint arXiv:2502.15764},
year = {2025}
}
备注
13 page,6 figures, submitted to npjCM