中文

基于(偏)信息分解的特征选择中冗余与相关性的严格信息论定义

信息论 2023-05-05 v4 机器学习 math.IT

摘要

选择关于目标变量信息量最大且最小的特征集是机器学习与统计中的核心任务。信息论为表述特征选择算法提供了强大框架——然而,一个严格的、信息论的特征相关性定义,能够解释冗余与协同贡献等特征交互作用,仍然缺失。我们认为这一缺失源于经典信息论无法将一组变量关于目标提供的信息分解为独特、冗余与协同贡献。此类分解直至最近才由偏信息分解(PID)框架引入。利用 PID,我们阐明了为何特征选择在使用信息论处理时是一个概念上困难的问题,并给出了 PID 术语下特征相关性与冗余的新定义。由该定义,我们证明条件互信息(CMI)在最大化相关性的同时最小化冗余,并提出了一种基于 CMI 的迭代算法用于实际特征选择。我们在基准示例上展示了基于 CMI 的算法相较于无条件互信息的优势,并提供相应 PID 估计以凸显 PID 如何量化特征及其交互在特征选择问题中的信息贡献。

关键词

引用

@article{arxiv.2105.04187,
  title  = {A Rigorous Information-Theoretic Definition of Redundancy and Relevancy in Feature Selection Based on (Partial) Information Decomposition},
  author = {Patricia Wollstadt and Sebastian Schmitt and Michael Wibral},
  journal= {arXiv preprint arXiv:2105.04187},
  year   = {2023}
}

备注

44 pages, 12 figures. Reorganization and shortening of manuscript, added Appendix with theoretical guarantees, background information on the algorithm used, and an additional example application on a larger problem. Minor text editing