大规模众数识别与数据驱动科学
统计方法学
2016-11-10 v4 统计理论
统计理论
摘要
凸包搜寻(bump-hunting)或众数识别(mode identification)是几乎每个数据驱动发现科学领域中出现的根本问题。令人惊讶的是,极少有数据建模工具可用于自动(不需要手动逐个案例调查)、客观(非主观)且非参数(不基于限制性参数模型假设)的众数发现,并能扩展到大型数据集。本文介绍LPMode——一种基于检测概率密度多模态新理论的算法。我们应用LPMode来回答从环境科学、生态学、计量经济学、分析化学到天文学和癌症基因组学等各个领域中的重要研究问题。
引用
@article{arxiv.1509.06428,
title = {Large-Scale Mode Identification and Data-Driven Sciences},
author = {Subhadeep Mukhopadhyay},
journal= {arXiv preprint arXiv:1509.06428},
year = {2016}
}
备注
I would like to express my sincere thanks to the Editor and the anonymous reviewers for their in-depth comments, which have greatly improved the manuscript