中文

大规模众数识别与数据驱动科学

统计方法学 2016-11-10 v4 统计理论 统计理论

摘要

凸包搜寻(bump-hunting)或众数识别(mode identification)是几乎每个数据驱动发现科学领域中出现的根本问题。令人惊讶的是,极少有数据建模工具可用于自动(不需要手动逐个案例调查)、客观(非主观)且非参数(不基于限制性参数模型假设)的众数发现,并能扩展到大型数据集。本文介绍LPMode——一种基于检测概率密度多模态新理论的算法。我们应用LPMode来回答从环境科学、生态学、计量经济学、分析化学到天文学和癌症基因组学等各个领域中的重要研究问题。

关键词

引用

@article{arxiv.1509.06428,
  title  = {Large-Scale Mode Identification and Data-Driven Sciences},
  author = {Subhadeep Mukhopadhyay},
  journal= {arXiv preprint arXiv:1509.06428},
  year   = {2016}
}

备注

I would like to express my sincere thanks to the Editor and the anonymous reviewers for their in-depth comments, which have greatly improved the manuscript