中文
相关论文

相关论文: An automatic approach to exclude interlopers from …

200 篇论文

Finding a set of nested partitions of a dataset is useful to uncover relevant structure at different scales, and is often dealt with a data-dependent methodology. In this paper, we introduce a general two-step methodology for model-based…

统计计算 · 统计学 2021-04-22 Etienne Côme , Nicolas Jouvin , Pierre Latouche , Charles Bouveyron

The problem of selecting a small, yet high quality subset of patterns from a larger collection of itemsets has recently attracted lot of research. Here we discuss an approach to this problem using the notion of decomposable families of…

机器学习 · 计算机科学 2020-06-18 Nikolaj Tatti , Hannes Heikinheimo

One of the great challenges of modern science is to faithfully model, and understand, matter at a wide range of scales. Starting with atoms, the vastness of the space of possible configurations poses a formidable challenge to any simulation…

材料科学 · 物理学 2017-08-28 Sebastian E. Ahnert , William P. Grant , Chris J. Pickard

We consider the problem of fast time-series data clustering. Building on previous work modeling the correlation-based Hamiltonian of spin variables we present an updated fast non-expensive Agglomerative Likelihood Clustering algorithm…

计算金融 · 定量金融 2022-03-22 Lionel Yelibi , Tim Gebbie

We generalize the setting of online clustering of bandits by allowing non-uniform distribution over user frequencies. A more efficient algorithm is proposed with simple set structures to represent clusters. We prove a regret bound for the…

机器学习 · 计算机科学 2019-07-03 Shuai Li , Wei Chen , Shuai Li , Kwong-Sak Leung

Clustering analysis identifies samples as groups based on either their mutual closeness or homogeneity. In order to detect clusters in arbitrary shapes, a novel and generic solution based on boundary erosion is proposed. The clusters are…

计算机视觉与模式识别 · 计算机科学 2018-04-16 Cheng-Hao Deng , Wan-Lei Zhao

We present a new fast online clustering algorithm that reliably recovers arbitrary-shaped data clusters in high throughout data streams. Unlike the existing state-of-the-art online clustering methods based on k-means or k-medoid, it does…

人工智能 · 计算机科学 2015-06-11 Krzysztof Choromanski , Sanjiv Kumar , Xiaofeng Liu

In order to fully understand the shapes of asteroids families in the 3-dimensional space of the proper elements $(a_{\rm p}, e_{\rm p}, \sin I_{\rm p})$ it is necessary to compare observed asteroids with N-body simulations. To this point,…

地球与行星天体物理 · 物理学 2018-10-10 Miroslav Brož , Alessandro Morbidelli

We have proposed a model based upon flocking on a complex network, and then developed two clustering algorithms on the basis of it. In the algorithms, firstly a \textit{k}-nearest neighbor (knn) graph as a weighted and directed graph is…

机器学习 · 计算机科学 2008-12-31 Qiang Li , Yan He , Jing-ping Jiang

We report on the results of a systematic search for associated asteroid families for all active asteroids known to date. We find that 10 out of 12 main-belt comets (MBCs) and 5 out of 7 disrupted asteroids are linked with known or candidate…

地球与行星天体物理 · 物理学 2018-02-14 Henry H. Hsieh , Bojan Novakovic , Yoonyoung Kim , Ramon Brasser

We introduce a technique to filter out complex data-sets by extracting a subgraph of representative links. Such a filtering can be tuned up to any desired level by controlling the genus of the resulting graph. We show that this technique is…

无序系统与神经网络 · 物理学 2007-05-23 M. Tumminello , T. Aste , T. Di Matteo , R. N. Mantegna

A good clustering can help a data analyst to explore and understand a data set, but what constitutes a good clustering may depend on domain-specific and application-specific criteria. These criteria can be difficult to formalize, even when…

机器学习 · 统计学 2016-06-21 Akash Srivastava , James Zou , Ryan P. Adams , Charles Sutton

In many applications of clustering (for example, ontologies or clusterings of animal or plant species), hierarchical clusterings are more descriptive than a flat clustering. A hierarchical clustering over $n$ elements is represented by a…

数据结构与算法 · 计算机科学 2018-04-18 Ehsan Emamjomeh-Zadeh , David Kempe

In this paper, we propose a technique for time series clustering using community detection in complex networks. Firstly, we present a method to transform a set of time series into a network using different distance functions, where each…

机器学习 · 统计学 2015-08-20 Leonardo N. Ferreira , Liang Zhao

Physical data layout is an important performance factor for modern databases. Clustering, i.e., storing similar values in proximity, can lead to performance gains in several ways. We present an automated model to determine beneficial…

数据库 · 计算机科学 2021-03-30 Alexander Löser

Hierarchical clustering over graphs is a fundamental task in data mining and machine learning with applications in domains such as phylogenetics, social network analysis, and information retrieval. Specifically, we consider the recently…

数据结构与算法 · 计算机科学 2022-06-16 Arpit Agarwal , Sanjeev Khanna , Huan Li , Prathamesh Patil

Anomaly detection to recognize unusual events in large scale systems in a time sensitive manner is critical in many industries, eg. bank fraud, enterprise systems, medical alerts, etc. Large-scale systems often grow in size and complexity…

机器学习 · 计算机科学 2022-10-31 Srishti Mishra , Tvarita Jain , Dinkar Sitaram

Asteroid diameters are traditionally difficult to estimate. When a direct measurement of the diameter cannot be made through either occultation or direct radar observation, the most common method is to approximate the diameter from infrared…

地球与行星天体物理 · 物理学 2023-05-29 Zachary Murray

We consider the classic Correlation Clustering problem: Given a complete graph where edges are labelled either $+$ or $-$, the goal is to find a partition of the vertices that minimizes the number of the \pedges across parts plus the number…

数据结构与算法 · 计算机科学 2023-10-02 Vincent Cohen-Addad , Euiwoong Lee , Shi Li , Alantha Newman

We address the problem of un-supervised soft-clustering called micro-clustering. The aim of the problem is to enumerate all groups composed of records strongly related to each other, while standard clustering methods separate records at…

数据结构与算法 · 计算机科学 2016-06-07 Takeaki Uno , Hiroki Maegawa , Takanobu Nakahara , Yukinobu Hamuro , Ryo Yoshinaka , Makoto Tatsuta