中文
相关论文

相关论文: Characterizing how 'distributional' NLP corpora di…

200 篇论文

Computing the similarity between two probability distributions is a recurring theme across control. We introduce a unified family of distances between the probability distributions of two random variables that is based on the discrepancy…

系统与控制 · 电气工程与系统科学 2025-10-03 Alexandros E. Tzikas , Arec Jamgochian , Nazim Kemal Ure , Mykel J. Kochenderfer , Stephen P. Boyd

In this paper, we develop a finite mixture of convolutional distributions, a statistical model to analyze continuous data distributed approximately on a mixture of low-dimensional affine subspaces. The observations are assumed independent…

统计理论 · 数学 2026-04-21 Sunrit Chakraborty , XuanLong Nguyen

This work considers the asymptotic behavior of the distance between two sample covariance matrices (SCM). A general result is provided for a class of functionals that can be expressed as sums of traces of functions that are separately…

统计理论 · 数学 2023-12-25 Roberto Pereira , Xavier Mestre , David Gregoratti

The degree distribution is an important characteristic of complex networks. In many applications, quantification of degree distribution in the form of a fixed-length feature vector is a necessary step. On the other hand, we often need to…

社会与信息网络 · 计算机科学 2013-12-24 Sadegh Aliakbary , Jafar Habibi , Ali Movaghar

We develop techniques to quantify the degree to which a given (training or testing) example is an outlier in the underlying distribution. We evaluate five methods to score examples in a dataset by how well-represented the examples are, for…

机器学习 · 计算机科学 2019-10-30 Nicholas Carlini , Úlfar Erlingsson , Nicolas Papernot

We consider the problem of allocating samples to a finite set of discrete distributions in order to learn them uniformly well in terms of four common distance measures: $\ell_2^2$, $\ell_1$, $f$-divergence, and separation distance. To…

机器学习 · 统计学 2019-12-10 Shubhanshu Shekhar , Tara Javidi , Mohammad Ghavamzadeh

Elliptically symmetric distributions are a classic example of a semiparametric model where the location vector and the scatter matrix (or a parameterization of them) are the two finite-dimensional parameters of interest, while the density…

统计理论 · 数学 2026-03-18 Stefano Fortunati , Jean-Pierre Delmas , Esa Ollila

The concentration of a distribution toward a lower bound is a conceptually simple property that closely relates to concepts of rarity and poverty, but that lacks a global descriptive statistic. We term this property 'shift' and define it as…

统计方法学 · 统计学 2025-06-16 Kenneth J. Locey , Brian D. Stein

Machine learning (ML) has employed various discretization methods to partition numerical attributes into intervals. However, an effective discretization technique remains elusive in many ML applications, such as association rule mining.…

机器学习 · 计算机科学 2023-11-07 Minakshi Kaushik , Rahul Sharma , Dirk Draheim

Distributed algorithms, particularly Diffusion Least Mean Square, are widely favored for their reliability, robustness, and fast convergence in various industries. However, limited observability of the target can compromise the integrity of…

信号处理 · 电气工程与系统科学 2023-10-18 Mahdi Shamsi , Farokh Marvasti

Domain specific (dis-)similarity or proximity measures used e.g. in alignment algorithms of sequence data, are popular to analyze complex data objects and to cover domain specific data properties. Without an underlying vector space these…

数据结构与算法 · 计算机科学 2014-11-07 Andrej Gisbrecht , Frank-Michael Schleif

Which neural networks are similar is a fundamental question for both machine learning and neuroscience. Here, it is proposed to base comparisons on the predictive distributions of linear readouts from intermediate representations. In…

机器学习 · 计算机科学 2025-05-27 Heiko H. Schütt

We study the question of how reliably one can distinguish two quantum field theories (QFTs). Each QFT defines a probability distribution on the space of fields. The relative entropy provides a notion of proximity between these distributions…

高能物理 - 理论 · 物理学 2015-05-08 Vijay Balasubramanian , Jonathan J. Heckman , Alexander Maloney

To what extent can we distinguish one probability distribution from another? Are there quantitative measures of distinguishability? The goal of this tutorial is to approach such questions by introducing the notion of the "distance" between…

数据分析、统计与概率 · 物理学 2015-06-23 Ariel Caticha

When we represent a network of sensors in Euclidean space by a graph, there are two distances between any two nodes that we may consider. One of them is the Euclidean distance. The other is the distance between the two nodes in the graph,…

网络与互联网体系结构 · 计算机科学 2009-06-10 Rodrigo S. C. Leao , Valmir C. Barbosa

There are many open questions pertaining to the statistical analysis of random objects, which are increasingly encountered. A major challenge is the absence of linear operations in such spaces. A basic statistical task is to quantify…

统计方法学 · 统计学 2025-08-01 Wookyeong Song , Hans-Georg Müller

Distance metric learning is a successful way to enhance the performance of the nearest neighbor classifier. In most cases, however, the distribution of data does not obey a regular form and may change in different parts of the feature…

计算机视觉与模式识别 · 计算机科学 2018-03-19 Hossein Rajabzadeh , Mansoor Zolghadri Jahromi , Mohammad Sadegh Zare , Mostafa Fakhrahmad

Network data arises through observation of relational information between a collection of entities. Recent work in the literature has independently considered when (i) one observes a sample of networks, connectome data in neuroscience being…

统计方法学 · 统计学 2022-06-22 George Bolt , Simón Lunagómez , Christopher Nemeth

The ability to measure similarity between documents enables intelligent summarization and analysis of large corpora. Past distances between documents suffer from either an inability to incorporate semantic similarities between words or from…

机器学习 · 计算机科学 2019-11-05 Mikhail Yurochkin , Sebastian Claici , Edward Chien , Farzaneh Mirzazadeh , Justin Solomon

In this paper we introduce the idea of partially sorting data to design nonparametric tests. This approach gives rise to tests that are sensitive to both the order and the underlying distribution of the data. We focus in particular on a…

统计理论 · 数学 2022-10-27 Krzysztof Bisewski , H. M. Jansen , Yoni Nazarathy