English
Related papers

Related papers: A Non-Parametric Approach to Detect Patterns in Bi…

200 papers

Much of the on-going statistical analysis of DNA sequences is focused on the estimation of characteristics of coding and non-coding regions that would possibly allow discrimination of these regions. In the current approach, we concentrate…

Genomics · Quantitative Biology 2009-11-10 D. Kugiumtzis , A. Provata

A sequence S is nonrepetitive if no two adjacent blocks of S are the same. In 1906 Thue proved that there exist arbitrarily long nonrepetitive sequences over 3 symbols. We consider the online variant of this result in which a nonrepetitive…

Combinatorics · Mathematics 2012-05-01 Jarosław Grytczuk , Piotr Szafruga , Michał Zmarz

The well-known Gumbel-Max trick for sampling from a categorical distribution can be extended to sample $k$ elements without replacement. We show how to implicitly apply this 'Gumbel-Top-$k$' trick on a factorized distribution over…

Machine Learning · Computer Science 2019-05-31 Wouter Kool , Herke van Hoof , Max Welling

The Win Ratio has gained significant traction in cardiovascular trials as a novel method for analyzing composite endpoints (Pocock and others, 2012). Compared with conventional approaches based on time to the first event, the Win Ratio…

Methodology · Statistics 2024-10-10 Baoshan Zhang , Yuan Wu

We propose a general framework for constructing powerful, sequential hypothesis tests for a large class of nonparametric testing problems. The null hypothesis for these problems is defined in an abstract form using the action of two known…

Machine Learning · Statistics 2023-10-31 Teodora Pandeva , Patrick Forré , Aaditya Ramdas , Shubhanshu Shekhar

Repetitive elements are important in genomic structures, functions and regulations, yet effective methods in precisely identifying repetitive elements in DNA sequences are not fully accessible, and the relationship between repetitive…

Genomics · Quantitative Biology 2016-08-03 Changchuan Yin

We revisit outlier hypothesis testing, propose exponentially consistent low complexity fixed-length and sequential tests and show that our tests achieve better tradeoff between detection performance and computational complexity than…

Information Theory · Computer Science 2026-01-09 Jun Diao , Jingjing Wang , Lin Zhou

Model misspecification can create significant challenges for the implementation of probabilistic models, and this has led to development of a range of robust methods which directly account for this issue. However, whether these more…

Machine Learning · Statistics 2025-04-22 Oscar Key , Arthur Gretton , François-Xavier Briol , Tamara Fernandez

The goal of group testing is to efficiently identify a few specific items, called positives, in a large population of items via tests. A test is an action on a subset of items which returns positive if the subset contains at least one…

Information Theory · Computer Science 2021-11-08 Thach V. Bui , Mahdi Cheraghchi , An T. H. Nguyen , Thuc D. Nguyen

The field of causal discovery develops model selection methods to infer cause-effect relations among a set of random variables. For this purpose, different modelling assumptions have been proposed to render cause-effect relations…

Methodology · Statistics 2023-11-09 Daniela Schkoda , Mathias Drton

We present a Bayesian framework for learning probabilistic specifications from large, unstructured code corpora, and a method to use this framework to statically detect anomalous, hence likely buggy, program behavior. The distinctive…

Software Engineering · Computer Science 2017-03-07 Vijayaraghavan Murali , Swarat Chaudhuri , Chris Jermaine

The scope of this paper is the presentation of a test that enables to detect heteroscedasticity in univariate regression model. The test is simple to compute and very general since no hypothesis is made on the regularity of the response…

Methodology · Statistics 2010-03-23 Jean-Baptiste Aubin , Samuela Leoni-Aubin

In this article we introduce a dual of the uniform boundedness principle which does not require completeness and gives an indirect means for testing the boundedness of a set. The dual principle, although known to the analyst and despite its…

Functional Analysis · Mathematics 2020-11-30 Ehssan Khanmohammadi , Omid Khanmohamadi

How can researchers test for heterogeneity in the local structure of a network? In this paper, we present a framework that utilizes random sampling to give subgraphs which are then used in a goodness of fit test to test for heterogeneity.…

Methodology · Statistics 2015-12-04 Jonathan Tuke , Matthew Roughan

We use the martingale-theoretic approach of game-theoretic probability to incorporate imprecision into the study of randomness. In particular, we define several notions of randomness associated with interval, rather than precise,…

Probability · Mathematics 2021-06-24 Gert de Cooman , Jasper De Bock

We consider sequential hypothesis testing based on observations which are received in groups of random size. The observations are assumed to be independent both within and between the groups. We assume that the group sizes are independent…

Methodology · Statistics 2021-10-11 Andrey Novikov , Xóchitl Itxel Popoca-Jiménez

To render a sequence testable, namely capable of identifying and detecting errors, it is necessary to apply a transformation that increases its length by introducing statistical dependence among symbols, as commonly exemplified by the…

Information Theory · Computer Science 2025-07-08 Aida Koch , Alix Petit

What advantage do \emph{sequential} procedures provide over batch algorithms for testing properties of unknown distributions? Focusing on the problem of testing whether two distributions $\mathcal{D}_1$ and $\mathcal{D}_2$ on $\{1,\dots,…

Data Structures and Algorithms · Computer Science 2022-05-13 Omar Fawzi , Nicolas Flammarion , Aurélien Garivier , Aadil Oufkir

Multiple metrics have been developed to detect causality relations between data describing the elements constituting complex systems, all of them considering their evolution through time. Here we propose a metric able to detect causality…

Data Analysis, Statistics and Probability · Physics 2016-05-20 Massimiliano Zanin

Universal outlier hypothesis testing is studied in a sequential setting. Multiple observation sequences are collected, a small subset of which are outliers. A sequence is considered an outlier if the observations in that sequence are…

Statistics Theory · Mathematics 2014-11-27 Yun Li , Sirin Nitinawarat , Venugopal V. Veeravalli