Related papers: Track Fit Hypothesis Testing and Kink Selection us…
Modern genomics experiments measure functional behaviors for many thousands of DNA sequences. We suggest that, especially when these sequences are chosen at random, it is natural to compute correlation functions between sequences and…
We describe an approach to improving model fitting and model generalization that considers the entropy of distributions of modelling residuals. We use simple simulations to demonstrate the observational signatures of overfitting on ordered…
A simple test is proposed for examining the correctness of a given completely specified response function against unspecified general alternatives in the context of univariate regression. The usual diagnostic tools based on residuals plots…
Analysis of data from particle physics experiments traditionally sacrifices some sensitivity to new particles for the sake of practical computability, effectively ignoring some potentially striking signatures. However, recent advances in…
Higher order correlation measurements involve multiple event averages which must run over unequal events to avoid statistical bias. We derive correction formulas for small event samples, where the bias is largest, and utilize the results to…
Most particle detectors are based on the hypothesis that particles are emitted randomly upon nuclear decay. In the present work, we tested the hypothesis of the existence of correlation in the random trajectories of alpha particles emitted…
We describe two families of statistical tests to detect partial correlation in vectorial timeseries. The tests measure whether an observed timeseries Y can be predicted from a second series X, even after accounting for a third series Z…
Concerning bivariate least squares linear regression, the classical approach pursued for functional models in earlier attempts is reviewed using a new formalism in terms of deviation (matrix) traces. Within the framework of classical error…
The ability that one system immediately affects another one by using local measurements is regarded as quantum steering, which can be detected by various steering criteria. Recently, Mondal et al. [Phys. Rev. A 98, 052330 (2018)] derived…
Linear theory provides a reasonable description of the velocity correlations of biased tracers both perpendicular and parallel to the line of separation, provided one accounts for the fact that the measurement is almost always made using…
We discuss fitting correlated data - with the example of hadron mass spectroscopy in mind. The main conclusion is that the method of minimising correlated $\chi^2$ is unreliable if the data sample is too small.
A novel combination of established data analysis techniques for reconstructing all charged-particle tracks in high energy collisions is proposed. It uses all information available in a collision event while keeping competing choices open as…
We argue that precise measurements of charge and magnetic radii can meaningfully constrain diquark models of the nucleon. We construct properly symmetrized, nonrelativistic three-quark wave functions that interpolate between the limits of a…
The simplicity of a question such as wondering if correlations characterize or not a certain system collides with the experimental difficulty of accessing such information. Here we present a low demanding experimental approach which refers…
Given a set of sequences, the distance between pairs of them helps us to find their similarity and derive structural relationship amongst them. For genomic sequences such measures make it possible to construct the evolution tree of…
We derive a simple consistency relation from the running of the tensor-to-scalar ratio. This new relation is first order in the slow-roll approximation. While for single field models we can obtain what can be found by using other…
Distribution shifts are common in real-world datasets and can affect the performance and reliability of deep learning models. In this paper, we study two types of distribution shifts: diversity shifts, which occur when test samples exhibit…
Choosing which properties of the data to use as input to multivariate decision algorithms -- a.k.a. feature selection -- is an important step in solving any problem with machine learning. While there is a clear trend towards training…
Chi-squared tests for lack of fit are traditionally employed to find evidence against a hypothesized model, with the model accepted if the Karl Pearson statistic comparing observed and expected numbers of observations falling within cells…
The detection of similarities between long DNA and protein sequences is studied using concepts of statistical physics. It is shown that mutual similarities can be detected by sequence alignment methods only if their amount exceeds a…