Related papers: Supervised Contamination Detection, with Flow Cyto…
We study a class of hypothesis testing problems in which, upon observing the realization of an $n$-dimensional Gaussian vector, one has to decide whether the vector was drawn from a standard normal distribution or, alternatively, whether…
We study a statistical procedure based on higher criticism (HC) to address the sparse multi-stream quickest change-point detection problem. Namely, we aim to detect a potential change in the distribution of multiple data streams at some…
Cells, other than their biological properties, have different electric and physical properties. In an impedance cytometer, cells should pass one by one in the detection region where pairs of electrodes are located. When cells are located…
Distributed Systems involve two or more computer systems which may be situated at geographically distinct locations and are connected by a communication network. Due to failures in the communication link, faults arise which may make the…
Specimens are collected from $N$ different sources. Each specimen has probability $p$ of being contaminated, independently of the other specimens. We assume group testing is applicable, namely one can take small portions from several…
This paper explores a Gaussian process emulator based approach for rapid Bayesian inference of contaminant source location and characteristics in an indoor environment. In the pre-event detection stage, the proposed approach represents…
The problem of joint sequential detection and isolation is considered in the context of multiple, not necessarily independent, data streams. A multiple testing framework is proposed, where each hypothesis corresponds to a different subset…
Time series anomaly detection is instrumental in maintaining system availability in various domains. Current work in this research line mainly focuses on learning data normality deeply and comprehensively by devising advanced neural network…
Label free tracking of small bio-particles such as proteins or viruses is of great utility in the study of biological processes, however such experiments are frequently hindered by weak signal strengths and a susceptibility to scattering…
Image classification models deployed in the real world may receive inputs outside the intended data distribution. For critical applications such as clinical decision making, it is important that a model can detect such out-of-distribution…
Flux sampling is an analysis that, based on a distribution, picks randomly an efficient number of points from the solution space of a metabolic model. Unlike most constraint-based analyses, flux sampling does not require an objective…
Discrete element method simulations of confined bidisperse granular shear flows elucidate the balance between diffusion and segregation that can lead to either mixed or segregated states, depending on confining pressure. Results indicate…
Pinched flow fractionation is shown to be an efficient and selective way to quickly separate particles by size in a very polydisperse semi-concentrated suspension. In an effort to optimize the method, we discuss the quantitative influence…
There has been significant study on the sample complexity of testing properties of distributions over large domains. For many properties, it is known that the sample complexity can be substantially smaller than the domain size. For example,…
Large language models (LLMs) are known to be trained on vast amounts of data, which may unintentionally or intentionally include data from commonly used benchmarks. This inclusion can lead to cheatingly high scores on model leaderboards,…
Given an aftermath of a cascade in the network, i.e. a set $V_I$ of "infected" nodes after an epidemic outbreak or a propagation of rumors/worms/viruses, how can we infer the sources of the cascade? Answering this challenging question is…
The risk of false positive results in noninvasive prenatal diagnosis focused on fetal gender and RhD status determination could be a problem in clinical routine. This is because these tests are based on detection of presence of DNA…
Detection of change-points in a sequence of high-dimensional observations is a very challenging problem, and this becomes even more challenging when the sample size (i.e., the sequence length) is small. In this article, we propose some…
We introduce Time-Conditioned Contraction Matching (TCCM), a novel method for semi-supervised anomaly detection in tabular data. TCCM is inspired by flow matching, a recent generative modeling framework that learns velocity fields between…
Sampling from high-dimensional distributions is a fundamental problem in statistical research and practice. However, great challenges emerge when the target density function is unnormalized and contains isolated modes. We tackle this…