Related papers: Reliability Correction is Key for Robust Kepler Oc…
We describe the Reversibility Error Method (REM) and its applications to planetary dynamics. REM is based on the time-reversibility analysis of the phase-space trajectories of conservative Hamiltonian systems. The round-off errors break the…
Conformal prediction is a flexible framework for calibrating machine learning predictions, providing distribution-free statistical guarantees. In outlier detection, this calibration relies on a reference set of labeled inlier data to…
Assessment of replicability is critical to ensure the quality and rigor of scientific research. In this paper, we discuss inference and modeling principles for replicability assessment. Targeting distinct application scenarios, we propose…
In the first three years of operation the Kepler mission found 3,697 planet candidates from a set of 18,406 transit-like features detected on over 200,000 distinct stars. Vetting candidate signals manually by inspecting light curves and…
Galaxy peculiar velocities are excellent cosmological probes provided that biases inherent to their measurements are contained before any study. This paper proposes a new algorithm based on an object point process model whose probability…
A subvector of predictor that satisfies the ignorability assumption, whose index set is called a sufficient adjustment set, is crucial for conducting reliable causal inference based on observational data. In this paper, we propose a general…
We apply multiple testing procedures to the validation of estimated default probabilities in credit rating systems. The goal is to identify rating classes for which the probability of default is estimated inaccurately, while still…
To fully leverage the statistical strength of the large number of planets found by projects such as the Kepler survey, the properties of planets and their host stars must be measured as accurately as possible. One key population for planet…
We study sequential testing for a binary disease outcome when risk follows an unknown logistic model. At each round, the decision maker may either pay for a test revealing the true label or predict the outcome based on patient features and…
Machine learning classification tasks often benefit from predicting a set of possible labels with confidence scores to capture uncertainty. However, existing methods struggle with the high-dimensional nature of the data and the lack of…
To check the accuracy of Bayesian computations, it is common to use rank-based simulation-based calibration (SBC). However, SBC has drawbacks: The test statistic is somewhat ad-hoc, interactions are difficult to examine, multiple testing is…
The Ripper algorithm is designed to generate rule sets for large datasets with many features. However, it was shown that the algorithm struggles with classification performance in the presence of missing data. The algorithm struggles to…
This paper aims to derive a map of relative planet occurrence rates that can provide constraints on the overall distribution of terrestrial planets around FGK stars. Based on the planet candidates in the Kepler DR25 data release, I first…
Reparameterization from the standard set of orbital elements to Cartesian position-velocity vectors can be computationally advantageous for orbit inference problems, particularly when orbital elements are weakly constrained. Here we present…
The Kepler Mission was designed to identify and characterize transiting planets in the Kepler Field of View and to determine their occurrence rates. Emphasis was placed on identification of Earth-size planets orbiting in the Habitable Zone…
A composite likelihood is a non-genuine likelihood function that allows to make inference on limited aspects of a model, such as marginal or conditional distributions. Composite likelihoods are not proper likelihoods and need therefore…
Many two-sample problems call for a comparison of two distributions from an exponential family. Density ratio estimation methods provide ways to solve such problems through direct estimation of the differences in natural parameters. The…
Most computer vision application rely on algorithms finding local correspondences between different images. These algorithms detect and compare stable local invariant descriptors centered at scale-invariant keypoints. Because of the…
We measure planet occurrence rates using the planet candidates discovered by the Q1-Q16 Kepler pipeline search. This study examines planet occurrence rates for the Kepler GK dwarf target sample for planet radii, 0.75<Rp<2.5 Rearth, and…
Confounding is a significant obstacle to unbiased estimation of causal effects from observational data. For settings with high-dimensional covariates -- such as text data, genomics, or the behavioral social sciences -- researchers have…