English
Related papers

Related papers: Breaking the Trend: How to Avoid Cherry-Picked Sig…

200 papers

A common approach to detect multiple changepoints is to minimise a measure of data fit plus a penalty that is linear in the number of changepoints. This paper shows that the general finite sample behaviour of such a method can be related to…

Statistics Theory · Mathematics 2022-08-15 Chao Zheng , Idris A. Eckley , Paul Fearnhead

Recent results in compressed sensing showed that the optimal subsampling strategy should take into account the sparsity pattern of the signal at hand. This oracle-like knowledge, even though desirable, nevertheless remains elusive in most…

Information Theory · Computer Science 2023-06-28 Simon Ruetz

We consider the problem of sparse canonical correlation analysis (CCA), i.e., the search for two linear combinations, one for each multivariate, that yield maximum correlation using a specified number of variables. We propose an efficient…

Computation · Statistics 2008-01-18 Ami Wiesel , Mark Kliger , Alfred O. Hero

We study the profitability of optimal mean reversion trading strategies in the US equity market. Different from regular pair trading practice, we apply maximum likelihood method to construct the optimal static pairs trading portfolio that…

Portfolio Management · Quantitative Finance 2016-02-19 Peng Huang , Tianxiang Wang

The paper tackles the problem of deriving a topological structure among stock prices from high frequency historical values. Similar studies using low frequency data have already provided valuable insights. However, in those cases data need…

Statistical Finance · Quantitative Finance 2008-12-02 Donatello Materassi , Giacomo Innocenti

In this paper, matching pairs of random graphs under the community structure model is considered. The problem emerges naturally in various applications such as privacy, image processing and DNA sequencing. A pair of randomly generated…

Cryptography and Security · Computer Science 2018-11-01 F. Shirani , S. Garg , E. Erkip

For the binary regression, the use of symmetrical link functions are not appropriate when we have evidence that the probability of success increases at a different rate than decreases. In these cases, the use of link functions based on the…

Methodology · Statistics 2024-07-23 João Victor B. de Freitas , Caio L. N. Azevedo

Many machine learning algorithms are based on the assumption that training examples are drawn independently. However, this assumption does not hold anymore when learning from a networked sample because two or more training examples may…

Artificial Intelligence · Computer Science 2017-06-06 Yuyi Wang , Jan Ramon , Zheng-Chu Guo

We develop new statistics for robustly filtering corrupted keypoint matches in the structure from motion pipeline. The statistics are based on consistency constraints that arise within the clustered structure of the graph of keypoint…

Computer Vision and Pattern Recognition · Computer Science 2022-01-19 Yunpeng Shi , Shaohan Li , Tyler Maunu , Gilad Lerman

Symbolic regression is a machine learning technique, and it has seen many advancements in recent years, especially in genetic programming approaches (GPSR). Furthermore, it has been known for many years that constant optimization of…

Machine Learning · Computer Science 2024-12-04 L. G. A dos Reis , V. L. P. S. Caminha , T. J. P. Penna

This paper is about how we study statistical methods. As an example, it uses the random regressions model, in which the intercept and slope of cluster-specific regression lines are modeled as a bivariate random effect. Maximizing this…

Other Statistics · Statistics 2019-05-22 James S. Hodges

An important task in machine learning and statistics is the approximation of a probability measure by an empirical measure supported on a discrete point set. Stein Points are a class of algorithms for this task, which proceed by…

All 21-cm signal experiments rely on electronic receivers that affect the data via both multiplicative and additive biases through the receiver's gain and noise temperature. While experiments attempt to remove these biases, the residuals of…

Cosmology and Nongalactic Astrophysics · Physics 2021-07-14 Keith Tauscher , David Rapetti , Bang D. Nhan , Alec Handy , Neil Bassett , Joshua Hibbard , David Bordenave , Richard F. Bradley , Jack O. Burns

Sequentially obtained dataset usually exhibits different behavior at different data resolutions/scales. Instead of inferring from data at each scale individually, it is often more informative to interpret the data as an ensemble of time…

Mesoscale and Nanoscale Physics · Physics 2021-03-19 Yuan Yang , Jie Ding

In this paper we revisit the problem of decomposing a signal into a tendency and a residual. The tendency describes an executive summary of a signal that encapsulates its notable characteristics while disregarding seemingly random, less…

Signal Processing · Electrical Eng. & Systems 2024-01-10 Caio Alves , Juan M. Restrepo , Jorge M. Ramirez

Consensus Sequences of event logs are often used in process mining to quickly grasp the core sequence of events to be performed in a process, or to represent the backbone of the process for doing other analyses. However, it is still not…

Artificial Intelligence · Computer Science 2021-07-05 Zhichao Xu , Shuhong Chen

The success of a cross-sectional systematic strategy depends critically on accurately ranking assets prior to portfolio construction. Contemporary techniques perform this ranking step either with simple heuristics or by sorting outputs from…

Trading and Market Microstructure · Quantitative Finance 2020-12-15 Daniel Poh , Bryan Lim , Stefan Zohren , Stephen Roberts

For many real data, long term observation consists of different processes that coexist or occur one after the other. Those processes very often exhibit different statistical properties and thus before the further analysis the observed data…

Statistics Theory · Mathematics 2016-05-30 Kucharczyk Daniel. Wyłomańska Agnieszka , Zimroz Radosław

We develop new methods to integrate experimental and observational data in causal inference. While randomized controlled trials offer strong internal validity, they are often costly and therefore limited in sample size. Observational data,…

Econometrics · Economics 2025-11-04 Xuelin Yang , Licong Lin , Susan Athey , Michael I. Jordan , Guido W. Imbens

In this paper, selection of an active sensor subset for tracking a discrete time, finite state Markov chain having an unknown transition probability matrix (TPM) is considered. A total of N sensors are available for making observations of…

Machine Learning · Computer Science 2020-11-02 Mrigank Raman , Ojal Kumar , Arpan Chattopadhyay