Related papers: Long memory and multifractality: A joint test
While many have shown how Large Language Models (LLMs) can be applied to a diverse set of tasks, the critical issues of data contamination and memorization are often glossed over. In this work, we address this concern for tabular data.…
Several phenomena are available representing market activity: volumes, number of trades, durations between trades or quotes, volatility - however measured - all share the feature to be represented as positive valued time series. When…
Large language models (LLMs) cannot be trusted for economic forecasts during periods covered by their training data. Counterfactual forecasting ability is non-identified when the model has seen the realized values: any observed output is…
We develop the mathematical properties of a multifractal analysis of data based on the weak scaling exponent. The advantage of this analysis is that it does not require any a priori global regularity assumption on the analyzed signal, in…
In LLM evaluations, reasoning is often distinguished from recall/memorization by performing numerical variations to math-oriented questions. Here we introduce a general variation method for multiple-choice questions that completely…
The Mike-Farmer (MF) model was constructed empirically based on the continuous double auction mechanism in an order-driven market, which can successfully reproduce the cubic law of returns and the diffusive behavior of stock prices at the…
The inversion formula for conservative multifractal measures was unveiled mathematically a decade ago, which is however not well tested in real complex systems. In this Letter, we propose to verify the inversion formula using high-frequency…
This article develops nonparametric cointegrating regression models with endogeneity and semi-long memory. We assume that semi-long memory is produced in the regressor process by tempering of random shock coefficients. The fundamental…
We study the properties of memory of a financial time series adopting two different methods of analysis, the detrended fluctuation analysis (DFA) and the analysis of the power spectrum (PSA). The methods are applied on three time series:…
Fractal behavior and long-range dependence have been observed in an astonishing number of physical systems. Either phenomenon has been modeled by self-similar random functions, thereby implying a linear relationship between fractal…
We systematically evaluate the reproducibility of data analysis conducted by Large Language Models (LLMs). We evaluate two prompting strategies, six models, and four temperature settings, with ten independent executions per configuration,…
Researchers have used many different methods to detect the possibility of long-term dependence (long memory) in stock market returns, but evidence is in general mixed. In this paper, three different tests, (namely Rescaled Range (R/S), its…
We study the long memory of order flow for each of three liquid currency pairs on a large electronic trading platform in the foreign exchange (FX) spot market. Due to the extremely high levels of market activity on the platform, and in…
The Maximum Mean Discrepancy (MMD) is a widely used multivariate distance metric for two-sample testing. The standard MMD test statistic has an intractable null distribution typically requiring costly resampling or permutation approaches…
This paper explores seasonal and long-memory time series properties by using the seasonal fractional ARIMA model when the seasonal data has one and two seasonal periods and short-memory counterparts. The stationarity and invertibility…
Current large language models (LLMs) often perform poorly on simple fact retrieval tasks. Here we investigate if coupling a dynamically adaptable external memory to a LLM can alleviate this problem. For this purpose, we test Larimar, a…
Recent multimodal large language models (MLLMs) have demonstrated significant potential in open-ended conversation, generating more accurate and personalized responses. However, their abilities to memorize, recall, and reason in sustained…
In this paper, we use the generalized Hurst exponent approach to study the multi- scaling behavior of different financial time series. We show that this approach is robust and powerful in detecting different types of multiscaling. We…
We develop a novel approach to tackle the common but challenging problem of conformal inference for missing data in machine learning, focusing on Missing at Random (MAR) data. We propose a new procedure Conformal prediction for Missing data…
This article deals with detection of nonconstant long memory parameter in time series. The null hypothesis presumes stationary or nonstationary time series with constant long memory parameter, typically an I(d) series with d>-.5. The…