Ashoka Prabashwara, Patricia Menéndez, Liam Hodgkinson, Stuart Lee
Modern time series are often long, serially dependent, and non-stationary. Existing change-point methods either target specific changes or become computationally intensive when using nonparametric costs on long series. Many also require thresholds to be carefully calibrated under serial dependence. We introduce SCAN, an offline method for detecting multiple distributional change-points in long, serially dependent univariate time series. SCAN compares adjacent windows using an integral probability metric, calibrates local discrepancies with a dependence-aware bootstrap, and refines candidate locations using a scaled 1-Wasserstein criterion, enabling detection of changes in mean, variance, and broader distributional structure within a unified framework. An ensemble over multiple window sizes reduces sensitivity to window size and threshold specification. We establish consistency of the estimated number and locations of change-points under exponential alpha-mixing dependence, and show that the localization statistic reduces to a CUSUM-type statistic under pure mean shifts. In simulations with up to one million observations, SCAN generally achieves higher covering and F1-scores than competing methods across mean and joint mean-variance shifts, particularly under serial dependence. On real data, SCAN identifies labeled activity transitions in sensor data and interpretable structural changes in hourly Bitcoin prices. Implementations are available in the Python package scan-cpd and R package scanr.
In this paper, we propose a distribution-free test for detecting changepoint in the mean direction of angular data. The uncertainty in angular measurements is quantified through the \textit{square of an angle}, derived from the intrinsic geometry of the torus. It is established that, under the null hypothesis, the test statistic distributionally converges to the Kolmogorov distribution, while under the alternative hypothesis, both the consistency of the test and the asymptotic properties of the changepoint estimator are established. Through extensive simulations, we compare the empirical performance of the proposed method with two existing approaches for angular data and further benchmark it against a test based on the circular arc length distance. Finally, we demonstrate the practical utility of our approach by analyzing the timestamps of extreme events in Bitcoin, Ethereum, and Gold price datasets, where the continuous, high-frequency nature of the data is modeled in the circular framework.
Simultaneous occurrences of extreme events need not imply symmetric or reciprocal tail dependence. However, most existing measures of extremal dependence are inherently symmetric and hence often fail to capture directional influence in tail association. We introduce a rank-based measure of Extreme Tail Association (ETA) for bivariate data quantifying such directional influence of one variable on another in extreme tail regions. The proposed estimator is easily computable, consistent with its population counterpart, and asymptotically normal under mild conditions, allowing for statistical inference. We further develop a formal test for asymmetry in tail association based on a multiplier bootstrap procedure. The practical relevance of the methodology is illustrated using data on extreme price movements in major cryptocurrencies. Beyond providing a flexible tool for extremal association, the proposed framework offers a substantive argument for investigating causal relationships in extreme scenarios.
Time series forecasting enables early warning and has driven asset performance management from traditional planned maintenance to predictive maintenance. However, the lack of interpretability in forecasting methods undermines users' trust and complicates debugging for developers. Consequently, interpretable time-series forecasting has attracted increasing research attention. Nevertheless, existing methods suffer from several limitations, including insufficient modeling of temporal dependencies, lack of feature-level interpretability to support early warning, and difficulty in simultaneously achieving the accuracy and interpretability. This paper proposes the interpretable polynomial learning (IPL) method, which integrates interpretability into the model structure by explicitly modeling original features and their interactions of arbitrary order through polynomial representations. This design preserves temporal dependencies, provides feature-level interpretability, and offers a flexible trade-off between prediction accuracy and interpretability by adjusting the polynomial degree. We evaluate IPL on simulated and Bitcoin price data, showing that it achieves high prediction accuracy with superior interpretability compared with widely used explainability methods. Experiments on field-collected antenna data further demonstrate that IPL yields simpler and more efficient early warning mechanisms.
Structural changes and outliers often coexist, complicating statistical inference. This paper addresses the problem of testing for parameter changes in conditionally heteroscedastic time series models, particularly in the presence of outliers. To mitigate the impact of outliers, we introduce a two-step procedure comprising robust estimation and residual truncation. Based on this procedure, we propose a residual-based robust CUSUM test and its self-normalized counterpart. We derive the limiting null distributions of the proposed robust tests and establish their consistency. Simulation results demonstrate the strong robustness of the tests against outliers. To illustrate the practical application, we analyze Bitcoin data.
Do Ethereum's Layer-2 (L2) rollups actually decongest the Layer-1 (L1) mainnet once protocol upgrades and demand are held constant? Using a 1245-day daily panel from August 5, 2021 to December 31, 2024 that spans the London, Merge, and Dencun upgrades, we link Ethereum fee and congestion metrics to L2 user activity, macro-demand proxies, and targeted event indicators. We estimate a regime-aware error-correction model that treats posting-clean L2 user share as a continuous treatment. Over the pre-Dencun (London+Merge) window, a 10 percentage point increase in L2 adoption lowers median base fees by about 13% -- roughly 5 Gwei at pre-Dencun levels -- and deviations from the long-run relation decay with an 11-day half-life. Block utilization and a scarcity index show similar congestion relief. After Dencun, L2 adoption is already high and treatment support narrows, so blob-era estimates are statistically imprecise and we treat them as exploratory. The pre-Dencun window therefore delivers the first cross-regime causal estimate of how aggregate L2 adoption decongests Ethereum, together with a reusable template for monitoring rollup-centric scaling strategies.
Betas from spot regressions are central to asset pricing and risk management, as measures of systematic risk. This paper develops a new estimation and inference framework for spot regressions by leveraging high-frequency candlesticks, extending conventional (open-to-close) returns with intra-period high/low prices. Specifically, I construct candlestick-based estimators of regression parameters, including spot beta, by minimizing a quadratic risk under a fixed-k asymptotic framework. I then develop a feasible hypothesis testing procedure for spot betas with correct asymptotic size. Simulation results show that the proposed estimator reduces estimation risk relative to return-based estimators, especially in small samples, and the test achieves notably higher power. I apply the framework to assess the market neutrality of Bitcoin using 1-minute data on IBIT and SPY, finding deviations from neutrality, particularly in high-volatility periods.
Financial fraud has been growing exponentially in recent years. The rise of cryptocurrencies as an investment asset has simultaneously seen a parallel growth in cryptocurrency scams. To detect possible cryptocurrency fraud, and in particular market manipulation, previous research focused on the detection of changes in the network of trades; however, market manipulators are now trading across multiple cryptocurrency platforms, making their detection more difficult. Hence, it is important to consider the identification of changes across several trading networks or a `network of networks' over time. To this end, in this article, we propose a new change-point detection method in the network structure of tensor-variate data. This new method, labeled TenSeg, first employs a tensor decomposition, and second detects multiple change-points in the second-order (cross-covariance or network) structure of the decomposed data. It allows for change-point detection in the presence of frequent changes of possibly small magnitudes and is computationally fast. We apply our method to several simulated datasets and to a cryptocurrency dataset, which consists of network tensor-variate data from the Ethereum blockchain. We demonstrate that our approach substantially outperforms other state-of-the-art change-point techniques, and the detected change-points in the Ethereum data set coincide with changes across several trading networks or a `network of networks' over time. Finally, all the relevant \textsf{R} code implementing the method in the article are available on https://github.com/Anastasiou-Andreas/TenSeg.
Federico P. Cortese, Antonio Pievatolo, Elisa Maria Alessi
Statistical jump models have been recently introduced to detect persistent regimes by clustering temporal features and discouraging frequent regime changes. However, they are limited to hard clustering and thereby do not account for uncertainty in state assignments. This work presents an extension of the statistical jump model that incorporates uncertainty estimation in cluster membership. Leveraging the similarities between statistical jump models and the fuzzy c-means framework, our fuzzy jump model sequentially estimates time-varying state probabilities. Our approach offers high flexibility, as it supports both soft and hard clustering through the tuning of a fuzziness parameter, and it naturally accommodates multivariate time series data of mixed types. Through a simulation study, we evaluate the ability of the proposed model to accurately estimate the true latent-state distribution, demonstrating that it outperforms competing approaches under high cluster assignment uncertainty. We further demonstrate its utility on two empirical applications: first, by automatically identifying co-orbital regimes in the three-body problem, a novel application with important implications for understanding asteroid behavior and designing interplanetary mission trajectories; and second, on a financial dataset of five assets representing distinct market sectors (equities, bonds, foreign exchange, cryptocurrencies, and utilities), where the model accurately tracks both bull and bear market phases.
Given a universe of N assets, investors often form equally weighted portfolios (EWPs) by selecting subsets of assets. EWPs are simple, robust, and competitive out-of-sample, yet the uncertainty about which subset truly performs best is largely ignored. Traditional approaches typically rely on a single selected portfolio, but this fails to consider alternative investment strategies that may perform just as well when accounting for statistical uncertainty. To address this selection uncertainty, we introduce the Selection Confidence Set (SCS) for EWPs: the set of all portfolios that, under a given loss function and at a specified confidence level, contains the unknown set of optimal portfolios under repeated sampling. The SCS quantifies selection uncertainty by identifying a range of plausible portfolios, challenging the idea of a uniquely optimal choice. Like a confidence set, its size reflects uncertainty -- growing with noisy or limited data, and shrinking as the sample size increases. Theoretically, we establish that the SCS covers the unknown optimal selection with high probability and characterize how its size grows with underlying uncertainty, corroborating these results through Monte Carlo experiments. Applications to the French 17-Industry Portfolios and Layer-1 cryptocurrencies underscore the importance of accounting for selection uncertainty when comparing equally weighted strategies.
Abstract This study provides essential insights into how diffusion processes unfold in complex networks, with a focus on cryptocurrency blockchains and infrastructure networks. The structural properties of these networks, such as hub-dominated, heavy-tailed topology, network motifs, and node centrality, significantly influence diffusion speed and reach. Using epidemic diffusion models, specifically the Kertesz threshold model and the Susceptible-Infected (SI) model, we analyze key factors affecting diffusion dynamics. To assess the uncertainty in the fraction of infected nodes over time, we employ bootstrap confidence intervals, while Bayesian credible intervals are constructed to quantify parameter uncertainties in the SI models. Our findings reveal substantial variations across different network types, including Erdős-Rényi networks, Geometric Random Graphs, and Delaunay Triangulation networks, emphasizing the role of network architecture in failure propagation. We identify that network motifs are crucial in diffusion. We highlight that hub-dominated networks, which dominate blockchain ecosystems, provide resilience against random failures but remain vulnerable to targeted attacks, posing significant risks to network stability. Furthermore, centrality measures such as degree, betweenness, and clustering coefficient strongly influence the transmissibility of diffusion in both blockchain and critical infrastructure networks.
Estimation of mean shift in a temporally ordered sequence of random variables with a possible existence of change-point is an important problem in many disciplines. In the available literature of more than fifty years the estimation methods of the mean shift is usually dealt as a two-step problem. A test for the existence of a change-point is followed by an estimation process of the mean shift, which is known as testimator. The problem suffers from over parametrization. When viewed as an estimation problem, we establish that the maximum likelihood estimator (MLE) always gives a false alarm indicting an existence of a change-point in the given sequence even though there is no change-point at all. After modelling the parameter space as a modified horn torus. We introduce a new method of estimation of the parameters. The newly introduced estimation method of the mean shift is assessed with a proper Riemannian metric on that conic manifold. It is seen that its performance is superior compared to that of the MLE. The proposed method is implemented on Bitcoin data and compared its performance with the performance of the MLE.
This work focuses on a self-exciting point process defined by a Hawkes-like intensity and a switching mechanism based on a hidden Markov chain. Previous works in such a setting assume constant intensities between consecutive events. We extend the model to general Hawkes excitation kernels that are piecewise constant between events. We develop an expectation-maximization algorithm for the statistical inference of the Hawkes intensities parameters as well as the state transition probabilities. The numerical convergence of the estimators is extensively tested on simulated data. Using high-frequency cryptocurrency data on a top centralized exchange, we apply the model to the detection of anomalous bursts of trades. We benchmark the goodness-of-fit of the model with the Markov-modulated Poisson process and demonstrate the relevance of the model in detecting suspicious activities.
Ayush Jha, Abootaleb Shirvani, Ali Jaffri, Svetlozar T. Rachev · 5 authors
This study presents the Adaptive Minimum-Variance Portfolio (AMVP) framework and the Adaptive Minimum-Risk Rate (AMRR) metric, innovative tools designed to optimize portfolios dynamically in volatile and nonstationary financial markets. Unlike traditional minimum-variance approaches, the AMVP framework incorporates real-time adaptability through advanced econometric models, including ARFIMA-FIGARCH processes and non-Gaussian innovations. Empirical applications on cryptocurrency and equity markets demonstrate the proposed framework's superior performance in risk reduction and portfolio stability, particularly during periods of structural market breaks and heightened volatility. The findings highlight the practical implications of using the AMVP and AMRR methodologies to address modern investment challenges, offering actionable insights for portfolio managers navigating uncertain and rapidly changing market conditions.
In this paper we develop a novel hidden Markov graphical model to investigate time-varying interconnectedness between different financial markets. To identify conditional correlation structures under varying market conditions and accommodate stylized facts embedded in financial time series, we rely upon the generalized hyperbolic family of distributions with time-dependent parameters evolving according to a latent Markov chain. We exploit its location-scale mixture representation to build a penalized EM algorithm for estimating the state-specific sparse precision matrices by means of an $L_1$ penalty. The proposed approach leads to regime-specific conditional correlation graphs that allow us to identify different degrees of network connectivity of returns over time. The methodology's effectiveness is validated through simulation exercises under different scenarios. In the empirical analysis we apply our model to daily returns of a large set of market indexes, cryptocurrencies and commodity futures over the period 2017-2023.
This paper introduces a novel regression model designed for angular response variables with linear predictors, utilizing a generalized Möbius transformation to define the regression curve. By mapping the real axis to the circle, the model effectively captures the relationship between linear and angular components. A key innovation is the introduction of an area-based loss function, inspired by the geometry of a curved torus, for efficient parameter estimation. The semi-parametric nature of the model eliminates the need for specific distributional assumptions about the angular error, enhancing its versatility. Extensive simulation studies, incorporating von Mises and wrapped Cauchy distributions, highlight the robustness of the framework. The model's practical utility is demonstrated through real-world data analysis of Bitcoin and Ethereum, showcasing its ability to derive meaningful insights from complex data structures.
Jimmy Cheung, Smruthi Rangarajan, Amelia Maddocks, Rohitash Chandra
Uncertainty quantification is crucial in time series prediction, and quantile regression offers a valuable mechanism for uncertainty quantification which is useful for extreme value forecasting. Although deep learning models have been prominent in multi-step ahead prediction, the development and evaluation of quantile deep learning models have been limited. We present a novel quantile regression deep learning framework for multi-step time series prediction. In this way, we elevate the capabilities of deep learning models by incorporating quantile regression, thus providing a more nuanced understanding of predictive values. We provide an implementation of prominent deep learning models for multi-step ahead time series prediction and evaluate their performance under high volatility and extreme conditions. We include multivariate and univariate modelling, strategies and provide a comparison with conventional deep learning models from the literature. Our models are tested on two cryptocurrencies: Bitcoin and Ethereum, using daily close-price data and selected benchmark time series datasets. The results show that integrating a quantile loss function with deep learning provides additional predictions for selected quantiles without a loss in the prediction accuracy when compared to the literature. Our quantile model has the ability to handle volatility more effectively and provides additional information for decision-making and uncertainty quantification through the use of quantiles when compared to conventional deep learning models.
We propose a novel nonparametric test to detect structural breaks in the conditional mean and/or variance of a time series. Our method does not assume any specific parametric form for the dependence structure of the regressor, the time series model, or the distribution of the noise. This flexibility allows our algorithm to be applicable to a wide range of framework. We further apply the proposed test to accurately localize the changepoints and establish theoretical guarantees showing that the estimated structural breaks are consistent, meaning they lie sufficiently close to the true breakpoints when a sufficiently large sample is available. The effectiveness of the proposed algorithm is demonstrated through an extensive simulation study encompassing a diverse range of time series structures, including light, moderately heavy, and heavy tailed distributions. We also show a real-life example, where an application to Bitcoin prices and Google search volume illustrates how the procedure can identify changes in the conditional relationship between market attention and price dynamics.
In many online domains, Sybil networks -- or cases where a single user assumes multiple identities -- is a pervasive feature. This complicates experiments, as off-the-shelf regression estimators at least assume known network topologies (if not fully independent observations) when Sybil network topologies in practice are often unknown. The literature has exclusively focused on techniques to detect Sybil networks, leading many experimenters to subsequently exclude suspected networks entirely before estimating treatment effects. I present a more efficient solution in the presence of these suspected Sybil networks: a weighted regression framework that applies weights based on the probabilities that sets of observations are controlled by single actors. I show in the paper that the MSE-minimizing solution is to set the weight matrix equal to the inverse of the expected network topology. I demonstrate the methodology on simulated data, and then I apply the technique to a competition with suspected Sybil networks run on the Sui blockchain and show reductions in the standard error of the estimate by 6 - 24%.
Based on a continuous-time stochastic volatility model with a linear drift, we develop a test for explosive behavior in financial asset prices at a low frequency when prices are sampled at a higher frequency. The test exploits the volatility information in the high-frequency data. The method consists of devolatizing log-asset price increments with realized volatility measures and performing a supremum-type recursive Dickey-Fuller test on the devolatized sample. The proposed test has a nuisance-parameter-free asymptotic distribution and is easy to implement. We study the size and power properties of the test in Monte Carlo simulations. A real-time date-stamping strategy based on the devolatized sample is proposed for the origination and conclusion dates of the explosive regime. Conditions under which the real-time date-stamping strategy is consistent are established. The test and the date-stamping strategy are applied to study explosive behavior in cryptocurrency and stock markets.
In this study, we introduce the first-of-its-kind class of tests for detecting change points in the distribution of a sequence of independent matrix-valued random variables. The tests are constructed using the weighted square integral difference of the empirical orthogonal Hankel transforms. The test statistics have a convenient closed-form expression, making them easy to implement in practice. We present their limiting properties and demonstrate their quality through an extensive simulation study. We utilize these tests for change point detection in cryptocurrency markets to showcase their practical use. The detection of change points in this context can have various applications in constructing and analyzing novel trading systems.
Given the high volatility and susceptibility to extreme events in the cryptocurrency market, forecasting tail risk is of paramount importance. Value-at-Risk (VaR), a quantile-based risk measure, is widely used for assessing tail risk and is central to monitoring financial market stability. In data-rich environments, functional data from various domains are employed to forecast conditional quantiles. However, the infinite-dimensional nature of functional data introduces uncertainty. This paper addresses this uncertainty problem by proposing a novel data-driven conditional quantile model averaging (MA) approach. With a set of candidate models varying by the number of components, MA assigns weights to each model determined by a K-fold cross-validation criterion. We prove the asymptotic optimality of the selected weights in terms of minimizing the excess final prediction error when all candidate models are misspecified. Additionally, when the true regression relationship belongs to the set of candidate models, we provide consistency results for the averaged estimators. Numerical studies indicate that, in most cases, the proposed method outperforms other model selection and averaging methods, particularly for extreme quantiles in cryptocurrency markets.
This paper introduces a novel two-sample test for a broad class of orthogonally equivalent positive definite symmetric matrix distributions. Our test is the first of its kind and we derive its asymptotic distribution. To estimate the test power, we use a warp-speed bootstrap method and consider the most common matrix distributions. We provide several real data examples, including the data for main cryptocurrencies and stock data of major US companies. The real data examples demonstrate the applicability of our test in the context closely related to algorithmic trading. The popularity of matrix distributions in many applications and the need for such a test in the literature are reconciled by our findings.
Data objects taking value in a general metric space have become increasingly common in modern data analysis. In this paper, we study two important statistical inference problems, namely, two-sample testing and change-point detection, for such non-Euclidean data under temporal dependence. Typical examples of non-Euclidean valued time series include yearly mortality distributions, time-varying networks, and covariance matrix time series. To accommodate unknown temporal dependence, we advance the self-normalization (SN) technique (Shao, 2010) to the inference of non-Euclidean time series, which is substantially different from the existing SN-based inference for functional time series that reside in Hilbert space (Zhang et al., 2011). Theoretically, we propose new regularity conditions that could be easier to check than those in the recent literature, and derive the limiting distributions of the proposed test statistics under both null and local alternatives. For change-point detection problem, we also derive the consistency for the change-point location estimator, and combine our proposed change-point test with wild binary segmentation to perform multiple change-point estimation. Numerical simulations demonstrate the effectiveness and robustness of our proposed tests compared with existing methods in the literature. Finally, we apply our tests to two-sample inference in mortality data and change-point detection in cryptocurrency data.