Michaël Allouche, Mnacho Echenim, Emmanuel Gobet, Anne-Claire Maurice
ABSTRACT We study price aggregation methodologies applied to crypto‐currency prices with quotations fragmented on different platforms. An intrinsic difficulty is that the price returns and volumes are heavy‐tailed, with many outliers, making averaging and aggregation challenging. While conventional methods rely on volume‐weighted average prices (called VWAPs), or volume‐weighted median prices (called VWMs), we develop a new robust weighted median (RWM) estimator that is robust to price and volume outliers. Our study is based on new probabilistic concentration inequalities for weighted means and weighted quantiles under different tail assumptions (heavy tails, sub‐gamma tails, sub‐Gaussian tails). This justifies that fluctuations of VWAP and VWM are statistically important given the heavy‐tailed properties of volumes and/or prices. We show that our RWM estimator overcomes this problem and also satisfies all the desirable properties of a price aggregator. We illustrate the behavior of RWM on synthetic data (within a parametric model close to real data): Our estimator achieves a statistical accuracy twice as good as its competitors, and also allows to recover realized volatilities in a very accurate way. Tests on real data are also performed and confirm the good behavior of the estimator on various use cases.
Ali Yeganeh, Xuelong Hu, Sandile Charles Shongwe, Frans F. Koning
In the area of multivariate process quality control, it is sometimes important to monitor the ratio of two normal random variables denoted by RZ over time. The concept of control charts has often been harnessed in this field, leading to the application of various types of statistical models, including Shewhart, Exponentially Weighted Moving Average (EWMA), and so forth. However, there is little attention to implementation of machine learning-based control charts. To bridge this gap, a novel machine learning based model incorporating the attention mechanism approach, as an implemented Artificial Intelligence (AI) model, is proposed to monitor the RZ in Phase II applications. The proposed RZ method not only provides quicker Out-of-Control (OC) shift detection than conventional RZ control charts but also does not require the quality controller to have any prior information about the upward or downward shift patterns, which is a major assumption in most of the previous RZ models. We provide extensive performance comparison results to discuss the statistical performance of our proposed method through Monte Carlo simulations. Moreover, a comprehensive real example about surveillance of the cryptocurrency market is provided to illustrate the practical application of our proposed method. Through simulation and back-testing results, it is shown how the proposed method can lead to an automated trading strategy.
Most of cryptocurrencies such as Bitcoin and Ethereum are produced by mining pools, and Block Withholding attack (BWH) takes the mining pools as the target. As far as we know, the detection of the BWH attack is still an open problem. In this work, we propose a new method-Hybrid Statistical Test and Cross-Check (HSTCC) for detecting BWH attack. In the statistical test phase, we introduce the Poisson distribution cumulative probability test function to judge the expected income and actual income of the miners, so as to find out the suspicious miners with possible BWH attacks. In the cross-checking phase, inspired by the honeypot idea, we distributed a template that could quickly find the full proof of work to the miner to cross-check whether the miner was an attacker. The experimental results on the dataset of real bitcoin mining pool Eligius show that compared with the traditional statistical test method of 7.7% precision, the proposed HSTCC method can accurately find the miners who implement BWH attack; compared with the cross-check method that all samples need to be verified, the HSTCC method reduces the check samples by 77.6%.
Given the high volatility and susceptibility to extreme events in the cryptocurrency market, forecasting tail risk is of paramount importance. Value-at-Risk (VaR), a quantile-based risk measure, is widely used for assessing tail risk and is central to monitoring financial market stability. In data-rich environments, functional data from various domains are employed to forecast conditional quantiles. However, the infinite-dimensional nature of functional data introduces uncertainty. This paper addresses this uncertainty problem by proposing a novel data-driven conditional quantile model averaging (MA) approach. With a set of candidate models varying by the number of components, MA assigns weights to each model determined by a K-fold cross-validation criterion. We prove the asymptotic optimality of the selected weights in terms of minimizing the excess final prediction error when all candidate models are misspecified. Additionally, when the true regression relationship belongs to the set of candidate models, we provide consistency results for the averaged estimators. Numerical studies indicate that, in most cases, the proposed method outperforms other model selection and averaging methods, particularly for extreme quantiles in cryptocurrency markets.
The implementation of statistical techniques in on-line surveillance of financial markets has been frequently studied more recently. As a novel approach, statistical control charts which are famous tools for monitoring industrial processes, have been applied in various financial applications in the last three decades. The aim of this study is to propose a novel application of control charts called profile monitoring in the surveillance of the cryptocurrency markets. In this way, a new control chart is proposed to monitor the price variation of a pair of two most famous cryptocurrencies i.e., Bitcoin (BTC) and Ethereum (ETH). Parameter estimation, tuning and sensitivity analysis are conducted assuming that the random explanatory variable follows a symmetric normal distribution. The triggered signals from the proposed method are interpreted to convert the BTC and ETH at proper times to increase their total value. Hence, the proposed method could be considered a financial indicator so that its signal can lead to a tangible increase of the pair of assets. The performance of the proposed method is investigated through different parameter adjustments and compared with some common technical indicators under a real data set. The results show the acceptable and superior performance of the proposed method.
In this thesis, our objective is to study the relationship between transaction price and volume in the BTC/USD Coinbase exchange. In the second chapter, we develop a consecutive CUSUM algorithm to detect instantaneous changes in the arrival rate of market orders. We begin by estimating a baseline rate using the assumption of a local time-homogeneous Poisson process. Our observations lead us to reject the plausibility of a time-homogeneous Poisson model on a more global scale by using a chi squared test. We thus proceed to use CUSUM-based alarms to detect consecutive upward and downward changes in the arrival rate of market orders. In the third chapter we identify active periods from the number of consecutive upward CUSUM alarms, leading to the classification of active versus inactive periods. Finally we use One-Way ANOVA to assess the level effect on price swings for periods classified as containing at least two or three consecutive CUSUM up alarms. We show that in these active periods, price swings are significantly larger than in inactive periods.
Professor Keeping's book is a text for a one-year course (90-100 hours) for students having a knowledge of elementary calculus-second or third year students.It is unusually complete in that it is difficult to think of a topic which is not treated, at least briefly, but which one might like to see included in such a course.As would be expected of a widely ranging book at this level, many results are stated without proof but it is by no means a " how to do it" book.In addition to the usual elementary probability theory, standard distributions, and classical estimation and testing, one finds, e.g. the cumulants and Ar-statistics, sampling techniques, sequential and nonparametric procedures, fixed, random and mixed models as well as latin square and incomplete block designs considered, and a last chapter which looks at multivariate problems and introduces stochastic processes.There is a laudable concern for the power of the tests discussed and the required non-central distributions are introduced.Appropriate tables, a large number of exercises (with answers) and a thirty page appendix on various mathematical topics are included as well.The price paid for the virtue of comprehensiveness is, of course, the brevity of some particular parts; one cannot have everything.However, one might reasonably suggest that the briefer the treatment the more precise should be the statements.This book is somewhat marred by puzzling, misleading, or false statements, e.g. both the sample and population moments are defined to be the " rth moment of X about zero"; "If T is sufficient, so is any function of T" (p.125); the variance of a maximum likelihood estimator is asserted to be the Cramer-Rao lower bound; in discussing the Mann-Whitney U-test, it is not clear at given points just what alternatives are being considered and while one statistic is described as the test statistic, we are instructed to reject for small values of another.