Non-fungible tokens (NFT) have recently emerged as a novel blockchain-hosted financial asset class that has attracted major transaction volumes. However, preprocessing and analysis of NFT transaction data, which investors often rely on for their investment decisions, pose several challenges not commonly encountered in traditional financial data. These challenges arise mainly due to the non-fungible nature of NFTs as well as the intrinsic characteristics of the blockchain, the primary data source for NFT transactions. Using data consisting of the transaction history of eight highly valued NFT collections, a selection of such challenges is illustrated. These include price differentiation by token traits, the possible existence of lateral swaps and wash trades in the transaction history, and finally, severe price volatility. This paper provides an overall summary of the challenges associated with data analytics on NFT transaction data and lay a foundation for future research on the topic.
Abstract Recent studies about cryptocurrency returns show that their distribution can be highly-peaked, skewed, and heavy-tailed, with a large excess kurtosis. To accommodate all these peculiarities, we propose the asymmetric Laplace scale mixture (ALSM) family of distributions. Each member of the family is obtained by dividing the scale parameter of the conditional asymmetric Laplace (AL) distribution by a convenient mixing random variable taking values on all or part of the positive real line and whose distribution depends on a parameter vector $$\varvec{\theta }$$ <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML"> <mml:mrow> <mml:mi>θ</mml:mi> </mml:mrow> </mml:math> providing greater flexibility to the resulting ALSM. Advantageously concerning the AL distribution, our family members allow for a wider range of values for skewness and kurtosis. For illustrative purposes, we consider different mixing distributions; they give rise to ALSMs having a closed-form probability density function where the AL distribution is obtained as a special case under a convenient choice of $$\varvec{\theta }$$ <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML"> <mml:mrow> <mml:mi>θ</mml:mi> </mml:mrow> </mml:math> . We examine some properties of our ALSMs such as hierarchical and stochastic representations and moments of practical interest. We describe an EM algorithm to obtain maximum likelihood estimates of the parameters for all the considered ALSMs. We fit these models to the returns of two cryptocurrencies, considering several classical distributions for comparison. The analysis shows how our models represent a valid alternative to the considered competitors in terms of AIC, BIC, and likelihood-ratio tests.
Marcin Wątorek, Jarosław Kwapień, Stanisław Drożdż
Unlike price fluctuations, the temporal structure of cryptocurrency trading has seldom been a subject of systematic study. In order to fill this gap, we analyse detrended correlations of the price returns, the average number of trades in time unit, and the traded volume based on high-frequency data representing two major cryptocurrencies: bitcoin and ether. We apply the multifractal detrended cross-correlation analysis, which is considered the most reliable method for identifying nonlinear correlations in time series. We find that all the quantities considered in our study show an unambiguous multifractal structure from both the univariate (auto-correlation) and bivariate (cross-correlation) perspectives. We looked at the bitcoin--ether cross-correlations in simultaneously recorded signals, as well as in time-lagged signals, in which a time series for one of the cryptocurrencies is shifted with respect to the other. Such a shift suppresses the cross-correlations partially for short time scales, but does not remove them completely. We did not observe any qualitative asymmetry in the results for the two choices of a leading asset. The cross-correlations for the simultaneous and lagged time series became the same in magnitude for the sufficiently long scales.
Stablecoins, digital assets pegged to a specific currency or commodity value, are heavily involved in transactions of major cryptocurrencies. The effects of deviations from their desired fixed values (depeggings) on the cryptocurrencies for which they are frequently used in transactions are therefore of interest to study. We propose a model for this phenomenon using a multivariate mutually-exciting Hawkes process, and present a numerical example applying this model to Tether (USDT) and Bitcoin (BTC).
We propose a portfolio allocation method based on risk factor budgeting using convex Nonnegative Matrix Factorization (NMF). Unlike classical factor analysis, PCA, or ICA, NMF ensures positive factor loadings to obtain interpretable long-only portfolios. As the NMF factors represent separate sources of risk, they have a quasi-diagonal correlation matrix, promoting diversified portfolio allocations. We evaluate our method in the context of volatility targeting on two long-only global portfolios of cryptocurrencies and traditional assets. Our method outperforms classical portfolio allocations regarding diversification and presents a better risk profile than hierarchical risk parity (HRP). We assess the robustness of our findings using Monte Carlo simulation.
The paper constructs a multi-variate Hawkes process model of Bitcoin block arrivals and price jumps. Hawkes processes are selfexciting point processes that can capture the self- and cross-excitation effects of block mining and Bitcoin price volatility. We use publicly available blockchain datasets to estimate the model parameters via maximum likelihood estimation. The results show that Bitcoin price volatility boost block mining rate and Bitcoin investment return demonstrates mean reversion. Quantile-Quantile plots show that the proposed Hawkes process model is a better fit to the blockchain datasets than a Poisson process model.
Sandra Johnson, David Hyland-Wood, Anders L. Madsen, Kerrie Mengersen
The concept of 'Stateless Ethereum' was conceived with the primary aim of mitigating Ethereum's unbounded state growth. The key facilitator of Stateless Ethereum is through the introduction of 'witnesses' into the ecosystem. The changes and potential consequences that these additional data packets pose on the network need to be identified and analysed to ensure that the Ethereum ecosystem can continue operating securely and efficiently. In this paper we propose a Bayesian Network model, a probabilistic graphical modelling approach, to capture the key factors and their interactions in Ethereum mainnet, the public Ethereum blockchain, focussing on the changes being introduced by Stateless Ethereum to estimate the health of the resulting Ethereum ecosystem. We use a mixture of empirical data and expert knowledge, where data are unavailable, to quantify the model. Based on the data and expert knowledge available to use at the time of modelling, the Ethereum ecosystem is expected to remain healthy following the introduction of Stateless Ethereum.
Rebekka Buse, Konstantin Görgen, Melanie Schienle
We study the prediction of Value at Risk (VaR) for cryptocurrencies. In contrast to classic assets, returns of cryptocurrencies are often highly volatile and characterized by large fluctuations around single events. Analyzing a comprehensive set of 105 major cryptocurrencies, we show that Generalized Random Forests (GRF) (Athey, Tibshirani & Wager, 2019) adapted to quantile prediction have superior performance over other established methods such as quantile regression, GARCH-type and CAViaR models. This advantage is especially pronounced in unstable times and for classes of highly-volatile cryptocurrencies. Furthermore, we identify important predictors during such times and show their influence on forecasting over time. Moreover, a comprehensive simulation study also indicates that the GRF methodology is at least on par with existing methods in VaR predictions for standard types of financial returns and clearly superior in the cryptocurrency setup.
Manuel Febrero–Bande, Wenceslao González–Manteiga, Brenda Prallon, Yuri F. Saporito
This paper proposes a classification model for predicting the main activity of bitcoin addresses based on their balances. Since the balances are functions of time, we apply methods from functional data analysis; more specifically, the features of the proposed classification model are the functional principal components of the data. Classifying bitcoin addresses is a relevant problem for two main reasons: to understand the composition of the bitcoin market, and to identify addresses used for illicit activities. Although other bitcoin classifiers have been proposed, they focus primarily on network analysis rather than curve behavior. Our approach, on the other hand, does not require any network information for prediction. Furthermore, functional features have the advantage of being straightforward to build, unlike expert-built features. Results show improvement when combining functional features with scalar features, and similar accuracy for the models using those features separately, which points to the functional model being a good alternative when domain-specific knowledge is not available.
Value at risk and expected shortfall are increasingly popular tail risk measures in the financial risk management field. Both academia and financial institutions are working to improve tail risk forecasts in order to meet the requirements of the Basel Capital Accord; it states that one purpose of risk management and measuring risk accuracy is, since extreme movements cannot always be avoided, financial institutions can prepare for these extreme returns by capital allocation, and putting aside the appropriate amount of capital so as to avoid default in times of extreme price or index movements. Forecast combination has drawn much attention, as a combined forecast can outperform the individual forecasts under certain conditions. We propose two methodology, one is a semiparametric combination framework that can jointly produce combined value at risk and expected shortfall forecasts, another one is a parametric regression framework named as Quantile-ES regression that can produce combined expected shortfall forecasts. The favourability of the semiparametric combination framework has been presented via an empirical study - application in cryptocurrency markets with high-frequency data where the necessity of risk management application increases as the cryptocurrency market becomes more popular and mature. Additionally, the general framework of the parametric Quantile-ES regression has been presented via a simulation study, whereas it still need to be improved in the future. The contributions of this work include but are not limited to the enabling of the combination of expected shortfall forecasts and the application of risk management procedures in the cryptocurrency market with high-frequency data.
We investigate the instantaneous and limiting behavior of an n-node blockchain which is under continuous monitoring of the IT department of a company but faces non-stop cyber attacks from a single hacker. The blockchain is functional as far as no data stored on it has been changed, deleted, or locked. Once the IT department detects the attack from the hacker, it will immediately re-set the blockchain, rendering all previous efforts of the hacker in vain. The hacker will not stop until the blockchain is dysfunctional. For arbitrary distributions of the hacking times and detecting times, we derive the limiting functional probability, instantaneous functional probability, and mean functional time of the blockchain. We also show that all these quantities are increasing functions of the number of nodes, substantiating the intuition that the more nodes a blockchain has, the harder it is for a hacker to succeed in a cyber attack.
Non-Fungible Token (NFT) markets are one of the fastest growing digital markets today, with the sales during the third quarter of 2021 exceeding $10 billions! Nevertheless, these emerging markets - similar to traditional emerging marketplaces - can be seen as a great opportunity for illegal activities (e.g., money laundering, sale of illegal goods etc.). In this study we focus on a specific marketplace, namely NBA TopShot, that facilitates the purchase and (peer-to-peer) trading of sports collectibles. Our objective is to build a framework that is able to label peer-to-peer transactions on the platform as anomalous or not. To achieve our objective we begin by building a model for the profit to be made by selling a specific collectible on the platform. We then use RFCDE - a random forest model for the conditional density of the dependent variable - to model the errors from the profit models. This step allows us to estimate the probability of a transaction being anomalous. We finally label as anomalous any transaction whose aforementioned probability is less than 1%. Given the absence of ground truth for evaluating the model in terms of its classification of transactions, we analyze the trade networks formed from these anomalous transactions and compare it with the full trade network of the platform. Our results indicate that these two networks are statistically different when it comes to network metrics such as, edge density, closure, node centrality and node degree distribution. This network analysis provides additional evidence that these transactions do not follow the same patterns that the rest of the trades on the platform follow. However, we would like to emphasize here that this does not mean that these transactions are also illegal. These transactions will need to be further audited from the appropriate entities to verify whether or not they are illicit.
Jarosław Kwapień, Marcin Wątorek, Stanisław Drożdż
Time series of price returns for 80 of the most liquid cryptocurrencies listed on Binance are investigated for the presence of detrended cross-correlations. A spectral analysis of the detrended correlation matrix and a topological analysis of the minimal spanning trees calculated based on this matrix are applied for different positions of a moving window. The cryptocurrencies become more strongly cross-correlated among themselves than they used to be before. The average cross-correlations increase with time on a specific time scale in a way that resembles the Epps effect amplification when going from past to present. The minimal spanning trees also change their topology and, for the short time scales, they become more centralized with increasing maximum node degrees, while for the long time scales they become more distributed, but also more correlated at the same time. Apart from the inter-market dependencies, the detrended cross-correlations between the cryptocurrency market and some traditional markets, like the stock markets, commodity markets, and Forex, are also analyzed. The cryptocurrency market shows higher levels of cross-correlations with the other markets during the same turbulent periods, in which it is strongly cross-correlated itself.
Danial Saef, Odett Nagy, Sergej Sizov, Wolfgang Karl Härdle
While attention is a predictor for digital asset prices, and jumps in Bitcoin prices are well-known, we know little about its alternatives. Studying high frequency crypto data gives us the unique possibility to confirm that cross market digital asset returns are driven by high frequency jumps clustered around black swan events, resembling volatility and trading volume seasonalities. Regressions show that intra-day jumps significantly influence end of day returns in size and direction. This provides fundamental research for crypto option pricing models. However, we need better econometric methods for capturing the specific market microstructure of cryptos. All calculations are reproducible via the quantlet.com technology.
Decentralized control, low-complexity, flexible and efficient communications are the requirements of an architecture that aims to scale blockchains beyond the current state. Such properties are attainable by reducing ledger size and providing parallel operations in the blockchain. Sharding is one of the approaches that lower the burden of the nodes and enhance performance. However, the current solutions lack the features for resolving concurrency during cross-shard communications. With multiple participants belonging to different shards, handling concurrent operations is essential for optimal sharding. This issue becomes prominent due to the lack of architectural support and requires additional consensus for cross-shard communications. Relying on the advantages of hybrid Proof-of-Work/Proof-of-Stake (PoW/PoS), like Ethereum , hybrid consensus and 2-hop blockchain , we propose Reinshard , a new blockchain that inherits the properties of hybrid consensus for optimal sharding. Reinshard uses PoW and PoS chain-pairs with PoS sub-chains for all the valid chain-pairs where the hybrid consensus is attained through Verifiable Delay Function (VDF). Our architecture provides a secure method of arranging nodes in shards and resolves concurrency conflicts using the delay factor of VDF. The applicability of Reinshard is demonstrated through security and experimental evaluations. A practical concurrency problem is considered to show the efficacy of Reinshard in providing optimal sharding.
We consider the problem of correctly identifying the \textit{mode} of a discrete distribution $\mathcal{P}$ with sufficiently high probability by observing a sequence of i.i.d. samples drawn from $\mathcal{P}$. This problem reduces to the estimation of a single parameter when $\mathcal{P}$ has a support set of size $K = 2$. After noting that this special case is tackled very well by prior-posterior-ratio (PPR) martingale confidence sequences \citep{waudby-ramdas-ppr}, we propose a generalisation to mode estimation, in which $\mathcal{P}$ may take $K \geq 2$ values. To begin, we show that the "one-versus-one" principle to generalise from $K = 2$ to $K \geq 2$ classes is more efficient than the "one-versus-rest" alternative. We then prove that our resulting stopping rule, denoted PPR-1v1, is asymptotically optimal (as the mistake probability is taken to $0$). PPR-1v1 is parameter-free and computationally light, and incurs significantly fewer samples than competitors even in the non-asymptotic regime. We demonstrate its gains in two practical applications of sampling: election forecasting and verification of smart contracts in blockchains.
The tutor-web drilling system is designed for learning so there are typically no limits on the number of attempts at improving performance. This system is used at multiple schools and universities in Iceland and Kenya, mostly for mathematics and statistics. Students earn SmileyCoin, a cryptocurrency, while studying. In Iceland the system has typically been used by students who use their own devices to solve homework assignments during the semester, accessing the Internet-based tutor-web at http://tutor-web.net. These students typically take final exams on paper at the end of the semester. In Kenya the system is a part of a plan to enhance mathematics education using educational technology, organised by the Smiley Charity with the African Maths Initiative. This has been done by donating servers running the tutor-web to schools and tablets to students. Typically these schools do not have Internet access so the cryptocurrency can not be used. Innovative redesign was needed during COVID-19 in spring, 2020, since universities in Iceland were not able to host in-house finals and schools in Kenya were closed so tablets could not be donated directly to students. Remote finals were held in Iceland but the implementation was largely in the hands of the instructors. In Kenya, community libraries remained open and became a place for students to come in to study. Innovations included using the tutor-web as a remote drilling system in place of final exams in a large undergraduate course in statistics and donating tablets to libraries in Kenya. These libraries all have access to the Internet and the students have therefore been given the option to purchase the tablet using their SmileyCoin. This paper describes these implementations and how this unintended experiment will likely affect the future development and use of the tutor-web in both countries.
Aug 7, 2021·In: Awan, I., Benbernou, S., Younas, M., Aleksy, M. (eds) The International Conference on Deep Learning, Big Data and Blockchain (Deep-BDB 2021). Deep-BDB 2021. Lecture Notes in Networks and Systems, vol 309. Springer, Cham
On the Ethereum network, it is challenging to determine a gas price that ensures a transaction will be included in a block within a user's required timeline without overpaying. One way of addressing this problem is through the use of gas price oracles that utilize historical block data to recommend gas prices. However, when transaction volumes increase rapidly, these oracles often underestimate or overestimate the price. In this paper, we demonstrate how Gaussian process models can predict the distribution of the minimum price in an upcoming block when transaction volumes are increasing. This is effective because these processes account for time correlations between blocks. We performed an empirical analysis using the Gaussian process model on historical block data and compared the performance with GasStation-Express and Geth gas price oracles. The results suggest that when transactions volumes fluctuate greatly, the Gaussian process model offers a better estimation. Further, we demonstrated that GasStation-Express and Geth can be improved upon by using a smaller training sample size which is properly pre-processed. Based on the results of empirical analysis, we recommended a gas price oracle made up of a hybrid model consisting of both the Gaussian process and GasStation-Express. This oracle provides efficiency, accuracy, and better cost.
Dorcas Ofori-Boateng, Ignacio Segovia Dominguez, Murat Kantarcioglu, Cuneyt G. Akcora · 5 authors
Motivated by the recent surge of criminal activities with cross-cryptocurrency trades, we introduce a new topological perspective to structural anomaly detection in dynamic multilayer networks. We postulate that anomalies in the underlying blockchain transaction graph that are composed of multiple layers are likely to also be manifested in anomalous patterns of the network shape properties. As such, we invoke the machinery of clique persistent homology on graphs to systematically and efficiently track evolution of the network shape and, as a result, to detect changes in the underlying network topology and geometry. We develop a new persistence summary for multilayer networks, called stacked persistence diagram, and prove its stability under input data perturbations. We validate our new topological anomaly detection framework in application to dynamic multilayer networks from the Ethereum Blockchain and the Ripple Credit Network, and demonstrate that our stacked PD approach substantially outperforms state-of-art techniques.
Jun 2, 2021·Transactions on Mass-Data Analysis of Images and Signals P-ISSN1868-6451, E-ISSN 2509-9353, ISBN 978-3-942952-80-4 Volume 11 - Number 1 - September 2020 - Page 3-26
Data transmission between two or more digital devices in industry and government demands secure and agile technology. Digital information distribution often requires deployment of Internet of Things (IoT) devices and Data Fusion techniques which have also gained popularity in both, civilian and military environments, such as, emergence of Smart Cities and Internet of Battlefield Things (IoBT). This usually requires capturing and consolidating data from multiple sources. Because datasets do not necessarily originate from identical sensors, fused data typically results in a complex Big Data problem. Due to potentially sensitive nature of IoT datasets, Blockchain technology is used to facilitate secure sharing of IoT datasets, which allows digital information to be distributed, but not copied. However, blockchain has several limitations related to complexity, scalability, and excessive energy consumption. We propose an approach to hide information (sensor signal) by transforming it to an image or an audio signal. In one of the latest attempts to the military modernization, we investigate sensor fusion approach by investigating the challenges of enabling an intelligent identification and detection operation and demonstrates the feasibility of the proposed Deep Learning and Anomaly Detection models that can support future application for specific hand gesture alert system from wearable devices.
In this study, we study the price dynamics of cryptocurrencies using adaptive complementary ensemble empirical mode decomposition (ACE-EMD) and Hilbert spectral analysis. This is a multiscale noise-assisted approach that decomposes any time series into a number of intrinsic mode functions, along with the corresponding instantaneous amplitudes and instantaneous frequencies. The decomposition is adaptive to the time-varying volatility of each cryptocurrency price evolution. Different combinations of modes allow us to reconstruct the time series using components of different timescales. We then apply Hilbert spectral analysis to define and compute the instantaneous energy-frequency spectrum of each cryptocurrency to illustrate the properties of various timescales embedded in the original time series.
Marco Ortu, Nicola Uras, Claudio Conversano, Giuseppe Destefanis · 5 authors
This work aims to analyse the predictability of price movements of cryptocurrencies on both hourly and daily data observed from January 2017 to January 2021, using deep learning algorithms. For our experiments, we used three sets of features: technical, trading and social media indicators, considering a restricted model of only technical indicators and an unrestricted model with technical, trading and social media indicators. We verified whether the consideration of trading and social media indicators, along with the classic technical variables (such as price's returns), leads to a significative improvement in the prediction of cryptocurrencies price's changes. We conducted the study on the two highest cryptocurrencies in volume and value (at the time of the study): Bitcoin and Ethereum. We implemented four different machine learning algorithms typically used in time-series classification problems: Multi Layers Perceptron (MLP), Convolutional Neural Network (CNN), Long Short Term Memory (LSTM) neural network and Attention Long Short Term Memory (ALSTM). We devised the experiments using the advanced bootstrap technique to consider the variance problem on test samples, which allowed us to evaluate a more reliable estimate of the model's performance. Furthermore, the Grid Search technique was used to find the best hyperparameters values for each implemented algorithm. The study shows that, based on the hourly frequency results, the unrestricted model outperforms the restricted one. The addition of the trading indicators to the classic technical indicators improves the accuracy of Bitcoin and Ethereum price's changes prediction, with an increase of accuracy from a range of 51-55% for the restricted model, to 67-84% for the unrestricted model.