We determine the amount of information contained in a time series of price returns at a given time scale, by using a widespread tool of the information theory, namely the Shannon entropy, applied to a symbolic representation of this time series. By deriving the exact and the asymptotic distribution of this market information indicator in the case where the efficient market hypothesis holds, we develop a statistical test of market efficiency. We apply it to a real dataset of stock indices, single stock, and cryptocurrency, for which we are able to determine at each date whether the efficient market hypothesis is to be rejected, with respect to a given confidence level.
Antonio Briola, David Vidal-Tomás, Yuanrong Wang, Tomaso Aste
We quantitatively describe the main events that led to the Terra project's failure in May 2022. We first review, in a systematic way, news from heterogeneous social media sources; we discuss the fragility of the Terra project and its vicious dependence on the Anchor protocol. We hence identify the crash's trigger events, analysing hourly and transaction data for Bitcoin, Luna, and TerraUSD. Finally, using state-of-the-art techniques from network science, we study the evolution of dependency structures for 61 highly capitalised cryptocurrencies during the down-market and we also highlight the absence of herding behaviour analysing cross-sectional absolute deviation of returns.
Marcin Wątorek, Jarosław Kwapień, Stanisław Drożdż
Unlike price fluctuations, the temporal structure of cryptocurrency trading has seldom been a subject of systematic study. In order to fill this gap, we analyse detrended correlations of the price returns, the average number of trades in time unit, and the traded volume based on high-frequency data representing two major cryptocurrencies: bitcoin and ether. We apply the multifractal detrended cross-correlation analysis, which is considered the most reliable method for identifying nonlinear correlations in time series. We find that all the quantities considered in our study show an unambiguous multifractal structure from both the univariate (auto-correlation) and bivariate (cross-correlation) perspectives. We looked at the bitcoin--ether cross-correlations in simultaneously recorded signals, as well as in time-lagged signals, in which a time series for one of the cryptocurrencies is shifted with respect to the other. Such a shift suppresses the cross-correlations partially for short time scales, but does not remove them completely. We did not observe any qualitative asymmetry in the results for the two choices of a leading asset. The cross-correlations for the simultaneous and lagged time series became the same in magnitude for the sufficiently long scales.
Bitcoin is the first and highest valued cryptocurrency that stores transactions in a publicly distributed ledger called the blockchain. Understanding the activity and behavior of Bitcoin actors is a crucial research topic as they are pseudonymous in the transaction network. In this article, we propose a method based on taint analysis to extract taint flows --dynamic networks representing the sequence of Bitcoins transferred from an initial source to other actors until dissolution. Then, we apply graph embedding methods to characterize taint flows. We evaluate our embedding method with taint flows from top mining pools and show that it can classify mining pools with high accuracy. We also found that taint flows from the same period show high similarity. Our work proves that tracing the money flows can be a promising approach to classifying source actors and characterizing different money flow patterns
This study examines the weak form of the efficient market hypothesis for Bitcoin using a feedforward neural network. Due to the increasing popularity of cryptocurrencies in recent years, the question has arisen, as to whether market inefficiencies could be exploited in Bitcoin. Several studies we refer to here discuss this topic in the context of Bitcoin using either statistical tests or machine learning methods, mostly relying exclusively on data from Bitcoin itself. Results regarding market efficiency vary from study to study. In this study, however, the focus is on applying various asset-related input features in a neural network. The aim is to investigate whether the prediction accuracy improves when adding equity stock indices (S&P 500, Russell 2000), currencies (EURUSD), 10 Year US Treasury Note Yield as well as Gold&Silver producers index (XAU), in addition to using Bitcoin returns as input feature. As expected, the results show that more features lead to higher training performance from 54.6% prediction accuracy with one feature to 61% with six features. On the test set, we observe that with our neural network methodology, adding additional asset classes, no increase in prediction accuracy is achieved. One feature set is able to partially outperform a buy-and-hold strategy, but the performance drops again as soon as another feature is added. This leads us to the partial conclusion that weak market inefficiencies for Bitcoin cannot be detected using neural networks and the given asset classes as input. Therefore, based on this study, we find evidence that the Bitcoin market is efficient in the sense of the efficient market hypothesis during the sample period. We encourage further research in this area, as much depends on the sample period chosen, the input features, the model architecture, and the hyperparameters.
Ziqiao Ao, Lin William Cong, Gergely Horváth, Luyao Zhang
Decentralized finance (DeFi) has the potential to disrupt centralized finance by validating peer-to-peer transactions through tamper-proof smart contracts, thus significantly lowering the transaction cost charged by financial intermediaries. However, the actual realization of peer-to-peer transactions and the levels and effects of decentralization are largely unknown. Our research pioneers a blockchain network study that applies social network analysis to measure the level, dynamics, and impacts of decentralization in DeFi token transactions on the Ethereum blockchain. First, we find a significant core-periphery structure in the AAVE token transaction network where the cores include the two largest centralized crypto exchanges. Second, we provide evidence that multiple network features consistently characterize decentralization dynamics. Finally, we document that a more decentralized network significantly predicts a higher return and lower volatility of the decentralized market of AAVE tokens on the Ethereum blockchain. We point out that our approach is seminal for inspiring future extensions related to the facets of application scenarios, research questions, and methodologies on the mechanics of blockchain decentralization.
Kwapie\'n, Jaros{\l}aw, W\k{a}torek, Marcin, Marija Bezbradica, Martin Crane · 6 authors
We analyse tick-by-tick data representing major cryptocurrencies traded on some different cryptocurrency trading platforms. We focus on such quantities like the inter-transaction times, the number of transactions in time unit, the traded volume, and volatility. We show that the inter-transaction times show long-range power-law autocorrelations. These lead to multifractality expressed by the right-side asymmetry of the singularity spectra $f(\alpha)$ indicating that the periods of increased market activity are characterised by richer multifractality compared to the periods of quiet market. We also show that neither the stretched exponential distribution nor the power-law-tail distribution are able to model universally the cumulative distribution functions of the quantities considered in this work. For each quantity, some data sets can be modeled by the former, some data sets by the latter, while both fail in other cases. An interesting, yet difficult to account for, observation is that parallel data sets from different trading platforms can show disparate statistical properties.
We investigate logarithmic price returns cross-correlations at different time horizons for a set of 25 liquid cryptocurrencies traded on the FTX digital currency exchange. We study how the structure of the Minimum Spanning Tree (MST) and the Triangulated Maximally Filtered Graph (TMFG) evolve from high (15 s) to low (1 day) frequency time resolutions. For each horizon, we test the stability, statistical significance and economic meaningfulness of the networks. Results give a deep insight into the evolutionary process of the time dependent hierarchical organization of the system under analysis. A decrease in correlation between pairs of cryptocurrencies is observed for finer time sampling resolutions. A growing structure emerges for coarser ones, highlighting multiple changes in the hierarchical reference role played by mainstream cryptocurrencies. This effect is studied both in its pairwise realizations and intra-sector ones.
Bitcoin, with its ever-growing popularity, has demonstrated extreme price volatility since its origin. This volatility, together with its decentralised nature, make Bitcoin highly subjective to speculative trading as compared to more traditional assets. In this paper, we propose a multimodal model for predicting extreme price fluctuations. This model takes as input a variety of correlated assets, technical indicators, as well as Twitter content. In an in-depth study, we explore whether social media discussions from the general public on Bitcoin have predictive power for extreme price movements. A dataset of 5,000 tweets per day containing the keyword `Bitcoin' was collected from 2015 to 2021. This dataset, called PreBit, is made available online. In our hybrid model, we use sentence-level FinBERT embeddings, pretrained on financial lexicons, so as to capture the full contents of the tweets and feed it to the model in an understandable way. By combining these embeddings with a Convolutional Neural Network, we built a predictive model for significant market movements. The final multimodal ensemble model includes this NLP model together with a model based on candlestick data, technical indicators and correlated asset prices. In an ablation study, we explore the contribution of the individual modalities. Finally, we propose and backtest a trading strategy based on the predictions of our models with varying prediction threshold and show that it can used to build a profitable trading strategy with a reduced risk over a `hold' or moving average strategy.
Stablecoins, digital assets pegged to a specific currency or commodity value, are heavily involved in transactions of major cryptocurrencies. The effects of deviations from their desired fixed values (depeggings) on the cryptocurrencies for which they are frequently used in transactions are therefore of interest to study. We propose a model for this phenomenon using a multivariate mutually-exciting Hawkes process, and present a numerical example applying this model to Tether (USDT) and Bitcoin (BTC).
Blockchain introduces decentralized trust in peer-to-peer networks, advancing security and democratizing systems. Yet, a unified definition for decentralization remains elusive. Our Systematization of Knowledge (SoK) seeks to bridge this gap, emphasizing quantification and methodological coherence. We've formulated a taxonomy defining blockchain decentralization across five facets: consensus, network, governance, wealth, and transaction. Despite the prevalent focus on consensus decentralization, our novel index, based on Shannon entropy, provides comprehensive insights. Moreover, we delve into alternative metrics like the Gini and Nakamoto Coefficients and the Herfindahl-Hirschman Index (HHI), supplemented by an open-source Python tool on GitHub. In terms of methodology, blockchain research has often bypassed stringent scientific methods. By employing descriptive, predictive, and causal methods, our study showcases the potential of structured research in blockchain. Descriptively, we observe a trend of converging decentralization levels over time. Examining DeFi platforms reveals exchange and lending applications as more decentralized than their payment and derivatives counterparts. Predictively, there's a notable correlation between Ether's returns and transaction decentralization in Ether-backed stablecoins. Causally, Ethereum's transition to the EIP-1559 transaction fee model has a profound impact on DeFi transaction decentralization. To conclude, our work outlines directions for blockchain research, emphasizing the delicate balance among decentralization facets, fostering long-term decentralization, and the ties between decentralization, security, privacy, and efficiency. We end by spotlighting challenges in grasping blockchain decentralization intricacies.
The paper investigates the rich class of Generalized Tempered Stable distribution, an alternative to Normal distribution and the $α$-Stable distribution for modelling asset return and many physical and economic systems. Firstly, we explore some important properties of the Generalized Tempered Stable (GTS) distribution. The theoretical tools developed are used to perform empirical analysis. The GTS distribution is fitted using S&P 500, SPY ETF and Bitcoin BTC. The Fractional Fourier Transform (FRFT) technique evaluates the probability density function and its derivatives in the maximum likelihood procedure. Based on the results from the statistical inference and the Kolmogorov-Smirnov (K-S) goodness-of-fit, the GTS distribution fits the underlying distribution of the SPY ETF return. The right side of the Bitcoin BTC return, and the left side of the S&P 500 return underlying distributions fit the Tempered Stable distribution; while the left side of the Bitcoin BTC return and the right side of the S&P 500 return underlying distributions are modelled by the compound Poisson process
With the proliferation of pump-and-dump schemes (P&Ds) in the cryptocurrency market, it becomes imperative to detect such fraudulent activities in advance to alert potentially susceptible investors. In this paper, we focus on predicting the pump probability of all coins listed in the target exchange before a scheduled pump time, which we refer to as the target coin prediction task. Firstly, we conduct a comprehensive study of the latest 709 P&D events organized in Telegram from Jan. 2019 to Jan. 2022. Our empirical analysis reveals some interesting patterns of P&Ds, such as that pumped coins exhibit intra-channel homogeneity and inter-channel heterogeneity. Here channel refers a form of group in Telegram that is frequently used to coordinate P&D events. This observation inspires us to develop a novel sequence-based neural network, dubbed SNN, which encodes a channel's P&D event history into a sequence representation via the positional attention mechanism to enhance the prediction accuracy. Positional attention helps to extract useful information and alleviates noise, especially when the sequence length is long. Extensive experiments verify the effectiveness and generalizability of proposed methods. Additionally, we release the code and P&D dataset on GitHub: https://github.com/Bayi-Hu/Pump-and-Dump-Detection-on-Cryptocurrency, and regularly update the dataset.
We quantify Non Fungible Token (NFT) rarity and investigate how it impacts market behaviour by analysing a dataset of 3.7M transactions collected between January 2018 and June 2022, involving 1.4M NFTs distributed across 410 collections. First, we consider the rarity of an NFT based on the set of human-readable attributes it possesses and show that most collections present heterogeneous rarity patterns, with few rare NFTs and a large number of more common ones. Then, we analyze market performance and show that, on average, rarer NFTs: (i) sell for higher prices, (ii) are traded less frequently, (iii) guarantee higher returns on investment (ROIs), and (iv) are less risky, i.e., less prone to yield negative returns. We anticipate that these findings will be of interest to researchers as well as NFT creators, collectors, and traders.
Ji Liu, Zheng Xu, Yanmei Zhang, Wei Dai · 6 authors
Since the emergence of blockchain technology, its application in the financial market has always been an area of focus and exploration by all parties. With the characteristics of anonymity, trust, tamper-proof, etc., blockchain technology can effectively solve some problems faced by the financial market, such as trust issues and information asymmetry issues. To deeply understand the application scenarios of blockchain in the financial market, the issue of securities issuance and trading in the primary market is a problem that must be studied clearly. We conducted an empirical study to investigate the main difficulties faced by primary market participants in their business practices and the potential challenges of the deepening application of blockchain technology in the primary market. We adopted a hybrid method combining interviews (qualitative methods) and surveys (quantitative methods) to conduct this research in two stages. In the first stage, we interview 15 major primary market participants with different backgrounds and expertise. In the second phase, we conducted a verification survey of 54 primary market practitioners to confirm various insights from the interviews, including challenges and desired improvements. Our interviews and survey results revealed several significant challenges facing blockchain applications in the primary market: complex due diligence, mismatch, and difficult monitoring. On this basis, we believe that our future research can focus on some aspects of these challenges.
Anticipating price developments in financial markets is a topic of continued interest in forecasting. Funneled by advancements in deep learning and natural language processing (NLP) together with the availability of vast amounts of textual data in form of news articles, social media postings, etc., an increasing number of studies incorporate text-based predictors in forecasting models. We contribute to this literature by introducing weak learning, a recently proposed NLP approach to address the problem that text data is unlabeled. Without a dependent variable, it is not possible to finetune pretrained NLP models on a custom corpus. We confirm that finetuning using weak labels enhances the predictive value of text-based features and raises forecast accuracy in the context of predicting cryptocurrency returns. More fundamentally, the modeling paradigm we present, weak labeling domain-specific text and finetuning pretrained NLP models, is universally applicable in (financial) forecasting and unlocks new ways to leverage text data.
Our study empirically predicts the bubble of non-fungible tokens (NFTs): transferable and unique digital assets on public blockchains. This topic is important because, despite their strong market growth in 2021, NFTs on a project basis have not been investigated in terms of bubble prediction. Specifically, we applied the logarithmic periodic power law (LPPL) model to time-series price data associated with four major NFT projects. The results indicate that, as of December 20, 2021, (i) NFTs, in general, are in a small bubble (a price decline is predicted), (ii) the Decentraland project is in a medium bubble (a price decline is predicted), and (iii) the Ethereum Name Service and ArtBlocks projects are in a small negative bubble (a price increase is predicted). A future work will involve a prediction refinement considering the heterogeneity of NFTs, comparison with other methods, and the use of more enriched data.
Criminals have become increasingly experienced in using cryptocurrencies, such as Bitcoin, for money laundering. The use of cryptocurrencies can hide criminal identities and transfer hundreds of millions of dollars of dirty funds through their criminal digital wallets. However, this is considered a paradox because cryptocurrencies are goldmines for open-source intelligence, giving law enforcement agencies more power when conducting forensic analyses. This paper proposed Inspection-L, a graph neural network (GNN) framework based on a self-supervised Deep Graph Infomax (DGI) and Graph Isomorphism Network (GIN), with supervised learning algorithms, namely Random Forest (RF), to detect illicit transactions for anti-money laundering (AML). To the best of our knowledge, our proposal is the first to apply self-supervised GNNs to the problem of AML in Bitcoin. The proposed method was evaluated on the Elliptic dataset and shows that our approach outperforms the state-of-the-art in terms of key classification metrics, which demonstrates the potential of self-supervised GNN in the detection of illicit cryptocurrency transactions.
Rebekka Buse, Konstantin Görgen, Melanie Schienle
We study the prediction of Value at Risk (VaR) for cryptocurrencies. In contrast to classic assets, returns of cryptocurrencies are often highly volatile and characterized by large fluctuations around single events. Analyzing a comprehensive set of 105 major cryptocurrencies, we show that Generalized Random Forests (GRF) (Athey, Tibshirani & Wager, 2019) adapted to quantile prediction have superior performance over other established methods such as quantile regression, GARCH-type and CAViaR models. This advantage is especially pronounced in unstable times and for classes of highly-volatile cryptocurrencies. Furthermore, we identify important predictors during such times and show their influence on forecasting over time. Moreover, a comprehensive simulation study also indicates that the GRF methodology is at least on par with existing methods in VaR predictions for standard types of financial returns and clearly superior in the cryptocurrency setup.
Zeyd Boukhers, Azeddine Bouabdallah, Cong Yang, Jan Jürjens
Since Bitcoin first appeared on the scene in 2009, cryptocurrencies have become a worldwide phenomenon as important decentralized financial assets. Their decentralized nature, however, leads to notable volatility against traditional fiat currencies, making the task of accurately forecasting the crypto-fiat exchange rate complex. In this study, we examine the various independent factors that affect the Bitcoin-Dollar exchange rate's volatility. To this end, we propose CoMForE, a multimodal AdaBoost-LSTM ensemble model, which not only utilizes historical trading data but also incorporates public sentiments from related tweets, public interest demonstrated by search volumes, and blockchain hash-rate data. Our developed model goes a step further by predicting fluctuations in the overall cryptocurrency value distribution, thus increasing its value for investment decision-making. We have subjected this method to extensive testing via comprehensive experiments, thereby validating the importance of multimodal combination over exclusive reliance on trading data. Further experiments show that our method significantly surpasses existing forecasting tools and methodologies, demonstrating a 19.29% improvement. This result underscores the influence of external independent factors on cryptocurrency volatility.
The goals of this paper are twofold: (1) to present a new method that is able to find linear laws governing the time evolution of Markov chains and (2) to apply this method for anomaly detection in Bitcoin prices. To accomplish these goals, first, the linear laws of Markov chains are derived by using the time embedding of their (categorical) autocorrelation function. Then, a binary series is generated from the first difference of Bitcoin exchange rate (against the United States Dollar). Finally, the minimum number of parameters describing the linear laws of this series is identified through stepped time windows. Based on the results, linear laws typically became more complex (containing an additional third parameter that indicates hidden Markov property) in two periods: before the crash of cryptocurrency markets inducted by the COVID-19 pandemic (12 March 2020), and before the record-breaking surge in the price of Bitcoin (Q4 2020 - Q1 2021). In addition, the locally high values of this third parameter are often related to short-term price peaks, which suggests price manipulation.
Uniswap, like other DEXs, has gained much attention this year because it is a non-custodial and publicly verifiable exchange that allows users to trade digital assets without trusted third parties. However, its simplicity and lack of regulation also makes it easy to execute initial coin offering scams by listing non-valuable tokens. This method of performing scams is known as rug pull, a phenomenon that already existed in traditional finance but has become more relevant in DeFi. Various projects such as [34,37] have contributed to detecting rug pulls in EVM compatible chains. However, the first longitudinal and academic step to detecting and characterizing scam tokens on Uniswap was made in [44]. The authors collected all the transactions related to the Uniswap V2 exchange and proposed a machine learning algorithm to label tokens as scams. However, the algorithm is only valuable for detecting scams accurately after they have been executed. This paper increases their data set by 20K tokens and proposes a new methodology to label tokens as scams. After manually analyzing the data, we devised a theoretical classification of different malicious maneuvers in Uniswap protocol. We propose various machine-learning-based algorithms with new relevant features related to the token propagation and smart contract heuristics to detect potential rug pulls before they occur. In general, the models proposed achieved similar results. The best model obtained an accuracy of 0.9936, recall of 0.9540, and precision of 0.9838 in distinguishing non-malicious tokens from scams prior to the malicious maneuver.
In this study, we investigate the BTC price time-series (17 August 2010-27 June 2021) and show that the 2017 pricing episode is not unique. We describe at least ten new events, which occurred since 2010-2011 and span more than five orders of price magnitudes ($US 1 -$US 60k). We find that those events have a similar duration of approx. 50-100 days. Although we are not able to predict times of a price peak, we however succeed to approximate the BTC price evolution using a function that is similar to a Fibonacci sequence. Finally, we complete a comparison with other types of financial instruments (equities, currencies, gold) which suggests that BTC may be classified as an illiquid asset.