Cryptocurrencies have become a trendy topic recently, primarily due to their disruptive potential and reports of unprecedented returns. In addition, academics increasingly acknowledge the predictive power of Social Media in many fields and, more specifically, for financial markets and economics. In this paper, we leverage the predictive power of Twitter and Reddit sentiment together with Google Trends indexes and volume to forecast the log returns of ten cryptocurrencies. Specifically, we consider $Bitcoin$, $Ethereum$, $Tether$, $Binance Coin$, $Litecoin$, $Enjin Coin$, $Horizen$, $Namecoin$, $Peercoin$, and $Feathercoin$. We evaluate the performance of LASSO-VAR using daily data from January 2018 to January 2022. In a 30 days recursive forecast, we can retrieve the correct direction of the actual series more than 50% of the time. We compare this result with the main benchmarks, and we see a 10% improvement in Mean Directional Accuracy (MDA). The use of sentiment and attention variables as predictors increase significantly the forecast accuracy in terms of MDA but not in terms of Root Mean Squared Errors. We perform a Granger causality test using a post-double LASSO selection for high-dimensional VARs. Results show no "causality" from Social Media sentiment to cryptocurrencies returns
Jan 1, 2022·Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
The rapid spread of information over social media influences quantitative trading and investments. The growing popularity of speculative trading of highly volatile assets such as cryptocurrencies and meme stocks presents a fresh challenge in the financial realm. Investigating such "bubbles" - periods of sudden anomalous behavior of markets are critical in better understanding investor behavior and market dynamics. However, high volatility coupled with massive volumes of chaotic social media texts, especially for underexplored assets like cryptocoins pose a challenge to existing methods. Taking the first step towards NLP for cryptocoins, we present and publicly release CryptoBubbles, a novel multi-span identification task for bubble detection, and a dataset of more than 400 cryptocoins from 9 exchanges over five years spanning over two million tweets. Further, we develop a set of sequence-to-sequence hyperbolic models suited to this multi-span identification task based on the power-law dynamics of cryptocurrencies and user behavior on social media. We further test the effectiveness of our models under zero-shot settings on a test set of Reddit posts pertaining to 29 "meme stocks'', which see an increase in trade volume due to social media hype. Through quantitative, qualitative, and zero-shot analyses on Reddit and Twitter spanning cryptocoins and meme-stocks, we show the practical applicability of CryptoBubbles and hyperbolic models.
Blockchain finance has become a part of the world financial system, most typically manifested in the attention to the price of Bitcoin. However, a great deal of work is still limited to using technical indicators to capture Bitcoin price fluctuation, with little consideration of historical relationships and interactions between related cryptocurrencies. In this work, we propose a generic Cross-Cryptocurrency Relationship Mining module, named C2RM, which can effectively capture the synchronous and asynchronous impact factors between Bitcoin and related Altcoins. Specifically, we utilize the Dynamic Time Warping algorithm to extract the lead-lag relationship, yielding Lead-lag Variance Kernel, which will be used for aggregating the information of Altcoins to form relational impact factors. Comprehensive experimental results demonstrate that our C2RM can help existing price prediction methods achieve significant performance improvement, suggesting the effectiveness of Cross-Cryptocurrency interactions on benefitting Bitcoin price prediction.
Crypto-coins (also known as cryptocurrencies) are tradable digital assets. Notable examples include Bitcoin, Ether and Litecoin. Ownerships of cryptocoins are registered on distributed ledgers (i.e., blockchains). Secure encryption techniques guarantee the security of the transactions (transfers of coins across owners), registered into the ledger. Cryptocoins are exchanged for specific trading prices. While history has shown the extreme volatility of such trading prices across all different sets of crypto-assets, it remains unclear what and if there are tight relations between the trading prices of different cryptocoins. Major coin exchanges (i.e., Coinbase) provide trend correlation indicators to coin owners, suggesting possible acquisitions or sells. However, these correlations remain largely unvalidated. In this paper, we shed lights on the trend correlations across a large variety of cryptocoins, by investigating their coin-price correlation trends over a period of two years. Our experimental results suggest strong correlation patterns between main coins (Ethereum, Bitcoin) and alt-coins. We believe our study can support forecasting techniques for time-series modeling in the context of crypto-coins. We release our dataset and code to reproduce our analysis to the research community.
This paper uses new and recently established methodologies to study the evolutionary dynamics of the cryptocurrency market, and compares the findings with that of the equity market. We begin by applying random matrix theory and principal components analysis (PCA) to correlation matrices of both collections, highlighting clear differences in the eigenspectra exhibited. We then explore the heterogeneity of both asset classes, studying the time-varying dynamics of underlying sector behaviours, and determine the collective similarity within each collection. We then turn to a study of structural break dynamics and evolutionary power spectra, where we quantify the collective affinity in structural breaks and evolutionary behaviours of underlying sector time series. Finally, we implement two algorithms simulating `portfolio choice' dynamics to compare the effectiveness of stock selection and sector allocation in cryptocurrency portfolios. There, we highlight the importance of both endeavours and comment on noteworthy implications for cryptocurrency portfolio management.
The objective of this paper is to assess the performances of dimensionality reduction techniques to establish a link between cryptocurrencies. We have focused our analysis on the two most traded cryptocurrencies: Bitcoin and Ethereum. To perform our analysis, we took log returns and added some covariates to build our data set. We first introduced the pearson correlation coefficient in order to have a preliminary assessment of the link between Bitcoin and Ethereum. We then reduced the dimension of our data set using canonical correlation analysis and principal component analysis. After performing an analysis of the links between Bitcoin and Ethereum with both statistical techniques, we measured their performance on forecasting Ethereum returns with Bitcoin s features.
Jarosław Kwapień, Marcin Wątorek, Stanisław Drożdż
Time series of price returns for 80 of the most liquid cryptocurrencies listed on Binance are investigated for the presence of detrended cross-correlations. A spectral analysis of the detrended correlation matrix and a topological analysis of the minimal spanning trees calculated based on this matrix are applied for different positions of a moving window. The cryptocurrencies become more strongly cross-correlated among themselves than they used to be before. The average cross-correlations increase with time on a specific time scale in a way that resembles the Epps effect amplification when going from past to present. The minimal spanning trees also change their topology and, for the short time scales, they become more centralized with increasing maximum node degrees, while for the long time scales they become more distributed, but also more correlated at the same time. Apart from the inter-market dependencies, the detrended cross-correlations between the cryptocurrency market and some traditional markets, like the stock markets, commodity markets, and Forex, are also analyzed. The cryptocurrency market shows higher levels of cross-correlations with the other markets during the same turbulent periods, in which it is strongly cross-correlated itself.
We study the information dynamics between the largest Bitcoin exchange markets during the bubble in 2017-2018. By analysing high-frequency market-microstructure observables with different information theoretic measures for dynamical systems, we find temporal changes in information sharing across markets. In particular, we study the time-varying components of predictability, memory, and synchronous coupling, measured by transfer entropy, active information storage, and multi-information. By comparing these empirical findings with several models we argue that some results could relate to intra-market and inter-market regime shifts, and changes in direction of information flow between different market observables.
Using the asymmetric stochastic volatility model, this study investigates the day-of-the-week and holiday effects on the returns and volatility of Bitcoin from January 1, 2013 to August 31, 2019; in this context, we also discuss the characteristics of Bitcoin as a financial asset. The results of the estimation are threefold. First, the finding shows a small day-of-the week effect in volatility on Saturday and Sunday than in the rest of the week. Second, although the holiday effects are examined in active trading countries, namely Japan, China, Germany, and the United States, the positive post-holiday effect on the returns and weak positive pre-holiday effect on the volatility are only observed in the United States. Finally, the asymmetry effect is not observed. A comparison of Bitcoin to several assets such as stock, currency, and gold shows Bitcoin's positioning between stock, currency, and gold in relation to the week and holiday effects, its reaction to federal funds and medium of exchange characteristics, and the lack of asymmetry effect.
We develop an analysis of the cryptocurrency market borrowing methods and concepts from ecology. This approach makes it possible to identify specific diversity patterns and their variation, in close analogy with ecological systems, and to characterize the cryptocurrency market in an effective way. At the same time, it shows how non-biological systems can have an important role in contrasting different ecological theories and in testing the use of neutral models. The study of the cryptocurrencies abundance distribution and the evolution of the community structure strongly indicates that these statistical patterns are not consistent with neutrality. In particular, the necessity to increase the temporal change in community composition when the number of cryptocurrencies grows, suggests that their interactions are not necessarily weak. The analysis of the intraspecific and interspecific interdependency supports this fact and demonstrates the presence of a market sector influenced by mutualistic relations. These latest findings challenge the hypothesis of weakly interacting symmetric species, the postulate at the heart of neutral models.
M. Eren Akbiyik, Mert Erkul, Killian Kaempf, Vaiva Vasiliauskaitė · 5 authors
Understanding the variations in trading price (volatility), and its response to exogenous information, is a well-researched topic in finance. In this study, we focus on finding stable and accurate volatility predictors for a relatively new asset class of cryptocurrencies, in particular Bitcoin, using deep learning representations of public social media data obtained from Twitter. For our experiments, we extracted semantic information and user statistics from over 30 million Bitcoin-related tweets, in conjunction with 15-minute frequency price data over a horizon of 144 days. Using this data, we built several deep learning architectures that utilized different combinations of the gathered information. For each model, we conducted ablation studies to assess the influence of different components and feature sets over the prediction accuracy. We found statistical evidences for the hypotheses that: (i) temporal convolutional networks perform significantly better than both classical autoregressive models and other deep learning-based architectures in the literature, and (ii) tweet author meta-information, even detached from the tweet itself, is a better predictor of volatility than the semantic content and tweet volume statistics. We demonstrate how different information sets gathered from social media can be utilized in different architectures and how they affect the prediction results. As an additional contribution, we make our dataset public for future research.
Abootaleb Shirvani, Stefan Mittnik, W. Brent Lindquist, Svetlozar T. Rachev
We propose a doubly subordinated Levy process, NDIG, to model the time series properties of the cryptocurrency bitcoin. NDIG captures the skew and fat-tailed properties of bitcoin prices and gives rise to an arbitrage free, option pricing model. In this framework we derive two bitcoin volatility measures. The first combines NDIG option pricing with the Cboe VIX model to compute an implied volatility; the second uses the volatility of the unit time increment of the NDIG model. Both are compared to a volatility based upon historical standard deviation. With appropriate linear scaling, the NDIG process perfectly captures observed, in-sample, volatility.
In December 2017, two leading derivative exchanges, CBOE and CME, introduced the first regulated Bitcoin futures. Our aim is estimating their causal impact on Bitcoin volatility and trading volume. Employing a new causal approach, C-ARIMA, we find that the CME future triggered an increase in both outcomes. There is also evidence of a positive volume-volatility relationship and that the effect on volatility was partially due to the higher trading volumes induced by the launch of the contract. After controlling for the effect on volumes, we find that the CME instrument caused Bitcoin volatility to increase by more than double.
This study examines the dynamic asset market linkages under the COVID-19 global pandemic based on market efficiency, in the sense of Fama (1970). Particularly, we estimate the joint degree of market efficiency by applying Ito et al.'s (2014; 2017) Generalized Least Squares-based time-varying vector autoregression model. The empirical results show that (1) the joint degree of market efficiency changes widely over time, as shown in Lo's (2004) adaptive market hypothesis, (2) the COVID-19 pandemic may eliminate arbitrage and improve market efficiency through enhanced linkages between the asset markets; and (3) the market efficiency has continued to decline due to the Bitcoin bubble that emerged at the end of 2020.
Bitcoin has attracted attention from different market participants due to unpredictable price patterns. Sometimes, the price has exhibited big jumps. Bitcoin prices have also had extreme, unexpected crashes. We test the predictive power of a wide range of determinants on bitcoins’ price direction under the continuous transfer entropy approach as a feature selection criterion. Accordingly, the statistically significant assets in the sense of permutation test on the nearest neighbour estimation of local transfer entropy are used as features or explanatory variables in a deep learning classification model to predict the price direction of bitcoin. The proposed variable selection do not find significative the explanatory power of NASDAQ and Tesla. Under different scenarios and metrics, the best results are obtained using the significant drivers during the pandemic as validation. In the test, the accuracy increased in the post-pandemic scenario of July 2020 to January 2021 without drivers. In other words, our results indicate that in times of high volatility, Bitcoin seems to self-regulate and does not need additional drivers to improve the accuracy of the price direction.
Natalia A. Van Heerden, Juan Cabral, Nadia Luczywo
In recent years, cryptocurrencies have gone from an obscure niche to a prominent place, with investment in these assets becoming increasingly popular. However, cryptocurrencies carry a high risk due to their high volatility. In this paper, criteria based on historical cryptocurrency data are defined in order to characterize returns and risks in different ways, in short time windows (7 and 15 days); then, the importance of criteria is analyzed by various methods and their impact is evaluated. Finally, the future plan is projected to use the knowledge obtained for the selection of investment portfolios by applying multi-criteria methods.
Renowned method of log-periodic power law(LPPL) is one of the few ways that a financial market crash could be predicted. Alongside with LPPL, this paper propose a novel method of stock market crash using white box model derived from simple assumptions about the state of rational bubble. By applying this model to Dow Jones Index and Bitcoin market price data, it is shown that the model successfully predicts some major crashes of both markets, implying the high sensitivity and generalization abilities of the model.
This paper introduces new methods to study behaviours among the 52 largest cryptocurrencies between 01-01-2019 and 30-06-2021. First, we explore evolutionary correlation behaviours and apply a recently proposed turning point algorithm to identify regimes in market correlation. Next, we inspect the relationship between collective dynamics and the cryptocurrency market size - revealing an inverse relationship between the size of the market and the strength of collective dynamics. We then explore the time-varying consistency of the relationships between cryptocurrencies' size and their returns and volatility. There, we demonstrate that there is greater consistency between size and volatility than size and returns. Finally, we study the spread of volatility behaviours across the market changing with time by examining the structure of Wasserstein distances between probability density functions of rolling volatility. We demonstrate a new phenomenon of increased uniformity in volatility during market crashes, which we term \emph{volatility dispersion}.
Price movement forecasting, aimed at predicting financial asset trends based on current market information, has achieved promising advancements through machine learning (ML) methods. Most existing ML methods, however, struggle with the extremely low signal-to-noise ratio and stochastic nature of financial data, often mistaking noises for real trading signals without careful selection of potentially profitable samples. To address this issue, we propose LARA, a novel price movement forecasting framework with two main components: Locality-Aware Attention (LA-Attention) and Iterative Refinement Labeling (RA-Labeling). (1) LA-Attention, enhanced by metric learning techniques, automatically extracts the potentially profitable samples through masked attention scheme and task-specific distance metrics. (2) RA-Labeling further iteratively refines the noisy labels of potentially profitable samples, and combines the learned predictors robust to the unseen and noisy samples. In a set of experiments on three real-world financial markets: stocks, cryptocurrencies, and ETFs, LARA significantly outperforms several machine learning based methods on the Qlib quantitative investment platform. Extensive ablation studies confirm LARA's superior ability in capturing more reliable trading opportunities.
The stock market has been a popular topic of interest in the recent past. The growth in the inflation rate has compelled people to invest in the stock and commodity markets and other areas rather than saving. Further, the ability of Deep Learning models to make predictions on the time series data has been proven time and again. Technical analysis on the stock market with the help of technical indicators has been the most common practice among traders and investors. One more aspect is the sentiment analysis - the emotion of the investors that shows the willingness to invest. A variety of techniques have been used by people around the globe involving basic Machine Learning and Neural Networks. Ranging from the basic linear regression to the advanced neural networks people have experimented with all possible techniques to predict the stock market. It's evident from recent events how news and headlines affect the stock markets and cryptocurrencies. This paper proposes an ensemble of state-of-the-art methods for predicting stock prices. Firstly sentiment analysis of the news and the headlines for the company Apple Inc, listed on the NASDAQ is performed using a version of BERT, which is a pre-trained transformer model by Google for Natural Language Processing (NLP). Afterward, a Generative Adversarial Network (GAN) predicts the stock price for Apple Inc using the technical indicators, stock indexes of various countries, some commodities, and historical prices along with the sentiment scores. Comparison is done with baseline models like - Long Short Term Memory (LSTM), Gated Recurrent Units (GRU), vanilla GAN, and Auto-Regressive Integrated Moving Average (ARIMA) model.
Marcin Wątorek, Jarosław Kwapień, Stanisław Drożdż
We analyze the price return distributions of currency exchange rates, cryptocurrencies, and contracts for differences (CFDs) representing stock indices, stock shares, and commodities. Based on recent data from the years 2017--2020, we model tails of the return distributions at different time scales by using power-law, stretched exponential, and $q$-Gaussian functions. We focus on the fitted function parameters and how they change over the years by comparing our results with those from earlier studies and find that, on the time horizons of up to a few minutes, the so-called "inverse-cubic power-law" still constitutes an appropriate global reference. However, we no longer observe the hypothesized universal constant acceleration of the market time flow that was manifested before in an ever faster convergence of empirical return distributions towards the normal distribution. Our results do not exclude such a scenario but, rather, suggest that some other short-term processes related to a current market situation alter market dynamics and may mask this scenario. Real market dynamics is associated with a continuous alternation of different regimes with different statistical properties. An example is the COVID-19 pandemic outburst, which had an enormous yet short-time impact on financial markets. We also point out that two factors -- speed of the market time flow and the asset cross-correlation magnitude -- while related (the larger the speed, the larger the cross-correlations on a given time scale), act in opposite directions with regard to the return distribution tails, which can affect the expected distribution convergence to the normal distribution.
Purpose This paper aims to demonstrate a dynamic cointegration-based pairs trading strategy, including an optimal look-back window framework in the cryptocurrency market and evaluate its return and risk by applying three different scenarios. Design/methodology/approach This study uses the Engle-Granger methodology, the Kapetanios-Snell-Shin test and the Johansen test as cointegration tests in different scenarios. This study calibrates the mean-reversion speed of the Ornstein-Uhlenbeck process to obtain the half-life used for the asset selection phase and look-back window estimation. Findings By considering the main limitations in the market microstructure, the strategy of this paper exceeds the naive buy-and-hold approach in the Bitmex exchange. Another significant finding is that this study implements a numerous collection of cryptocurrency coins to formulate the model’s spread, which improves the risk-adjusted profitability of the pairs trading strategy. Besides, the strategy’s maximum drawdown level is reasonably low, which makes it useful to be deployed. The results also indicate that a class of coins has better potential arbitrage opportunities than others. Originality/value This research has some noticeable advantages, making it stand out from similar studies in the cryptocurrency market. First is the accuracy of data in which minute-binned data create the signals in the formation period. Besides, to backtest the strategy during the trading period, this study simulates the trading signals using best bid/ask quotes and market trades. This study exclusively takes the order execution into account when the asset size is already available at its quoted price (with one or more period gaps after signal generation). This action makes the backtesting much more realistic.
In recent years, Bitcoin price prediction has attracted the interest of researchers and investors. However, the accuracy of previous studies is not well enough. Machine learning and deep learning methods have been proved to have strong prediction ability in this area. This paper proposed a method combined with Ensemble Empirical Mode Decomposition (EEMD) and a deep learning method called long short-term memory (LSTM) to research the problem of next-day Bitcoin price forecast.
Roberto Mota Navarro, Paulino Monroy Castillero, Francois Leyvraz
Several studies have shown that large changes in the returns of an asset are associated with the sized of the gaps present in the order book In general, these associations have been studied without explicitly considering the dynamics of either gaps or returns. Here we present a study of these relationships. Our results suggest that the causal relationship between gaps and returns is limited to instantaneous causation.