Blockchain Papers

Follow blockchain research across journals, conferences, and preprint repositories.

20 papersLast indexed Aug 31, 2026
Search papers

Paper index

20 results · page 1 of 1

Clear filters
Aug 26, 2025·Journal of theoretical and applied electronic commerce research
1 cites
RL-TweetGen: A Socio-Technical Framework for Engagement-Optimized Short Text Generation in Digital Commerce Using Large Language Models and Reinforcement Learning

S. Chitrakala, Pavithra S S

In the rapidly evolving landscape of digital marketing and electronic commerce, short-form content—particularly on platforms like Twitter (now X)—has become pivotal for real-time branding, community engagement, and product promotion. The rise of Non-Fungible Tokens (NFTs) and Web3 ecosystems further underscores the need for domain-specific, engagement-oriented social media content. However, automating the generation of such content while balancing linguistic quality, semantic relevance, and audience engagement remains a substantial challenge. To address this, we propose RL-TweetGen, a socio-technical framework that integrates instruction-tuned large language models (LLMs) with reinforcement learning (RL) to generate concise, impactful, and engagement-optimized tweets. The framework incorporates a structured pipeline comprising domain-specific data curation, semantic classification, and intent-aware prompt engineering, and leverages Parameter-Efficient Fine-Tuning (PEFT) with LoRA for scalable model adaptation. We fine-tuned and evaluated three LLMs—LLaMA-3.1-8B, Mistral-7B Instruct, and DeepSeek 7B Chat—guided by a hybrid reward function that blends XGBoost-predicted engagement scores with expert-in-the-loop feedback. To enhance lexical diversity and contextual alignment, we implemented advanced decoding strategies, including Tailored Beam Search, Enhanced Top-p Sampling, and Contextual Temperature Scaling. A case study focused on NFT-related tweet generation demonstrated the practical effectiveness of RL-TweetGen. Experimental results showed that Mistral-7B achieved the highest lexical fluency (BLEU: 0.2285), LLaMA-3.1 exhibited superior semantic precision (BERT-F1: 0.8155), while DeepSeek 7B provided balanced performance. Overall, RL-TweetGen presents a scalable and adaptive solution for marketers, content strategists, and Web3 platforms seeking to automate and optimize social media engagement. The framework advances the role of generative AI in digital commerce by aligning content generation with platform dynamics, user preferences, and marketing goals.

Open access
Topic Modeling
Sentiment Analysis and Opinion Mining
Advanced Text Analysis Techniques
Original source
Aug 8, 2024·Indonesian Journal of Computer Science
3 cites
Sentiment Analysis of User Reviews on Cryptocurrency Application: Evaluating the Impact of Dataset Split Scenarios Using Multinomial Naive Bayes

Chrisdion Andrew Ramaputra, Mohammad Hamim Zajuli Al Faroby, Berlian Rahmy Lidiawaty

The surge in cryptocurrency investors in Indonesia, reaching 18.83 million by January 2024, signifies an expanding interest in this market. This research conducts a sentiment analysis of user reviews on Indodax and Tokocrypto, the premier cryptocurrency trading platforms in Indonesia. Utilizing the Multinomial Naive Bayes method, the study examines the influence of various dataset split scenarios and random states on the model's performance. The findings reveal substantial variability in the model's accuracy based on different random states and test sizes. Notably, the Positive sentiment label consistently shows high-performance metrics, while the Neutral label underperforms. These insights are invaluable for developers aiming to improve user experience and for investors seeking to make informed decisions. This research underscores the significance of sentiment analysis in understanding user interactions and enhancing the credibility of cryptocurrency investment platforms.

Open access
Digital Marketing and Social Media
Advanced Text Analysis Techniques
Original source
Apr 13, 2024·Scientific Reports
8 cites
Psycholinguistic and emotion analysis of cryptocurrency discourse on X platform

Moein Shahiki Tash, Olga Kolesnikova, Zahra Ahani, Grigori Sidorov

This paper provides an extensive examination of a sizable dataset of English tweets focusing on nine widely recognized cryptocurrencies, specifically Cardano, Binance, Bitcoin, Dogecoin, Ethereum, Fantom, Matic, Shiba, and Ripple. Our goal was to conduct a psycholinguistic and emotional analysis of social media content associated with these cryptocurrencies. Such analysis can enable researchers and experts dealing with cryptocurrencies to make more informed decisions. Our work involved comparing linguistic characteristics across the diverse digital coins, shedding light on the distinctive linguistic patterns emerging in each coin's community. To achieve this, we utilized advanced text analysis techniques. Additionally, this work unveiled an understanding of the interplay between these digital assets. By examining which coin pairs are mentioned together most frequently in the dataset, we established co-mentions among different cryptocurrencies. To ensure the reliability of our findings, we initially gathered a total of 832,559 tweets from X. These tweets underwent a rigorous preprocessing stage, resulting in a refined dataset of 115,899 tweets that were used for our analysis. Overall, our research offers valuable perception into the linguistic nuances of various digital coins' online communities and provides a deeper understanding of their interactions in the cryptocurrency space.

Open access
Misinformation and Its Impacts
Blockchain Technology Applications and Security
Mental Health via Writing
Original source
Mar 26, 2024·arXiv (Cornell University)
2 cites
Chain-of-Action: Faithful and Multimodal Question Answering through Large Language Models

Zhenyu Pan, Haozheng Luo, Manling Li, Han Liu

We present a Chain-of-Action (CoA) framework for multimodal and retrieval-augmented Question-Answering (QA). Compared to the literature, CoA overcomes two major challenges of current QA applications: (i) unfaithful hallucination that is inconsistent with real-time or domain facts and (ii) weak reasoning performance over compositional information. Our key contribution is a novel reasoning-retrieval mechanism that decomposes a complex question into a reasoning chain via systematic prompting and pre-designed actions. Methodologically, we propose three types of domain-adaptable `Plug-and-Play' actions for retrieving real-time information from heterogeneous sources. We also propose a multi-reference faith score (MRFS) to verify and resolve conflicts in the answers. Empirically, we exploit both public benchmarks and a Web3 case study to demonstrate the capability of CoA over other methods.

Open access
2 source records
cs.CL
Topic Modeling
Natural Language Processing Techniques
Original source
Feb 4, 2024·arXiv (Cornell University)
1 cites
Safeguarding the Truth of High-Value Price Oracle Task: A Dynamically Adjusted Truth Discovery Method

Youquan Xian, Peng Liu, Dongcheng Li, Xueying Zeng

In recent years, the Decentralized Finance (DeFi) market has witnessed numerous attacks on the price oracle, leading to substantial economic losses. Despite the advent of truth discovery methods opening up new avenues for oracle development, it falls short in addressing high-value attacks on price oracle tasks. Consequently, this paper introduces a dynamically adjusted truth discovery method safeguarding the truth of high-value price oracle tasks. In the truth aggregation stage, we enhance future considerations to improve the precision of aggregated truth. During the credibility update phase, credibility is dynamically assessed based on the task's value and the Cumulative Potential Economic Contribution (CPEC) of information sources. Experimental results demonstrate a significant reduction in data deviation by 65.8\% and potential economic loss by 66.5\%, compared to the baseline scheme, in the presence of high-value attacks.

Open access
2 source records
cs.GT
cs.CE
cs.DC
Original source
Jan 15, 2024·arXiv
0 cites
Utilizing deep learning models for the identification of enhancers and super-enhancers based on genomic and epigenomic features

Zahra Ahani, Moein Shahiki Tash, Yoel Ledo Mezquita, Jason Angel

This paper provides an extensive examination of a sizable dataset of English tweets focusing on nine widely recognized cryptocurrencies, specifically Cardano, Binance, Bitcoin, Dogecoin, Ethereum, Fantom, Matic, Shiba, and Ripple. Our primary objective was to conduct a psycholinguistic and emotion analysis of social media content associated with these cryptocurrencies. To enable investigators to make more informed decisions. The study involved comparing linguistic characteristics across the diverse digital coins, shedding light on the distinctive linguistic patterns that emerge within each coin's community. To achieve this, we utilized advanced text analysis techniques. Additionally, our work unveiled an intriguing Understanding of the interplay between these digital assets within the cryptocurrency community. By examining which coin pairs are mentioned together most frequently in the dataset, we established correlations between different cryptocurrencies. To ensure the reliability of our findings, we initially gathered a total of 832,559 tweets from Twitter. These tweets underwent a rigorous preprocessing stage, resulting in a refined dataset of 115,899 tweets that were used for our analysis. Overall, our research offers valuable Perception into the linguistic nuances of various digital coins' online communities and provides a deeper understanding of their interactions in the cryptocurrency space.

Open access
cs.CL
cs.AI
Original source
Jan 1, 2024·Knowledge-Based Systems
4 cites
Enhancing large language models for bitcoin time series forecasting

Owen Chaffard, Pablo Mollá, Marc Cavazza, Helmut Prendinger

In the recent advancements in application of deep learning to time series forecasting, focus has shifted from training transformers end-to-end to efficiently leveraging the predictive capabilities of Large Language Models (LLMs). Models that encode the time series data to interact with a frozen LLM backbone have been shown to outperform transformers on all benchmark datasets. However, their efficiency on complex datasets, which do not show clear seasonality or trend, remains an open question. In this work, we seek to evaluate the performance of reprogrammed LLMs on the Bitcoin price chart, a financial time series known for its complexity and high volatility. We propose effective methods to improve the performance of Time-LLM, a State-of-the-art (SOTA) method, on such a time series. First, we propose structural improvements to Time-LLM. Second, we suggest an efficient way to handle the non-stationarity of the dataset. Finally, we propose an efficient method for passing additional financial information to the LLM. Our results demonstrate a 50% improvement on the average percentage loss and a 5% increase on accuracy of our adapted Time-LLM architecture on Bitcoin data when compared to SOTA models, including the original Time-LLM model. This highlights the impact on forecast accuracy of domain-specific decision making in data processing and feature selection.

Open access
2 source records
Stock Market Forecasting Methods
Time Series Analysis and Forecasting
Advanced Text Analysis Techniques
Original source
Mar 16, 2023·International Journal of Computing and Digital Systems
2 cites
Towards an Extractive Summarization for utilizing Learning Content using Deep Learning algorithm: Proposed Framework and Implementation

Yusra Mohammed AlRoshdi, Mohammed Al-Badawi, Abdullah Al-Hamdani, Mohamed Sarrab

Decentralized Web (Web3) and Finance (DeFi) have become the main discussion topic in research and industry fields.Cryptocurrencies, as an essential part of DeFi, enjoyed the interest of many stakeholders such as companies, professionals, researchers, and even common citizens eager to benefit from the proposed ecosystems.Although previous research studies focused on establishing price prediction systems using Sentiment Analysis (SA) techniques, the main focus of these studies was the performance of the predictive system rather than the accuracy and efficiency of the used models.In our work, we address two research questions; the predictability of cryptocurrency price based on past social and technical information, and the effect of social features on cryptocurrency price fluctuations using an SA and a Time Series approach.A combination of selected social and technical features was processed and reframed as a prediction problem, then studied to assess the ability of our model to predict the desired price.We noted that there is both an explicit correlation for some considered features and implicit for others, also social features including overall positive and neutral sentiment, and community engagement improved the performance of our model.

Open access
Topic Modeling
Advanced Text Analysis Techniques
Web Data Mining and Analysis
Original source
Mar 1, 2023·IT Professional
4 cites
Topic Modeling Based on Two-Step Flow Theory: Application to Tweets about Bitcoin

Aos Mulahuwaish, Matthew Loucks, Basheer Qolomany, Ala Al‐Fuqaha

Digital cryptocurrencies such as Bitcoin have exploded in recent years in both popularity and value. By their novelty, cryptocurrencies tend to be both volatile and highly speculative. The capricious nature of these coins is helped facilitated by social media networks such as Twitter. However, not everyone's opinion matters equally, with most posts garnering little to no attention. Additionally, the majority of tweets are retweeted from popular posts. We must determine whose opinion matters and the difference between influential and non-influential users. This study separates these two groups and analyzes the differences between them. It uses Hypertext-induced Topic Selection (HITS) algorithm, which segregates the dataset based on influence. Topic modeling is then employed to uncover differences in each group's speech types and what group may best represent the entire community. We found differences in language and interest between these two groups regarding Bitcoin and that the opinion leaders of Twitter are not aligned with the majority of users. There were 2559 opinion leaders (0.72% of users) who accounted for 80% of the authority and the majority (99.28%) users for the remaining 20% out of a total of 355,139 users.

Open access
2 source records
cs.SI
cs.AI
cs.CY
Original source
Jan 6, 2023·Informatics
11 cites
Cryptoblend: An AI-Powered Tool for Aggregation and Summarization of Cryptocurrency News

Andrea Pozzi, Enrico Barbierato, Daniele Toti

In the last decade, the techniques of news aggregation and summarization have been increasingly gaining relevance for providing users on the web with condensed and unbiased information. Indeed, the recent development of successful machine learning algorithms, such as those based on the transformers architecture, have made it possible to create effective tools for capturing and elaborating news from the Internet. In this regard, this work proposes, for the first time in the literature to the best of the authors’ knowledge, a methodology for the application of such techniques in news related to cryptocurrencies and the blockchain, whose quick reading can be deemed as extremely useful to operators in the financial sector. Specifically, cutting-edge solutions in the field of natural language processing were employed to cluster news by topic and summarize the corresponding articles published by different newspapers. The results achieved on 22,282 news articles show the effectiveness of the proposed methodology in most of the cases, with 86.8% of the examined summaries being considered as coherent and 95.7% of the corresponding articles correctly aggregated. This methodology was implemented in a freely accessible web application.

Open access
Stock Market Forecasting Methods
Blockchain Technology Applications and Security
Advanced Text Analysis Techniques
Original source
Jan 1, 2023·IEEE Access
4 cites
Annotators’ Selection Impact on the Creation of a Sentiment Corpus for the Cryptocurrency Financial Domain

Manoel Fernando Alonso Gadi, Miguel‐Ángel Sicilia

Well labeled natural language corpus data is essential for most natural language processing techniques, especially in specialized fields. However, cohort biases remain a significant challenge in machine learning. The narrow origin of data sampling or human annotators in cohorts is a prevalent issue for machine learning researchers due to its potential to induce bias in the final product. During the development of the CryptoLin corpus for another research project, the authors became concerned about the potential influence of cohort bias on the selection of annotators. Therefore, this paper addresses the question of whether cohort diversity improves the labeling result through the implementation of a repeated annotator process, involving two annotator cohorts and a statistically robust comparison methodology. The utilization of statistical tests, such as the Chi-Square Independence test for absolute frequency tables, and the construction of confidence intervals for Kappa point estimates, facilitates a rigorous analysis of the differences between Kappa estimates. Furthermore, the application of a two-proportion z-test to compare the accuracy scores of UTAD and IE annotators for various pre-trained models, including Vader Sentiment Analysis, TextBlob Sentiment Analysis, Flair NLP library, and FinBERT Financial Sentiment Analysis with BERT, contributes to the advancement of knowledge in this field. The paper utilizes Cryptocurrency Linguo (CryptoLin), a corpus containing 2683 cryptocurrency-related news articles spanning more than three years,and compares two different selection criteria for the annotators. CryptoLin was annotated twice with discrete values representing negative, neutral, and positive news respectively. The first annotation was done by twenty-seven annotators from the same cohort. Each news title was randomly assigned and blindly annotated by three human annotators. The second annotation was carried out by eighty-three annotators from three cohorts. Each news title was randomly assigned and blindly annotated by three human annotators, one in each different cohort. In both annotations, a consensus mechanism using simple voting was applied. The first annotation used the same cohort with students from the same nationality and background. The second used three cohorts with students from a very diverse set of nationalities and educational backgrounds. The results demonstrate that manual labeling done by both groups was acceptable according to inter-rater reliability coefficients Fleiss’s Kappa, Krippendorff’s Alpha, and Gwet’s AC1. Preliminary analysis utilizing Vader, Textblob, Flair, and FinBERT confirmed the utility of the data set labeling for further refinement of sentiment analysis algorithms. Our results also highlight that the more diverse annotator pool performed better in all measured aspects.

Open access
Topic Modeling
Advanced Text Analysis Techniques
Stock Market Forecasting Methods
Original source
Mar 17, 2022·arXiv (Cornell University)
1 cites
Short Text Topic Modeling: Application to tweets about Bitcoin

Hugo Schnoering

Understanding the semantic of a collection of texts is a challenging task. Topic models are probabilistic models that aims at extracting "topics" from a corpus of documents. This task is particularly difficult when the corpus is composed of short texts, such as posts on social networks. Following several previous research papers, we explore in this paper a set of collected tweets about bitcoin. In this work, we train three topic models and evaluate their output with several scores. We also propose a concrete application of the extracted topics.

Open access
2 source records
cs.IR
cs.LG
Advanced Text Analysis Techniques
Original source
Jun 2, 2021·arXiv (Cornell University)
1 cites
A weighted unified informetrics based on Scopus and WoS

Parul Khurana, Geetha Ganesan, Gulshan Kumar, Kiran Sharma

Numerous indexing databases keep track of the number of publications, citations, etc. in order to maintain the progress of science and individual. However, the choice of journals and articles varies among these indexing databases, hence the number of citations and h-index varies. There is no common platform exists that can provide a single count for the number of publications, citations, h-index, etc. To overcome this limitation, we have proposed a weighted unified informetrics, named "conflate". The proposed system takes into account the input from multiple indexing databases and generates a single output. Here, we have used the data from Scopus and WoS to generate a conflate dataset. Further, a comparative analysis of conflate has been performed with Scopus and WoS at three levels: author, organization, and journal. Finally, a mapping is proposed between research publications and distributed ledger technology in order to provide a transparent and distributed view to its stakeholders.

Open access
2 source records
cs.DL
cs.IR
Advanced Text Analysis Techniques
Original source
Jan 1, 2020·IEEE Access
6 cites
The Irruption of Cryptocurrencies Into Twitter Cashtags: A Classifying Solution

Ana Fernández Vilas, Rebeca P. Dı́az Redondo, Anton Lorenzo Garcia

There is a consensus about the good sensing characteristics of Twitter to mine and uncover knowledge in financial markets, being considered a relevant feeder for taking decisions about buying or holding stock shares and even for detecting stock manipulation. Although Twitter hashtags allow to aggregate topic-related content, a specific mechanism for financial information also exists: Cashtag (consisting of the company ticker preceded by $) is a supporting mechanism to track financial tweets referring to a company listed in a stock market. However, according to our experiments and due to the lack of conventions in cashtags usage, the irruption of cryptocurrencies has resulted in a significant degradation on the cashtag-based aggregation of posts. Unfortunately, Twitter' users may use homonym tickers to refer to cryptocurrencies and to companies in stock markets, which means that filtering by cashtag may result on both posts referring to stock companies and cryptocurrencies. This research proposes automated classifiers to distinguish conflicting cashtags and, so, their container tweets by analyzing the distinctive features of tweets referring to stock companies and cryptocurrencies. As experiment, this paper analyses the interference between cryptocurrencies and company tickers in the London Stock Exchange (LSE), specifically, companies in the main and alternative market indices FTSE-100 and AIM-100. Heuristic-based as well as supervised classifiers are proposed and their advantages and drawbacks, including their ability to self-adapt to Twitter usage changes, are discussed. The experiment confirms a significant distortion in collected data when colliding or homonym cashtags exist, i.e., the same $ acronym to refer to company tickers and cryptocurrencies. According to our results, the distinctive features of posts including cryptocurrencies or company tickers support accurate classification of colliding tweets (homonym cashtags) and Independent Models, as the most detached classifiers from training data, have the potential to be trans-applicability (in different stock markets) while retaining performance.

Open access
2 source records
Stock Market Forecasting Methods
Advanced Text Analysis Techniques
Complex Systems and Time Series Analysis
Original source
Dec 1, 2019·Acta Marisiensis Seria Technologica
10 cites
Cryptocurrency – Sentiment Analysis in Social Media

Tudor-Mircea Dulău, Mircea Dulău

Abstract The paper proposes the exploration, identification and development of a Java solution for extracting the sentiment related to the cryptocurrencies phenomenon, from the content of the posts of certain popular social networks. Detecting the positive, neutral or negative character of the sentiment is adopted as a relevant method of establishing the nature of the human perception on the topical issue defined by cryptocurrencies.

Open access
Sentiment Analysis and Opinion Mining
Advanced Text Analysis Techniques
Spam and Phishing Detection
Original source
Apr 16, 2019·arXiv (Cornell University)
10 cites
DSTP-RNN: a dual-stage two-phase attention-based recurrent neural networks for long-term and multivariate time series prediction

Yeqi Liu, Chuanyang Gong, Ling Yang, Yingyi Chen

Long-term prediction of multivariate time series is still an important but challenging problem. The key to solve this problem is to capture the spatial correlations at the same time, the spatio-temporal relationships at different times and the long-term dependence of the temporal relationships between different series. Attention-based recurrent neural networks (RNN) can effectively represent the dynamic spatio-temporal relationships between exogenous series and target series, but it only performs well in one-step time prediction and short-term time prediction. In this paper, inspired by human attention mechanism including the dual-stage two-phase (DSTP) model and the influence mechanism of target information and non-target information, we propose DSTP-based RNN (DSTP-RNN) and DSTP-RNN-2 respectively for long-term time series prediction. Specifically, we first propose the DSTP-based structure to enhance the spatial correlations between exogenous series. The first phase produces violent but decentralized response weight, while the second phase leads to stationary and concentrated response weight. Secondly, we employ multiple attentions on target series to boost the long-term dependence. Finally, we study the performance of deep spatial attention mechanism and provide experiment and interpretation. Our methods outperform nine baseline methods on four datasets in the fields of energy, finance, environment and medicine, respectively.

Open access
Time Series Analysis and Forecasting
Stock Market Forecasting Methods
Advanced Text Analysis Techniques
Original source
Jan 1, 2016·SSRN Electronic Journal
10 cites
Rapid Prototyping of a Text Mining Application for Cryptocurrency Market Intelligence

Marek Laskowski, Henry Kim

Blockchain represents a technology for establishing a shared, immutable version of the truth between a network of participants that do not trust one another, and therefore has the potential to disrupt any financial or other industries that rely on third-parties to establish trust. Recent trends in computing including: prevalence of Free and Open Source Software (FOSS); easy access to High Performance Computing (HPC i.e. 'The Cloud'); and increasingly advanced analytics capabilities such as Natural Language Processing (NLP) and Machine Learning (ML) allow for rapidly prototyping applications for analysis of trends in the emergence of Blockchain technology. A scaleable proof-of-concept pipeline that lays the groundwork for analysis of multiple streams of semi-structured data posted on social media is demonstrated. Preliminary analysis and performance metrics are presented and discussed. Future work is described that will scale the system to cloud-based, real-time, analysis of multiple data streams, with Information Extraction (IE) (ex. sentiment analysis) and Machine Learning capability.

Open access
3 source records
cs.CY
cs.ET
Blockchain Technology Applications and Security
Original source
Jul 14, 2004·Hosei University Repository (Hosei University)
0 cites
Webからの時制クラスタの解釈(Web3)(夏のデータベースワークショップDBWS2004)

正輝 森, 孝夫 三浦, 勇 塩谷

本稿では、Webページ集合からの事象抽出及び自動解釈を行うための新しいWebマイニングの手法を提案する。Webページを調査し有効時間を抽出しK-meansクラスタリングにより事象の抽出を行う。TDTのアプローチをWeb環境に適応し、提案する手法が時制Webページに対して有効であることをいくつかの実験により示す。

Open access
Web Data Mining and Analysis
Information Retrieval and Search Behavior
Advanced Text Analysis Techniques
Original source