In the rapidly evolving landscape of digital marketing and electronic commerce, short-form content—particularly on platforms like Twitter (now X)—has become pivotal for real-time branding, community engagement, and product promotion. The rise of Non-Fungible Tokens (NFTs) and Web3 ecosystems further underscores the need for domain-specific, engagement-oriented social media content. However, automating the generation of such content while balancing linguistic quality, semantic relevance, and audience engagement remains a substantial challenge. To address this, we propose RL-TweetGen, a socio-technical framework that integrates instruction-tuned large language models (LLMs) with reinforcement learning (RL) to generate concise, impactful, and engagement-optimized tweets. The framework incorporates a structured pipeline comprising domain-specific data curation, semantic classification, and intent-aware prompt engineering, and leverages Parameter-Efficient Fine-Tuning (PEFT) with LoRA for scalable model adaptation. We fine-tuned and evaluated three LLMs—LLaMA-3.1-8B, Mistral-7B Instruct, and DeepSeek 7B Chat—guided by a hybrid reward function that blends XGBoost-predicted engagement scores with expert-in-the-loop feedback. To enhance lexical diversity and contextual alignment, we implemented advanced decoding strategies, including Tailored Beam Search, Enhanced Top-p Sampling, and Contextual Temperature Scaling. A case study focused on NFT-related tweet generation demonstrated the practical effectiveness of RL-TweetGen. Experimental results showed that Mistral-7B achieved the highest lexical fluency (BLEU: 0.2285), LLaMA-3.1 exhibited superior semantic precision (BERT-F1: 0.8155), while DeepSeek 7B provided balanced performance. Overall, RL-TweetGen presents a scalable and adaptive solution for marketers, content strategists, and Web3 platforms seeking to automate and optimize social media engagement. The framework advances the role of generative AI in digital commerce by aligning content generation with platform dynamics, user preferences, and marketing goals.
Abstract Self-organized system (SOS) offers a compelling paradigm for enabling autonomous coordination in a multiagent system (MAS). By leveraging decentralized decision-making, these systems can dynamically adapt to evolving environments, making them highly suitable for complex engineering applications. While multiagent reinforcement learning (MARL) has advanced agents to perform collaboratively, the mechanisms driving such self-organization still remain largely unexplored. Thus, this paper aims to deepen the understanding how agents self-organize themselves into an intelligent team from the role specialization perspective. In the context of a precision assembly task with collision avoidance, agent teams are trained to collaborate without predefined role assignments. We apply physical motion analysis to quantify individual contributions to system dynamics as a team and utilize K-means clustering to further provide a data-driven view of emergent behaviors. By analyzing agent role’s scores based on their contributions to the system, the mechanisms of role specialization have been uncovered. Results show that the effective role differentiation can be naturally emergent from the MARL training process. Furthermore, the training leads the role assignments into a hierarchical structure, where some agents take primary roles while others provide dynamic support. The findings offer practical guidelines for designing an adaptive, efficient, and robust multiagent systems for applications such as autonomous robotics and advanced manufacturing.
This exploratory study introduces a sentiment-based framework for the social dynamics of hype around emerging technologies to support strategic investment decisions and contribute to innovation and strategic management research. Drawing on theories of herding behaviour and information cascades, we analysed social media sentiment of the Reddit discourse and investigated bubbles in representative financial assets across four emerging technology cases ‘Metaverse’, ‘Decentralized Finance (DeFi)’, ‘Non-fungible Token (NFT)’, and ‘Hydrogen Economy’ over one-year periods. Additionally, we compared results against measures of search volumes, news coverage, patent filings, and academic publications. We find two essential characteristics of the social dynamics of hype: intensifying positive sentiment and increasing conformity of sentiment towards a positive majority. Since hype can distort decision-making and hinder objective innovation assessments, this study offers practitioners a sentiment-based approach to navigate speculative hype and support decision-making.
Blockchain technology has rapidly expanded in recent years, impacting ICT and cryptocurrencies. It offers complete security for transactions, recording each transaction in a block that serves as a record sheet. Blockchain has changed how we interact with the internet, reducing transaction costs, eliminating middlemen, and increasing immutability and transparency. The community of blockchain technologies is multifaceted, with distinct frameworks meeting different users' needs and use cases. This study presents an in-depth comparison of blockchain frameworks, including Ethereum, Quorum, Corda, Ripple and Hyperledger Fabric to help developers and organizations in selecting the most preferable platform for their unique use cases. The study evaluates the blockchain frameworks on the basis of performance analysis metrics such as consensus techniques, throughput, latency, and transaction per second with Hyperledger Fabric being the most suitable choice for achieving throughput and scalability.
Natural Language Processing (NLP) has significantly advanced the ability to analyse and interpret textual data, playing a crucial role in understanding user sentiments. This study applies cutting-edge NLP techniques to perform sentiment analysis on reviews from India’s top cryptocurrency apps and Twitter data. Given the growing interest in cryptocurrencies, understanding user sentiment is vital for market insights and product improvement. We collected a diverse dataset from the Google Play store and Twitter, encompassing 7,197 reviews and numerous tweets. Utilizing the BERT (Bidirectional Encoder Representations from Transformers) model, known for its deep learning capabilities, we processed and analysed the data. The dataset underwent thorough pre-processing, including tokenization and the removal of irrelevant elements. Our analysis compared the BERT model’s performance with traditional classifiers such as Naive Bayes and Support Vector Machines (SVM). Findings show that BERT significantly outperformed other models, achieving superior precision, recall, accuracy and F1-scores. These results underscore the effectiveness of advanced NLP models in sentiment analysis, particularly for understanding public sentiment towards cryptocurrency apps in India. Future research will explore sentiment analysis on a broader range of platforms and fine-tuning BERT model parameters to further enhance performance and accuracy.
Chrisdion Andrew Ramaputra, Mohammad Hamim Zajuli Al Faroby, Berlian Rahmy Lidiawaty
The surge in cryptocurrency investors in Indonesia, reaching 18.83 million by January 2024, signifies an expanding interest in this market. This research conducts a sentiment analysis of user reviews on Indodax and Tokocrypto, the premier cryptocurrency trading platforms in Indonesia. Utilizing the Multinomial Naive Bayes method, the study examines the influence of various dataset split scenarios and random states on the model's performance. The findings reveal substantial variability in the model's accuracy based on different random states and test sizes. Notably, the Positive sentiment label consistently shows high-performance metrics, while the Neutral label underperforms. These insights are invaluable for developers aiming to improve user experience and for investors seeking to make informed decisions. This research underscores the significance of sentiment analysis in understanding user interactions and enhancing the credibility of cryptocurrency investment platforms.
This research introduces a cutting-edge Web3 literary analysis platform, harnessing the power of blockchain and deep learning technologies. By employing the immutable and transparent nature of blockchain, the platform ensures robust copyright protection while offering readers enhanced interactive features. It applies deep learning techniques for comprehensive analyses of sentiment, topic, and stylistic elements, which are instrumental in predicting potential Nobel Prize laureates. This methodology not only enhances the accuracy of predictions but also sheds light on the evaluation criteria and historical trends associated with the Nobel Prize. Moreover, the platform adopts a directed graph model alongside the struc2vec algorithm to create text vectors for comparative studies, uncovering similarities between works that have won awards and those that have been nominated. Utilizing the LESS model for detailed content examination, the platform delves into sequence relationships within semantic networks, thus improving interpretability and visualization. The integration of blockchain technology guarantees access to unbiased datasets, enabling more precise literary analyses and predictions. This innovative approach has been validated using works that have either won or been nominated for the Nobel Prize, proving its efficacy in identifying the textual characteristics favored by the Nobel Prize committee.
Moein Shahiki Tash, Olga Kolesnikova, Zahra Ahani, Grigori Sidorov
This paper provides an extensive examination of a sizable dataset of English tweets focusing on nine widely recognized cryptocurrencies, specifically Cardano, Binance, Bitcoin, Dogecoin, Ethereum, Fantom, Matic, Shiba, and Ripple. Our goal was to conduct a psycholinguistic and emotional analysis of social media content associated with these cryptocurrencies. Such analysis can enable researchers and experts dealing with cryptocurrencies to make more informed decisions. Our work involved comparing linguistic characteristics across the diverse digital coins, shedding light on the distinctive linguistic patterns emerging in each coin's community. To achieve this, we utilized advanced text analysis techniques. Additionally, this work unveiled an understanding of the interplay between these digital assets. By examining which coin pairs are mentioned together most frequently in the dataset, we established co-mentions among different cryptocurrencies. To ensure the reliability of our findings, we initially gathered a total of 832,559 tweets from X. These tweets underwent a rigorous preprocessing stage, resulting in a refined dataset of 115,899 tweets that were used for our analysis. Overall, our research offers valuable perception into the linguistic nuances of various digital coins' online communities and provides a deeper understanding of their interactions in the cryptocurrency space.
We present a Chain-of-Action (CoA) framework for multimodal and retrieval-augmented Question-Answering (QA). Compared to the literature, CoA overcomes two major challenges of current QA applications: (i) unfaithful hallucination that is inconsistent with real-time or domain facts and (ii) weak reasoning performance over compositional information. Our key contribution is a novel reasoning-retrieval mechanism that decomposes a complex question into a reasoning chain via systematic prompting and pre-designed actions. Methodologically, we propose three types of domain-adaptable `Plug-and-Play' actions for retrieving real-time information from heterogeneous sources. We also propose a multi-reference faith score (MRFS) to verify and resolve conflicts in the answers. Empirically, we exploit both public benchmarks and a Web3 case study to demonstrate the capability of CoA over other methods.
In recent years, the Decentralized Finance (DeFi) market has witnessed numerous attacks on the price oracle, leading to substantial economic losses. Despite the advent of truth discovery methods opening up new avenues for oracle development, it falls short in addressing high-value attacks on price oracle tasks. Consequently, this paper introduces a dynamically adjusted truth discovery method safeguarding the truth of high-value price oracle tasks. In the truth aggregation stage, we enhance future considerations to improve the precision of aggregated truth. During the credibility update phase, credibility is dynamically assessed based on the task's value and the Cumulative Potential Economic Contribution (CPEC) of information sources. Experimental results demonstrate a significant reduction in data deviation by 65.8\% and potential economic loss by 66.5\%, compared to the baseline scheme, in the presence of high-value attacks.
Zahra Ahani, Moein Shahiki Tash, Yoel Ledo Mezquita, Jason Angel
This paper provides an extensive examination of a sizable dataset of English tweets focusing on nine widely recognized cryptocurrencies, specifically Cardano, Binance, Bitcoin, Dogecoin, Ethereum, Fantom, Matic, Shiba, and Ripple. Our primary objective was to conduct a psycholinguistic and emotion analysis of social media content associated with these cryptocurrencies. To enable investigators to make more informed decisions. The study involved comparing linguistic characteristics across the diverse digital coins, shedding light on the distinctive linguistic patterns that emerge within each coin's community. To achieve this, we utilized advanced text analysis techniques. Additionally, our work unveiled an intriguing Understanding of the interplay between these digital assets within the cryptocurrency community. By examining which coin pairs are mentioned together most frequently in the dataset, we established correlations between different cryptocurrencies. To ensure the reliability of our findings, we initially gathered a total of 832,559 tweets from Twitter. These tweets underwent a rigorous preprocessing stage, resulting in a refined dataset of 115,899 tweets that were used for our analysis. Overall, our research offers valuable Perception into the linguistic nuances of various digital coins' online communities and provides a deeper understanding of their interactions in the cryptocurrency space.
Owen Chaffard, Pablo Mollá, Marc Cavazza, Helmut Prendinger
In the recent advancements in application of deep learning to time series forecasting, focus has shifted from training transformers end-to-end to efficiently leveraging the predictive capabilities of Large Language Models (LLMs). Models that encode the time series data to interact with a frozen LLM backbone have been shown to outperform transformers on all benchmark datasets. However, their efficiency on complex datasets, which do not show clear seasonality or trend, remains an open question. In this work, we seek to evaluate the performance of reprogrammed LLMs on the Bitcoin price chart, a financial time series known for its complexity and high volatility. We propose effective methods to improve the performance of Time-LLM, a State-of-the-art (SOTA) method, on such a time series. First, we propose structural improvements to Time-LLM. Second, we suggest an efficient way to handle the non-stationarity of the dataset. Finally, we propose an efficient method for passing additional financial information to the LLM. Our results demonstrate a 50% improvement on the average percentage loss and a 5% increase on accuracy of our adapted Time-LLM architecture on Bitcoin data when compared to SOTA models, including the original Time-LLM model. This highlights the impact on forecast accuracy of domain-specific decision making in data processing and feature selection.
Mehedi Hasan, Md. Tahmid Rahman, Kazi Ahnaf Alavee, Abu Hasnayen Zillanee · 6 authors
In the fast-paced realm of global financial markets, characterized by rapid trading of both stocks and cryptocurren-cies, it has become essential to grasp the influence of sentiment on market dynamics. With more than 630,000 publicly traded companies worldwide and major stock exchanges like the NYSE handling a substantial portion of global equity transactions, the inherent volatility of the stock market is well-established. Over the past decade, various factors have contributed to the consistent fluctuations in stock prices. One key factor is the influence of investor reviews sourced from diverse news outlets and social media platforms such as Twitter. Understanding how these reviews can be collected and effectively summarized is crucial. This paper centers on the intricate field of market sentiment analysis and its profound impact on user sentiment, subsequently affecting price fluctuations in both stocks and cryptocurrencies. In this study, we present a comprehensive exploration of the development and evaluation of an automated sentiment analysis system tailored for summarizing web-based news related to stocks and cryptocurrencies.We have implemented BERT (Bidirectional Encoder Representations from Transformers) in combination with NLTK for text summarization, a highly accurate model with a performance level of 95.84%, as part of our proposed approach.
Yusra Mohammed AlRoshdi, Mohammed Al-Badawi, Abdullah Al-Hamdani, Mohamed Sarrab
Decentralized Web (Web3) and Finance (DeFi) have become the main discussion topic in research and industry fields.Cryptocurrencies, as an essential part of DeFi, enjoyed the interest of many stakeholders such as companies, professionals, researchers, and even common citizens eager to benefit from the proposed ecosystems.Although previous research studies focused on establishing price prediction systems using Sentiment Analysis (SA) techniques, the main focus of these studies was the performance of the predictive system rather than the accuracy and efficiency of the used models.In our work, we address two research questions; the predictability of cryptocurrency price based on past social and technical information, and the effect of social features on cryptocurrency price fluctuations using an SA and a Time Series approach.A combination of selected social and technical features was processed and reframed as a prediction problem, then studied to assess the ability of our model to predict the desired price.We noted that there is both an explicit correlation for some considered features and implicit for others, also social features including overall positive and neutral sentiment, and community engagement improved the performance of our model.
This paper predicts sentiments of crypto currency news articles using BERT (Bidirectional Encoder Representation) model, as there is a lack of research in crypto currency price prediction using natural language processing. The text data obtained is unlabeled and it is labelled using a parsimonious rule-based model and then BERT is used to dassify news sentiment as “Positive”, “Negative” or “Neutral” which may be helpful in reading cryptocurrency market movement.
Aos Mulahuwaish, Matthew Loucks, Basheer Qolomany, Ala Al‐Fuqaha
Digital cryptocurrencies such as Bitcoin have exploded in recent years in both popularity and value. By their novelty, cryptocurrencies tend to be both volatile and highly speculative. The capricious nature of these coins is helped facilitated by social media networks such as Twitter. However, not everyone's opinion matters equally, with most posts garnering little to no attention. Additionally, the majority of tweets are retweeted from popular posts. We must determine whose opinion matters and the difference between influential and non-influential users. This study separates these two groups and analyzes the differences between them. It uses Hypertext-induced Topic Selection (HITS) algorithm, which segregates the dataset based on influence. Topic modeling is then employed to uncover differences in each group's speech types and what group may best represent the entire community. We found differences in language and interest between these two groups regarding Bitcoin and that the opinion leaders of Twitter are not aligned with the majority of users. There were 2559 opinion leaders (0.72% of users) who accounted for 80% of the authority and the majority (99.28%) users for the remaining 20% out of a total of 355,139 users.
In the last decade, the techniques of news aggregation and summarization have been increasingly gaining relevance for providing users on the web with condensed and unbiased information. Indeed, the recent development of successful machine learning algorithms, such as those based on the transformers architecture, have made it possible to create effective tools for capturing and elaborating news from the Internet. In this regard, this work proposes, for the first time in the literature to the best of the authors’ knowledge, a methodology for the application of such techniques in news related to cryptocurrencies and the blockchain, whose quick reading can be deemed as extremely useful to operators in the financial sector. Specifically, cutting-edge solutions in the field of natural language processing were employed to cluster news by topic and summarize the corresponding articles published by different newspapers. The results achieved on 22,282 news articles show the effectiveness of the proposed methodology in most of the cases, with 86.8% of the examined summaries being considered as coherent and 95.7% of the corresponding articles correctly aggregated. This methodology was implemented in a freely accessible web application.
Well labeled natural language corpus data is essential for most natural language processing techniques, especially in specialized fields. However, cohort biases remain a significant challenge in machine learning. The narrow origin of data sampling or human annotators in cohorts is a prevalent issue for machine learning researchers due to its potential to induce bias in the final product. During the development of the CryptoLin corpus for another research project, the authors became concerned about the potential influence of cohort bias on the selection of annotators. Therefore, this paper addresses the question of whether cohort diversity improves the labeling result through the implementation of a repeated annotator process, involving two annotator cohorts and a statistically robust comparison methodology. The utilization of statistical tests, such as the Chi-Square Independence test for absolute frequency tables, and the construction of confidence intervals for Kappa point estimates, facilitates a rigorous analysis of the differences between Kappa estimates. Furthermore, the application of a two-proportion z-test to compare the accuracy scores of UTAD and IE annotators for various pre-trained models, including Vader Sentiment Analysis, TextBlob Sentiment Analysis, Flair NLP library, and FinBERT Financial Sentiment Analysis with BERT, contributes to the advancement of knowledge in this field. The paper utilizes Cryptocurrency Linguo (CryptoLin), a corpus containing 2683 cryptocurrency-related news articles spanning more than three years,and compares two different selection criteria for the annotators. CryptoLin was annotated twice with discrete values representing negative, neutral, and positive news respectively. The first annotation was done by twenty-seven annotators from the same cohort. Each news title was randomly assigned and blindly annotated by three human annotators. The second annotation was carried out by eighty-three annotators from three cohorts. Each news title was randomly assigned and blindly annotated by three human annotators, one in each different cohort. In both annotations, a consensus mechanism using simple voting was applied. The first annotation used the same cohort with students from the same nationality and background. The second used three cohorts with students from a very diverse set of nationalities and educational backgrounds. The results demonstrate that manual labeling done by both groups was acceptable according to inter-rater reliability coefficients Fleiss’s Kappa, Krippendorff’s Alpha, and Gwet’s AC1. Preliminary analysis utilizing Vader, Textblob, Flair, and FinBERT confirmed the utility of the data set labeling for further refinement of sentiment analysis algorithms. Our results also highlight that the more diverse annotator pool performed better in all measured aspects.
Understanding the semantic of a collection of texts is a challenging task. Topic models are probabilistic models that aims at extracting "topics" from a corpus of documents. This task is particularly difficult when the corpus is composed of short texts, such as posts on social networks. Following several previous research papers, we explore in this paper a set of collected tweets about bitcoin. In this work, we train three topic models and evaluate their output with several scores. We also propose a concrete application of the extracted topics.
Numerous indexing databases keep track of the number of publications, citations, etc. in order to maintain the progress of science and individual. However, the choice of journals and articles varies among these indexing databases, hence the number of citations and h-index varies. There is no common platform exists that can provide a single count for the number of publications, citations, h-index, etc. To overcome this limitation, we have proposed a weighted unified informetrics, named "conflate". The proposed system takes into account the input from multiple indexing databases and generates a single output. Here, we have used the data from Scopus and WoS to generate a conflate dataset. Further, a comparative analysis of conflate has been performed with Scopus and WoS at three levels: author, organization, and journal. Finally, a mapping is proposed between research publications and distributed ledger technology in order to provide a transparent and distributed view to its stakeholders.