Blockchain Papers

Follow blockchain research across journals, conferences, and preprint repositories.

98 papersLast indexed Aug 31, 2026
Search papers

Paper index

98 results · page 4 of 5

Clear filters
Jul 1, 2023·Proceedings/Proceedings of the ... International Conference on Software Engineering and Knowledge Engineering
2 cites
CCGRA: Smart Contract Code Comment Generation with Retrieval-enhanced Approach

Zhenhua Zhang, Shizhan Chen, Guodong Fan, Guang Yang · 5 authors

Smart contracts are self-executing programs on the blockchain that are critical to a range of industries, including finance, supply chain management, and healthcare.However, comprehending smart contracts can be challenging due to a lack of effective comments in most user-defined code.To address this challenge, we propose a novel retrieval-enhanced approach CC-GRA that leverages retrieval knowledge to generate high-quality comments for Solidity language code.Our approach carefully eliminates duplicated data and template data in the widely-used smart contract dataset to ensure a high-quality corpus.Extensive experiments and comprehensive analysis demonstrate the effectiveness applicability of our approach after being compared with eight state-of-the-art baselines.Finally, we conduct a human study and find the comment quality generated by our approach is better than baselines in terms of similarity, naturalness, and informativeness.

Open access
Blockchain Technology Applications and Security
FinTech, Crowdfunding, Digital Finance
Topic Modeling
Original source
Jun 9, 2023·Information
16 cites
Trend Analysis of Decentralized Autonomous Organization Using Big Data Analytics

Hyejin Park, Ivan Ureta, Boyoung Kim

Decentralized Autonomous Organizations (DAOs) have gained widespread attention in academia and industry as potential future models for decentralized governance and organization. In order to understand the trends and future potential of this rapidly growing technology, it is crucial to conduct research in the field. This research aims at a data-driven approach for the objective content analysis of big data related to DAOs, using text mining and Latent Dirichlet Allocation (LDA)-based topic modeling. The study analyzed tweets with the hashtag #DAO and all Reddit data with “DAO”. The results were from the identification of the top 100 frequently appearing keywords, as well as the top 20 keywords with high network centrality, and key topics related to finance, gaming, and fundraising, from both Twitter and Reddit. The analysis revealed twelve topics from Twitter and eight topics from Reddit, with the term “community” frequently appearing across many of these topics. The findings provide valuable insights into the current trend and future potential of DAOs, and should be used by researchers to guide further research in the field and by decision makers to explore innovative ways to govern the organizations.

Open access
Digital Marketing and Social Media
FinTech, Crowdfunding, Digital Finance
Blockchain Technology Applications and Security
Original source
Mar 16, 2023·International Journal of Computing and Digital Systems
2 cites
Towards an Extractive Summarization for utilizing Learning Content using Deep Learning algorithm: Proposed Framework and Implementation

Yusra Mohammed AlRoshdi, Mohammed Al-Badawi, Abdullah Al-Hamdani, Mohamed Sarrab

Decentralized Web (Web3) and Finance (DeFi) have become the main discussion topic in research and industry fields.Cryptocurrencies, as an essential part of DeFi, enjoyed the interest of many stakeholders such as companies, professionals, researchers, and even common citizens eager to benefit from the proposed ecosystems.Although previous research studies focused on establishing price prediction systems using Sentiment Analysis (SA) techniques, the main focus of these studies was the performance of the predictive system rather than the accuracy and efficiency of the used models.In our work, we address two research questions; the predictability of cryptocurrency price based on past social and technical information, and the effect of social features on cryptocurrency price fluctuations using an SA and a Time Series approach.A combination of selected social and technical features was processed and reframed as a prediction problem, then studied to assess the ability of our model to predict the desired price.We noted that there is both an explicit correlation for some considered features and implicit for others, also social features including overall positive and neutral sentiment, and community engagement improved the performance of our model.

Open access
Topic Modeling
Advanced Text Analysis Techniques
Web Data Mining and Analysis
Original source
Mar 1, 2023·IT Professional
4 cites
Topic Modeling Based on Two-Step Flow Theory: Application to Tweets about Bitcoin

Aos Mulahuwaish, Matthew Loucks, Basheer Qolomany, Ala Al‐Fuqaha

Digital cryptocurrencies such as Bitcoin have exploded in recent years in both popularity and value. By their novelty, cryptocurrencies tend to be both volatile and highly speculative. The capricious nature of these coins is helped facilitated by social media networks such as Twitter. However, not everyone's opinion matters equally, with most posts garnering little to no attention. Additionally, the majority of tweets are retweeted from popular posts. We must determine whose opinion matters and the difference between influential and non-influential users. This study separates these two groups and analyzes the differences between them. It uses Hypertext-induced Topic Selection (HITS) algorithm, which segregates the dataset based on influence. Topic modeling is then employed to uncover differences in each group's speech types and what group may best represent the entire community. We found differences in language and interest between these two groups regarding Bitcoin and that the opinion leaders of Twitter are not aligned with the majority of users. There were 2559 opinion leaders (0.72% of users) who accounted for 80% of the authority and the majority (99.28%) users for the remaining 20% out of a total of 355,139 users.

Open access
2 source records
cs.SI
cs.AI
cs.CY
Original source
Feb 3, 2023·arXiv
37 cites
Show me your NFT and I tell you how it will perform: Multimodal representation learning for NFT selling price prediction

Davide Costa, Lucio La Cava, Andrea Tagarelli

Non-Fungible Tokens (NFTs) represent deeds of ownership, based on blockchain technologies and smart contracts, of unique crypto assets on digital art forms (e.g., artworks or collectibles). In the spotlight after skyrocketing in 2021, NFTs have attracted the attention of crypto enthusiasts and investors intent on placing promising investments in this profitable market. However, the NFT financial performance prediction has not been widely explored to date. In this work, we address the above problem based on the hypothesis that NFT images and their textual descriptions are essential proxies to predict the NFT selling prices. To this purpose, we propose MERLIN, a novel multimodal deep learning framework designed to train Transformer-based language and visual models, along with graph neural network models, on collections of NFTs' images and texts. A key aspect in MERLIN is its independence on financial features, as it exploits only the primary data a user interested in NFT trading would like to deal with, i.e., NFT images and textual descriptions. By learning dense representations of such data, a price-category classification task is performed by MERLIN models, which can also be tuned according to user preferences in the inference phase to mimic different risk-return investment profiles. Experimental evaluation on a publicly available dataset has shown that MERLIN models achieve significant performances according to several financial assessment criteria, fostering profitable investments, and also beating baseline machine-learning classifiers based on financial features.

Open access
2 source records
Stock Market Forecasting Methods
Music and Audio Processing
Topic Modeling
Original source
Jan 12, 2023·Informatics
52 cites
Blockchain Propels Tourism Industry—An Attempt to Explore Topics and Information in Smart Tourism Management through Text Mining and Machine Learning

Vikram Puri, Subhra R. Mondal, Subhankar Das, Vasiliki Vrana

Blockchain and immersive technology are the pioneers in bringing digitalization to tourism, and researchers worldwide are exploring many facets of these techniques. This paper analyzes the various aspects of blockchain technology and its potential use in tourism. We explore high-frequency keywords, perform network analysis of relevant publications to analyze patterns, and introduce machine learning techniques to facilitate systematic reviews. We focused on 94 publications from Web Science that dealt with blockchain implementation in tourism from 2017 to 2022. We used Vosviewer for network analysis and artificial intelligence models with the help of machine learning tools to predict the relevance of the work. Many reviewed articles mainly deal with blockchain in tourism and related terms such as smart tourism and crypto tourism. This study is the first attempt to use text analysis to improve the topic modeling of blockchain in tourism. It comprehensively analyzes the technology’s potential use in the hospitality, accommodation, and booking industry. In this context, the paper provides significant value to researchers by giving an insight into the trends and keyword patterns. Tourism still has many unexplored areas; journal articles should also feature special studies on this topic.

Open access
Digital Marketing and Social Media
Consumer Behavior in Brand Consumption and Identification
Diverse Aspects of Tourism Research
Original source
Jan 1, 2023·IEEE Access
4 cites
Annotators’ Selection Impact on the Creation of a Sentiment Corpus for the Cryptocurrency Financial Domain

Manoel Fernando Alonso Gadi, Miguel‐Ángel Sicilia

Well labeled natural language corpus data is essential for most natural language processing techniques, especially in specialized fields. However, cohort biases remain a significant challenge in machine learning. The narrow origin of data sampling or human annotators in cohorts is a prevalent issue for machine learning researchers due to its potential to induce bias in the final product. During the development of the CryptoLin corpus for another research project, the authors became concerned about the potential influence of cohort bias on the selection of annotators. Therefore, this paper addresses the question of whether cohort diversity improves the labeling result through the implementation of a repeated annotator process, involving two annotator cohorts and a statistically robust comparison methodology. The utilization of statistical tests, such as the Chi-Square Independence test for absolute frequency tables, and the construction of confidence intervals for Kappa point estimates, facilitates a rigorous analysis of the differences between Kappa estimates. Furthermore, the application of a two-proportion z-test to compare the accuracy scores of UTAD and IE annotators for various pre-trained models, including Vader Sentiment Analysis, TextBlob Sentiment Analysis, Flair NLP library, and FinBERT Financial Sentiment Analysis with BERT, contributes to the advancement of knowledge in this field. The paper utilizes Cryptocurrency Linguo (CryptoLin), a corpus containing 2683 cryptocurrency-related news articles spanning more than three years,and compares two different selection criteria for the annotators. CryptoLin was annotated twice with discrete values representing negative, neutral, and positive news respectively. The first annotation was done by twenty-seven annotators from the same cohort. Each news title was randomly assigned and blindly annotated by three human annotators. The second annotation was carried out by eighty-three annotators from three cohorts. Each news title was randomly assigned and blindly annotated by three human annotators, one in each different cohort. In both annotations, a consensus mechanism using simple voting was applied. The first annotation used the same cohort with students from the same nationality and background. The second used three cohorts with students from a very diverse set of nationalities and educational backgrounds. The results demonstrate that manual labeling done by both groups was acceptable according to inter-rater reliability coefficients Fleiss’s Kappa, Krippendorff’s Alpha, and Gwet’s AC1. Preliminary analysis utilizing Vader, Textblob, Flair, and FinBERT confirmed the utility of the data set labeling for further refinement of sentiment analysis algorithms. Our results also highlight that the more diverse annotator pool performed better in all measured aspects.

Open access
Topic Modeling
Advanced Text Analysis Techniques
Stock Market Forecasting Methods
Original source
Jan 1, 2023·IEEE Access
19 cites
Unveiling Cryptocurrency Conversations: Insights From Data Mining and Unsupervised Learning Across Multiple Platforms

Hae Sun Jung, Haein Lee, Jang Hyun Kim

The rapid growth of the cryptocurrency market has led to an increasing interest in the subject. Cryptocurrency is now recognized as an asset, and laws and financial regulations have begun to emerge for supporting its practical use. As a result, it has become essential to perform data mining and attain knowledge from text data related to cryptocurrency. Previous studies have focused on analyzing data from a single source such as Twitter. However, there are unique insights to be gained from data across multiple platforms. In the present study, we utilized data mining techniques to extract insights from LexisNexis, Web of Science, and Reddit, representing the media, academia, and general public, respectively. Among unsupervised learning technologies, topic modeling was employed for the analysis. Topic modeling is a methodology that uncovers hidden meanings within the collected data. Among the diverse topic modeling techniques available, bidirectional encoder representations from transformers topic was chosen for the analysis. BERTopic considered to be state-of-the-art in the field of topic modeling. Dynamic topic modeling was employed to track changes in themes over time. Our experimental results reveal a tendency in the news to cover major events related to cryptocurrencies, such as regulatory developments and market trends. Academic papers, on the other hand, tend to focus on the technology behind cryptocurrencies and related research. Finally, social media conversations center more around information delivery from an investor’s psychological perspective, such as market sentiment and investment strategies.

Open access
Blockchain Technology Applications and Security
FinTech, Crowdfunding, Digital Finance
Digital Marketing and Social Media
Original source
Nov 30, 2022·The Journal of the Korea Contents Association
1 cites
Analysis of Research Trends Related to Non-Fungible Tokens (NFT) Using Topic Modeling

Jinyoung Han, Min-Jeong Lee, Ji-In Lee

대체불가토큰(Non-Fungible Token: NFT)은 현재 예술, 엔터테인먼트, 게임을 중심으로 급성장하고 있으며, 메타버스를 중심으로 패션, 마케팅 산업에도 적극적으로 도입되고 있다. NFT는 이제 산업이 형성되고 있어 NFT와 관련된 연구는 아직 탐색적인 수준이다. 본 연구에서는 2018년부터 2022년 6월까지 출간된 NFT 관련 국내외 논문 각각 95편, 153편의 서지정보를 LDA 기법으로 토픽모델링을 수행하였다. 이를 통해 키워드간 관계성, NFT 관련 분야 전체 연구의 구조를 살펴보고 국내외 NFT 관련 연구 주제의 유사점과 차이점을 비교하였으며, 분석결과와 선행연구를 기반으로 NFT 분야의 향후 연구 방향을 제시하였다. 본 연구는 NFT 관련 국내외 연구를 통합하여 진행한 첫 리뷰 연구로서 의의가 있으며 실무적으로 NFT 산업 초기 단계에 정부가 관련 정책 수립하는데 현황을 파악하고, 정부의 역할을 인지하는데 기여할 것으로 기대한다.

Open access
Diverse Topics in Contemporary Research
Technology and Data Analysis
Diverse Approaches in Healthcare and Education Studies
Original source
Aug 26, 2022·International Journal of Information Management Data Insights
63 cites
Blockchain technology for cybersecurity: A text mining literature analysis

Ravi Prakash, V.S. Anoop, S. Asharaf

Blockchain, the technology infrastructure behind the famous cryptocurrency bitcoin, can take away the notion of trust from centralized organizations to a decentralized platform that is mathematically verifiable and cryptographically secure. It is gaining more significant momentum exponentially and disrupts the way businesses function beyond the digital currency aspects. This work presents a text mining literature analysis of research articles published in major digital libraries on blockchain technology and cybersecurity. This literature analysis employs automated text mining approaches such as topic modeling and keyphrase extraction for unearthing the themes from a vast body of literature. This analysis highlights the multidisciplinary nature of blockchain technology within the cybersecurity domain. The findings also show the cyber threats and vulnerabilities that evolve with blockchain technology developments. This analysis also showcases the computer security research community’s vulnerabilities and provides future research dimensions that are crucial for designing secure blockchain applications and platforms.

Open access
Blockchain Technology Applications and Security
Cybercrime and Law Enforcement Studies
Spam and Phishing Detection
Original source
May 28, 2022·arXiv (Cornell University)
4 cites
A New High-Performance Approach to Approximate Pattern-Matching for Plagiarism Detection in Blockchain-Based Non-Fungible Tokens (NFTs)

Ciprian Pungilă, Darius Galiş, Viorel Negru

We are presenting a fast and innovative approach to performing approximate pattern-matching for plagiarism detection, using an NDFA-based approach that significantly enhances performance compared to other existing similarity measures. We outline the advantages of our approach in the context of blockchain-based non-fungible tokens (NFTs). We present, formalize, discuss and test our proposed approach in several real-world scenarios and with different similarity measures commonly used in plagiarism detection, and observe significant throughput enhancements throughout the entire spectrum of tests, with little to no compromises on the accuracy of the detection process overall. We conclude that our approach is suitable and adequate to perform approximate pattern-matching for plagiarism detection, and outline research directions for future improvements.

Open access
2 source records
Academic integrity and plagiarism
Imbalanced Data Classification Techniques
Topic Modeling
Original source
Mar 17, 2022·arXiv (Cornell University)
1 cites
Short Text Topic Modeling: Application to tweets about Bitcoin

Hugo Schnoering

Understanding the semantic of a collection of texts is a challenging task. Topic models are probabilistic models that aims at extracting "topics" from a corpus of documents. This task is particularly difficult when the corpus is composed of short texts, such as posts on social networks. Following several previous research papers, we explore in this paper a set of collected tweets about bitcoin. In this work, we train three topic models and evaluate their output with several scores. We also propose a concrete application of the extracted topics.

Open access
2 source records
cs.IR
cs.LG
Advanced Text Analysis Techniques
Original source
Jan 1, 2022·ETRI Journal
5 cites
Exploring trends in blockchain publications with topic modeling: Implications for forecasting the emergence of industry applications

Jeongho Lee, Hangjung Zo, Tom Steinberger

Abstract Technological innovation generates products, services, and processes that can disrupt existing industries and lead to the emergence of new fields. Distributed ledger technology, or blockchain, offers novel transparency, security, and anonymity characteristics in transaction data that may disrupt existing industries. However, research attention has largely examined its application to finance. Less is known of any broader applications, particularly in Industry 4.0. This study investigates academic research publications on blockchain and predicts emerging industries using academia‐industry dynamics. This study adopts latent Dirichlet allocation and dynamic topic models to analyze large text data with a high capacity for dimensionality reduction. Prior studies confirm that research contributes to technological innovation through spillover, including products, processes, and services. This study predicts emerging industries that will likely incorporate blockchain technology using insights from the knowledge structure of publications.

Open access
2 source records
Blockchain Technology Applications and Security
FinTech, Crowdfunding, Digital Finance
Digital Marketing and Social Media
Original source
Jan 1, 2022·IEEE Access
25 cites
What Users Tweet on NFTs: Mining Twitter to Understand NFT-Related Concerns Using a Topic Modeling Approach

Chris Meyns, Fisnik Dalipi

Non-fungible token (NFT) trade has grown drastically over recent years. While scholarship on the technical aspects and potential applications of NFTs has been steadily increasing, less attention has been directed to the human perception of or attitudes toward this new type of digital asset. The aim of this research is to investigate what concerns are expressed in relation to non-fungible tokens by those who engage with NFTs on the social media platform Twitter. In this study, data was gathered through online social media data mining of NFT-related posts on Twitter. Two datasets (with 18,373 and 36,354 individual tweet records, respectively) were obtained. Topic modeling was used as a method of data analysis. Our results reveal 19 overall themes of concerns around NFTs as expressed on Twitter, which broadly fall into two categories: concerns about attacks and threats by third parties; and concerns about trading and the role of marketplaces. Overall, this study offers a better understanding of the expressions of concern, uncertainty, and the perception of possible barriers related to NFT trading. These findings contribute to theoretical insight and can, moreover, function as a basis for developing practical design and policy interventions.

Open access
Digital Marketing and Social Media
Consumer Behavior in Brand Consumption and Identification
Consumer Market Behavior and Pricing
Original source
Nov 28, 2021·arXiv (Cornell University)
5 cites
Semantic Code Search for Smart Contracts

Chaochen Shi, Yong Xiang, Jiangshan Yu, Longxiang Gao

Semantic code search technology allows searching for existing code snippets through natural language, which can greatly improve programming efficiency. Smart contracts, programs that run on the blockchain, have a code reuse rate of more than 90%, which means developers have a great demand for semantic code search tools. However, the existing code search models still have a semantic gap between code and query, and perform poorly on specialized queries of smart contracts. In this paper, we propose a Multi-Modal Smart contract Code Search (MM-SCS) model. Specifically, we construct a Contract Elements Dependency Graph (CEDG) for MM-SCS as an additional modality to capture the data-flow and control-flow information of the code. To make the model more focused on the key contextual information, we use a multi-head attention network to generate embeddings for code features. In addition, we use a fine-tuned pretrained model to ensure the model's effectiveness when the training data is small. We compared MM-SCS with four state-of-the-art models on a dataset with 470K (code, docstring) pairs collected from Github and Etherscan. Experimental results show that MM-SCS achieves an MRR (Mean Reciprocal Rank) of 0.572, outperforming four state-of-the-art models UNIF, DeepCS, CARLCS-CNN, and TAB-CS by 34.2%, 59.3%, 36.8%, and 14.1%, respectively. Additionally, the search speed of MM-SCS is second only to UNIF, reaching 0.34s/query.

Open access
2 source records
Software Engineering Research
Topic Modeling
Advanced Malware Detection Techniques
Original source
Jun 2, 2021·arXiv (Cornell University)
2 cites
multiPRover: Generating Multiple Proofs for Improved Interpretability in\n Rule Reasoning

Swarnadeep Saha, Prateek Yadav, Mohit Bansal

We focus on a type of linguistic formal reasoning where the goal is to reason\nover explicit knowledge in the form of natural language facts and rules (Clark\net al., 2020). A recent work, named PRover (Saha et al., 2020), performs such\nreasoning by answering a question and also generating a proof graph that\nexplains the answer. However, compositional reasoning is not always unique and\nthere may be multiple ways of reaching the correct answer. Thus, in our work,\nwe address a new and challenging problem of generating multiple proof graphs\nfor reasoning over natural language rule-bases. Each proof provides a different\nrationale for the answer, thereby improving the interpretability of such\nreasoning systems. In order to jointly learn from all proof graphs and exploit\nthe correlations between multiple proofs for a question, we pose this task as a\nset generation problem over structured output spaces where each proof is\nrepresented as a directed graph. We propose two variants of a proof-set\ngeneration model, multiPRover. Our first model, Multilabel-multiPRover,\ngenerates a set of proofs via multi-label classification and implicit\nconditioning between the proofs; while the second model, Iterative-multiPRover,\ngenerates proofs iteratively by explicitly conditioning on the previously\ngenerated proofs. Experiments on multiple synthetic, zero-shot, and\nhuman-paraphrased datasets reveal that both multiPRover models significantly\noutperform PRover on datasets containing multiple gold proofs.\nIterative-multiPRover obtains state-of-the-art proof F1 in zero-shot scenarios\nwhere all examples have single correct proofs. It also generalizes better to\nquestions requiring higher depths of reasoning where multiple proofs are more\nfrequent. Our code and models are publicly available at\nhttps://github.com/swarnaHub/multiPRover\n

Open access
2 source records
Natural Language Processing Techniques
Topic Modeling
Semantic Web and Ontologies
Original source
Mar 12, 2021·arXiv (Cornell University)
63 cites
A Multi-Modal Transformer-based Code Summarization Approach for Smart Contracts

Zhen Yang, Jacky Keung, Xiao Yu, Xiaodong Gu · 7 authors

Code comment has been an important part of computer programs, greatly facilitating the understanding and maintenance of source code. However, high-quality code comments are often unavailable in smart contracts, the increasingly popular programs that run on the blockchain. In this paper, we propose a Multi-Modal Transformer-based (MMTrans) code summarization approach for smart contracts. Specifically, the MMTrans learns the representation of source code from the two heterogeneous modalities of the Abstract Syntax Tree (AST), i.e., Structure-based Traversal (SBT) sequences and graphs. The SBT sequence provides the global semantic information of AST, while the graph convolution focuses on the local details. The MMTrans uses two encoders to extract both global and local semantic information from the two modalities respectively, and then uses a joint decoder to generate code comments. Both the encoders and the decoder employ the multi-head attention structure of the Transformer to enhance the ability to capture the long-range dependencies between code tokens. We build a dataset with over 300K pairs of smart contracts, and evaluate the MMTrans on it. The experimental results demonstrate that the MMTrans outperforms the state-of-the-art baselines in terms of four evaluation metrics by a substantial margin, and can generate higher quality comments.

Open access
3 source records
Software Engineering Research
Topic Modeling
Advanced Malware Detection Techniques
Original source
Mar 4, 2021·Symmetry
27 cites
Technology Hotspot Tracking: Topic Discovery and Evolution of China’s Blockchain Patents Based on a Dynamic LDA Model

Jinli Wang, Yong Fan, Hui Zhang, Libo Feng

Tracking scientific and technological (S&T) research hotspots can help scholars to grasp the status of current research and develop regular patterns in the field over time. It contributes to the generation of new ideas and plays an important role in promoting the writing of scientific research projects and scientific papers. Patents are important S&T resources, which can reflect the development status of the field. In this paper, we use topic modeling, topic intensity, and evolutionary computing models to discover research hotspots and development trends in the field of blockchain patents. First, we propose a time-based dynamic latent Dirichlet allocation (TDLDA) modeling method based on a probabilistic graph model and knowledge representation learning for patent text mining. Second, we present a computational model, topic intensity (TI), that expresses the topic strength and evolution. Finally, the point-wise mutual information (PMI) value is used to evaluate topic quality. We obtain 20 hot topics through TDLDA experiments and rank them according to the strength calculation model. The topic evolution model is used to analyze the topic evolution trend from the perspectives of rising, falling, and stable. From the experiments we found that 8 topics showed an upward trend, 6 topics showed a downward trend, and 6 topics became stable or fluctuated. Compared with the baseline method, TDLDA can have the best effect when K is 40 or less. TDLDA is an effective topic model that can extract hot topics and evolution trends of blockchain patent texts, which helps researchers to more accurately grasp the research direction and improves the quality of project application and paper writing in the blockchain technology domain.

Open access
Machine Learning in Materials Science
Intellectual Property and Patents
Computational and Text Analysis Methods
Original source
Jan 1, 2021·Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
15 cites
multiPRover: Generating Multiple Proofs for Improved Interpretability in Rule Reasoning

Swarnadeep Saha, Prateek Yadav, Mohit Bansal

We focus on a type of linguistic formal reasoning where the goal is to reason over explicit knowledge in the form of natural language facts and rules A recent work, named PROVER However, compositional reasoning is not always unique and there may be multiple ways of reaching the correct answer. Thus, in our work, we address a new and challenging problem of generating multiple proof graphs for reasoning over natural language rule-bases. Each proof provides a different rationale for the answer, thereby improving the interpretability of such reasoning systems. In order to jointly learn from all proof graphs and exploit the correlations between multiple proofs for a question, we pose this task as a set generation problem over structured output spaces where each proof is represented as a directed graph. We propose two variants of a proof-set generation model, MULTIPROVER. Our first model, Multilabel-MULTIPROVER, generates a set of proofs via multi-label classification and implicit conditioning between the proofs; while the second model, Iterative-MULTIPROVER, generates proofs iteratively by explicitly conditioning on the previously generated proofs. Experiments on multiple synthetic, zero-shot, and human-paraphrased datasets reveal that both MULTIPROVER models significantly outperform PROVER on datasets containing multiple gold proofs. Iterative-MULTIPROVER obtains state-of-the-art proof F1 in zero-shot scenarios where all examples have single correct proofs. It also generalizes better to questions requiring higher depths of reasoning where multiple proofs are more frequent.

Open access
Topic Modeling
Natural Language Processing Techniques
Multimodal Machine Learning Applications
Original source
Sep 22, 2020·PLoS Biology
59 cites
Quantifying and contextualizing the impact of bioRxiv preprints through automated social media audience segmentation

Jedidiah Carlson, Kelley Harris

Engagement with scientific manuscripts is frequently facilitated by Twitter and other social media platforms. As such, the demographics of a paper's social media audience provide a wealth of information about how scholarly research is transmitted, consumed, and interpreted by online communities. By paying attention to public perceptions of their publications, scientists can learn whether their research is stimulating positive scholarly and public thought. They can also become aware of potentially negative patterns of interest from groups that misinterpret their work in harmful ways, either willfully or unintentionally, and devise strategies for altering their messaging to mitigate these impacts. In this study, we collected 331,696 Twitter posts referencing 1,800 highly tweeted bioRxiv preprints and leveraged topic modeling to infer the characteristics of various communities engaging with each preprint on Twitter. We agnostically learned the characteristics of these audience sectors from keywords each user's followers provide in their Twitter biographies. We estimate that 96% of the preprints analyzed are dominated by academic audiences on Twitter, suggesting that social media attention does not always correspond to greater public exposure. We further demonstrate how our audience segmentation method can quantify the level of interest from nonspecialist audience sectors such as mental health advocates, dog lovers, video game developers, vegans, bitcoin investors, conspiracy theorists, journalists, religious groups, and political constituencies. Surprisingly, we also found that 10% of the preprints analyzed have sizable (>5%) audience sectors that are associated with right-wing white nationalist communities. Although none of these preprints appear to intentionally espouse any right-wing extremist messages, cases exist in which extremist appropriation comprises more than 50% of the tweets referencing a given preprint. These results present unique opportunities for improving and contextualizing the public discourse surrounding scientific research.

Open access
Academic Publishing and Open Access
Misinformation and Its Impacts
Social Media in Health Education
Original source
Apr 1, 2020·Expert Systems with Applications
174 cites
Word2vec-based latent semantic analysis (W2V-LSA) for topic modeling: A study on blockchain technology trend analysis

Suhyeon Kim, Haecheong Park, Junghye Lee

Blockchain has become one of the core technologies in Industry 4.0. To help decision-makers establish action plans based on blockchain, it is an urgent task to analyze trends in blockchain technology. However, most of existing studies on blockchain trend analysis are based on effort demanding full-text investigation or traditional bibliometric methods whose study scope is limited to a frequency-based statistical analysis. Therefore, in this paper, we propose a new topic modeling method called Word2vec-based Latent Semantic Analysis (W2V-LSA), which is based on Word2vec and Spherical k-means clustering to better capture and represent the context of a corpus. We then used W2V-LSA to perform an annual trend analysis of blockchain research by country and time for 231 abstracts of blockchain-related papers published over the past five years. The performance of the proposed algorithm was compared to Probabilistic LSA, one of the common topic modeling techniques. The experimental results confirmed the usefulness of W2V-LSA in terms of the accuracy and diversity of topics by quantitative and qualitative evaluation. The proposed method can be a competitive alternative for better topic modeling to provide direction for future research in technology trend analysis and it is applicable to various expert systems related to text mining.

Open access
Blockchain Technology Applications and Security
Innovation Diffusion and Forecasting
Digital Marketing and Social Media
Original source
Jan 1, 2020·Proceedings of the ... Annual Hawaii International Conference on System Sciences/Proceedings of the Annual Hawaii International Conference on System Sciences
14 cites
A Cross-Disciplinary Review of Blockchain Research Trends and Methodologies: Topic Modeling Approach

Muhammad Nauman Shahid

Given the increasing interest in blockchain technology, we present a large-scale cross-disciplinary literature analysis of research on the blockchain using topic modelling with the goal of identifying the major research trends, research methodologies, and fruitful areas for further research. In particular, the analysis focuses on abstracting out research trends from relevant terms and topics related to the research disciplines of Business, Computer Science, Economics, Social Sciences, Engineering, Healthcare, and Law. A total of 2,125 articles published between 2008 to up until early 2019 in academic journals and conferences were analyzed. Results of our analysis reveal that research is bipartite between practical and research domains, with academic research on blockchain not clearly aligning with organizational and social benefits. Also, we found – 1) few inter-disciplinary publications, and 2) a small number of studies that use surveys, experiments, and case studies as their research method. Our findings also reveal that research on Blockchain in the social sciences and law is still in the embryonic stage, thus making it essential to develop more direct research efforts for Blockchain to thrive in all research disciplines.

Open access
Blockchain Technology Applications and Security
FinTech, Crowdfunding, Digital Finance
Impact of AI and Big Data on Business and Society
Original source
Jan 1, 2020·Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)
8 cites
PRover: Proof Generation for Interpretable Reasoning over Rules

Swarnadeep Saha, Sayan Ghosh, Shashank Srivastava, Mohit Bansal

shows that transformers can act as "soft theorem provers" by answering questions over explicitly provided knowledge in natural language. In our work, we take a step closer to emulating formal theorem provers, by proposing PROVER, an interpretable transformer-based model that jointly answers binary questions over rule-bases and generates the corresponding proofs. Our model learns to predict nodes and edges corresponding to proof graphs in an efficient constrained training paradigm. During inference, a valid proof, satisfying a set of global constraints is generated. We conduct experiments on synthetic, hand-authored, and human-paraphrased rule-bases to show promising results for QA and proof generation, with strong generalization performance. First, PROVER generates proofs with an accuracy of 87%, while retaining or improving performance on the QA task, compared to RuleTakers (up to 6% improvement on zero-shot evaluation). Second, when trained on questions requiring lower depths of reasoning, it generalizes significantly better to higher depths (up to 15% improvement). Third, PROVER obtains near perfect QA accuracy of 98% using only 40% of the training data. However, generating proofs for questions requiring higher depths of reasoning becomes challenging, and the accuracy drops to 65% for "depth 5", indicating significant scope for future work.

Open access
Topic Modeling
Natural Language Processing Techniques
Explainable Artificial Intelligence (XAI)
Original source
Jan 1, 2020·IEEE Access
131 cites
Charting the Landscape of Online Cryptocurrency Manipulation

Leonardo Nizzoli, Serena Tardelli, Marco Avvenuti, Stefano Cresci · 6 authors

Cryptocurrencies represent one of the most attractive markets for financial speculation. As a consequence, they have attracted unprecedented attention on social media. Besides genuine discussions and legitimate investment initiatives, several deceptive activities have flourished. In this work, we chart the online cryptocurrency landscape across multiple platforms. To reach our goal, we collected a large dataset, composed of more than 50M messages published by almost 7M users on Twitter, Telegram and Discord, over three months. We performed bot detection on Twitter accounts sharing invite links to Telegram and Discord channels, and we discovered that more than 56% of them were bots or suspended accounts. Then, we applied topic modeling techniques to Telegram and Discord messages, unveiling two different deception schemes - “pump-and-dump” and “Ponzi” - and identifying the channels involved in these frauds. Whereas on Discord we found a negligible level of deception, on Telegram we retrieved 296 channels involved in pump-and-dump and 432 involved in Ponzi schemes, accounting for a striking 20% of the total. Moreover, we observed that 93% of the invite links shared by Twitter bots point to Telegram pump-and-dump channels, shedding light on a little-known social bot activity. Charting the landscape of online cryptocurrency manipulation can inform actionable policies to fight such abuse.

Open access
3 source records
Spam and Phishing Detection
Cybercrime and Law Enforcement Studies
Misinformation and Its Impacts
Original source