This working paper examines a structural shift in search behaviour affecting crypto and Web3 brands: the divergence between AI-summarised retrieval (zero-click) and deep-intent click-through traffic. Drawing on published zero-click search data and observed AI citation patterns, it identifies why high-volume content strategies are failing in AI-mediated search environments and outlines three asset classes - statistics hubs, original industry research, and high-utility tools - that generate the editorial backlink profiles AI models use to select citation sources. The paper defines the concept of "authority infrastructure" as a capital investment in linkable assets with compounding residual value, contrasting this with recurring spend on keyword-optimised content with no durable equity. Intended for crypto protocol marketing teams, Web3 founders, and DeFi growth leads evaluating content strategy for AI search visibility. Published by David Wood, CryptoContent.dev.
Due to a combination of both rigid binary structures within formal identification systems and ambiguous laws; along with discriminatory practices, third gender individuals continue to be excluded from formal identity systems. Most existing centralized identity verification systems have failed to provide non-binary identity solutions, resulting in limited access to banking services, education services, legal protection and access to health care. This paper examines the ability of decentralized identity systems using blockchain technologies to provide users with secure, private and self-managed identity options for third gender individuals. It evaluates the core technology of decentralized identity systems, specifically decentralized identifiers (DIDs); verifiable credentials (VCs); smart contracts; and zero knowledge proof; through an examination of real-world examples. Additionally, the paper outlines the technical and legal constraints associated with decentralized identity systems, specifically literacy requirements; lack of consistency in national frameworks for recognition; and risk of symbolic inclusion (i.e., "being included" rather than having the rights of recognition) rather than actual structural reforms.
Process attestation verifies human authorship by collecting behavioral biometric evidence, including keystroke dynamics, typing patterns, and editing behavior, during the creative process. However, the very data needed to prove authenticity can reveal intimate details about an author's cognitive state, health conditions, and identity, constituting sensitive biometric data under GDPR Article 9. We resolve this privacy-attestation paradox using zero-knowledge proofs. We present ZK-PoP, a construction that allows a verifier to confirm that (a) sequential work function chains were computed correctly, (b) behavioral feature vectors fall within human population distributions, and (c) content evolution is consistent with incremental human editing, all without learning the underlying behavioral data, exact timing, or intermediate content. Our construction uses Groth16 proofs over arithmetic circuits with Pedersen commitments and Bulletproof range proofs. We prove that ZK-PoP is computationally zero-knowledge, computationally sound, and achieves unlinkability across sessions. Evaluation shows proof generation in under 30 seconds for a 1-hour writing session, with 192-byte proofs verifiable in 8.2 ms, while incurring less than 5% accuracy loss in simulation at practical privacy levels (epsilon >= 1.0) compared to non-private baselines.
The "Identity Trilemma" posits that a decentralized network can enforce only two of the following three properties: Privacy (Anonymity), Accountability (Sybil Resistance), and Permissionlessness (No Central Gatekeeper). Traditional Web2 platforms resolve this by sacrificing Privacy (enforcing Real-Name Policies), while early Web3 platforms sacrificed Accountability, resulting in "Sybil Swarms" where single actors control thousands of wallets. This paper introduces the Klyrox solution to the trilemma: Pseudonymous Accountability. By utilizing Zero-Knowledge Proofs (ZKPs) and non-linear Time-Energy Cost Functions, the Klyrox Protocol enables users to mathematically prove they are unique, high-integrity actors without ever revealing their physical identity, biometric data, or government credentials. We define a new standard for "Proof of Personhood" based not on biology, but on consistent historical behavior recorded in a Soulbound Token (ERC-721M). Author's Note: This paper is a foundational pillar of the Klyrox Protocol architecture, expanding upon the core framework published in The Klyrox Protocol: A Decentralized Framework for Optimistic Content Verification and Epistemic Reputation (available at: https://doi.org/10.5281/zenodo.18729968). It outlines the specific mechanics underpinning the concept of "Epistemic Capital," as explored in the complete five-volume series, The Algorithmic Monographs (The Algorithmic Invisible Hand, The Republic of Code, The Market for Truth, The Heavy Metal Intelligence, and The Synthetic C-Suite).
LIU Ronglong, LI Ziwei, WAN Yue, WU Jiajing, JIANG Zigui
As the paradigm of âłdecentralized next-generation Internet,âł Web3, relying on blockchain technology, has become an emerging field with great potential in the digital intelligence service ecosystem. However, Web3 phishing websites pose a serious threat to ecological health. Phishers carefully design domain names as the primary bait, inducing users to visit and engage in high-risk operations to steal digital assets. Currently, the antiphishing works of Web3 primarily focus on phishing account detection, phishing transaction detection, and phishing gang mining, whereas the existing phishing website domain name detection primarily targets traditional phishing websites, which have limitations such as insufficient adaptability and a lack of systematic analysis. To this end, a detection method called WPWHunter is proposed for Web3 phishing website domain names, which conducts multidimensional analysis on the detected real Web3 phishing websites and explores the potential application of Large Language Model (LLM) in web page analysis. The WPWHunter algorithm detects three features in Web3 phishing website domain names: inducing words, visual deception, and item name imitation. The experimental results show that WPWHunter can effectively detect suspicious Web3 phishing domains with a G-means index of 0.769 on a test set, which is 0.048 higher than that of the best-performing baseline method. Additionally, as a supplementary exploratory experiment, three universal LLM are used to analyze the content of Web3 phishing websites that WPWHunter failed to detect and the logic used by LLM to determine Web3 phishing websites is summarized.
Luis deâMarcos, AdriĂĄn DomĂnguezâDĂaz, Javier Junquera-SĂĄnchez, Carlos Cilleruelo ¡ 5 authors
The Dark Web, a hidden segment of the internet, has become a hub for illicit activities, facilitated by various forms of digital identification (IDs) such as email addresses, Telegram accounts, and cryptocurrency wallets. This study conducts a comprehensive analysis of the Dark Webâs identification and communication patterns, focusing on the roles of different ID types and their associated activities. Using a dataset of Dark Web documents, we construct and analyze a bipartite network to model the relationships between IDs and web documents, employing graphâtheoretical metrics such as degree centrality, closeness centrality, betweenness centrality, and k-core decomposition, while analyzing subnetworks formed by ID type. Our findings reveal that Telegram forms the backbone of the network, serving as the primary communication tool for hacking-related activities, particularly within Russian-speaking communities. In contrast, email plays a more decentralized role, facilitating financeâcrypto and other activities but with a high level of fragmentation and English as the predominant language. XMR (Monero) wallets emerge as a key component in financial transactions, forming a cohesive subnetwork focused on cryptocurrency-related activities. The analysis also highlights the modular and hierarchical nature of the Dark Web, with distinct clusters for hacking, financeâcrypto, and drugsânarcotics, often operating independently but with some cross-topic interactions. This study provides a foundation for understanding the Dark Webâs structure and dynamics, offering insights that can inform strategies for monitoring and mitigating its risks.
Phishing attacks in Web3 ecosystems are increasingly sophisticated, exploiting deceptive contract logic, malicious frontend scripts, and token approval patterns. We present DeepTx, a real-time transaction analysis system that detects such threats before user confirmation. DeepTx simulates pending transactions, extracts behavior, context, and UI features, and uses multiple large language models (LLMs) to reason about transaction intent. A consensus mechanism with self-reflection ensures robust and explainable decisions. Evaluated on our phishing dataset, DeepTx achieves high precision and recall (demo video: https://youtu.be/4OfK9KCEXUM).
Academic publishing, integral to knowledge dissemination and scientific advancement, increasingly faces threats from unethical practices such as unconsented authorship, gift authorship, author ambiguity, and undisclosed conflicts of interest. While existing infrastructures like ORCID effectively disambiguate researcher identities, they fall short in enforcing explicit authorship consent, accurately verifying contributor roles, and robustly detecting conflicts of interest during peer review. To address these shortcomings, this paper introduces a decentralized framework leveraging Self-Sovereign Identity (SSI) and blockchain technology. The proposed model uses Decentralized Identifiers (DIDs) and Verifiable Credentials (VCs) to securely verify author identities and contributions, reducing ambiguity and ensuring accurate attribution. A blockchain-based trust registry records authorship consent and peer-review activity immutably. Privacy-preserving cryptographic techniques, especially Zero-Knowledge Proofs (ZKPs), support conflict-of-interest detection without revealing sensitive data. Verified authorship metadata and consent records are embedded in publications, increasing transparency. A stakeholder survey of researchers, editors, and reviewers suggests the framework improves ethical compliance and confidence in scholarly communication. This work represents a step toward a more transparent, accountable, and trustworthy academic publishing ecosystem.
The increasing use of cryptocurrencies in criminal activities presents significant challenges to society and the judicial system, particularly in tracking and seizing illicit digital assets. Among all relevant digital evidence, mnemonic phrases, which are critical for accessing cryptocurrency wallets, are crucial digital evidence for confiscating criminal proceeds and conducting investigations. However, traditional digital forensics tools, such as the Mnemonic Library Matching Method, lack flexibility and efficiency when handling cryptocurrency-related data. This study introduces an innovative Natural Language Processing (NLP) and deep learning approach for rapid mnemonic identification across 11 languages, including English, Spanish, and Japanese. We trained and compared four NLP deep learning models: RNN, LSTM, BiLSTM, and TextCNN, on a large-scale, real-world dataset. Our analysis reveals that the Text Convolutional Neural Network (TextCNN) model exhibits superior performance, achieving a 99.9993% accuracy rate, nearly matching the 100% accuracy of the Mnemonic Library Matching Method. Crucially, our TextCNN-driven approach processes data 40.47 times faster than the traditional method, significantly enhancing efficiency in time-sensitive forensic environments. This NLP-driven method not only maintains high accuracy while dramatically reducing processing time but also offers greater adaptability for diverse forensic needs compared to traditional techniques. By enabling more effective tracking and seizure of criminal assets, this approach aims to address the broader societal and judicial challenges posed by cryptocurrency-related criminal activities. Our research showcases the potential of NLP and deep learning in digital forensics, providing law enforcement with advanced tools for investigating cryptocurrency-related crimes and curbing the misuse of cryptocurrencies in illicit activities.
Manoel Fernando Alonso Gadi, MiguelâĂngel Sicilia
Abstract The objective of this paper is to describe Cryptocurrency Linguo (CryptoLin), a novel corpus containing 2683 cryptocurrency-related news articles covering more than a three-year period. CryptoLin was human-annotated with discrete values representing negative, neutral, and positive news respectively. Eighty-three people participated in the annotation process; each news title was randomly assigned and blindly annotated by three human annotators, one in each different cohort, followed by a consensus mechanism using simple voting. The selection of the annotators was intentionally made using three cohorts with students from a very diverse set of nationalities and educational backgrounds to minimize bias as much as possible. In case one of the annotators was in total disagreement with the other two (e.g., one negative vs two positive or one positive vs two negative), we considered this minority report and defaulted the labeling to neutral. Fleissâs Kappa, Krippendorffâs Alpha, and Gwetâs AC1 inter-rater reliability coefficients demonstrate CryptoLinâs acceptable quality of inter-annotator agreement. The dataset also includes a text span with the three manual label annotations for further auditing of the annotation mechanism. To further assess the quality of the labeling and the usefulness of CryptoLin dataset, it incorporates four pretrained Sentiment Analysis models: Vader, Textblob, Flair, and FinBERT. Vader and FinBERT demonstrate reasonable performance in the CryptoLin dataset, indicating that the data was not annotated randomly and is therefore useful for further research1. FinBERT (negative) presents the best performance, indicating an advantage of being trained with financial news. Both the CryptoLin dataset and the Jupyter Notebook with the analysis, for reproducibility, are available at the projectâs Github. Overall, CryptoLin aims to complement the current knowledge by providing a novel and publicly available Gadi and Ăngel Sicilia (Cryptolin dataset and python jupyter notebooks reproducibility codes, 2022) cryptocurrency sentiment corpus and fostering research on the topic of cryptocurrency sentiment analysis and potential applications in behavioral science. This can be useful for businesses and policymakers who want to understand how cryptocurrencies are being used and how they might be regulated. Finally, the rules for selecting and assigning annotators make CryptoLin unique and interesting for new research in annotator selection, assignment, and biases.
This study explores the unique linguistic characteristics of bitcoin, which has significantly changed the financial world over the past few years. As the concept of bitcoin, cryptocurrency and the digital network behind them are not dealt with in LSP (Language for Specific Purposes) coursebooks yet, this small-scale research is intended to fill this niche. In order to see what terminology has to be acquired to be able to understand the basic issues about bitcoin, online sources dealing with this innovative technology and its regulatory systems have been used. The selected online texts are analysed by TextStat software, which is capable of making word counts and collocation frequency. The results show us the most common collocations with bitcoin, and blockchain processing within context, as well as the most frequently used words (e.g., cryptocurrency or exchange), which definitely need to be learned by students majoring in Business English. My aim with this research is that LSP teachers get a comprehensive picture of what terminology to teach to their student when dealing with the topic of cryptocurrencies. In addition, bitcoin-related vocabulary can be integrated into other subjects, such as economics, finance or technology, allowing students to explore the connections between different fields of knowledge. Keywords: Bitcoin, cryptocurrency, blockchain, collocation, online
Summary Blockchain users are identified by addresses (public keys), which cannot be easily linked back to them without outâofânetwork information. This provides pseudoâanonymity, which is amplified when the user generates a new address for each transaction. Since all transaction history is visible to all users in public blockchains, finding affiliation between related addresses undermines pseudoâanonymity. Such affiliation information can be used to discriminate against addresses linked with undesired activities or can lead to deâanonymization if outâofânetwork information becomes available. In this work, we propose an approach to undermine pseudoâanonymity of blockchain transactions by linking together addresses that were used to deploy smart contracts, which were produced by the same authors. In our approach, we leverage stylometry techniques, widely used in the social science field for attribution of literary texts to their corresponding authors. The assumption underlying authorship attribution is the existence of a distinctive writing style, unique to an author and easily distinguishable from others. Drawing an analogy between literary text and smart contracts' source code, we explore the extent to which unique features of source code and byte code of Ethereum smart contracts can represent the coding style of smart contract developers. We show that even a small number of representative features leads to a sufficiently high accuracy in attributing smart contracts' code to its deployer's address. We further validate our approach on realâworld scammers' data and Ponzi schemeârelated contracts. Additionally, we provide an algorithm to extract distinctly contributing features per an entire dataset or per specific authors. We use this algorithm to extract and explore such features in our dataset and in the Ponzi schemeârelated dataset.
Due to increasing popularity of Bitcoin and other cryptocurrencies, proliferation of deceptive cryptocurrencies over the internet is a global concern. In this paper, we have identified a set of 24 features through analyzing Cryptocurrency Market Capitalization (CMC) data and propose a Multilayer Perceptron (MLP) architecture for detecting deceptive cryptocurrencies. The proposed MLP architecture is compared with three traditional machine learning algorithms over a real cryptocurrency dataset crawled from CMC website, and it performs significantly better.
Public software repositories such as GitHub make transparent the development history of an open source software system. Source code commits, discussions about new features and bugs, and code reviews are stored and carefully attributed to the appropriate developers. However, sometimes governments may seek to analyze these repositories, to identify citizens who contribute to projects they disapprove of, such as those involving cryptography or social media. While developers who seek anonymity may contribute under assumed identities, their body of public work may be characteristic enough to betray who they really are. The ability to contribute anonymously to public bodies of knowledge is extremely important to the future of technological and intellectual freedoms. Just as in security hacking, the only way to protect vulnerable individuals is by demonstrating the means and strength of available attacks so that those concerned may know of the need and develop the means to protect themselves. \n \nIn this work, we present a method to de-anonymize source code contributors based on the authors' intrinsic programming style. First, we present a partial replication study wherein we attempt to de-anonymize a large number of entries into the Google Code Jam competition. We base our approach on Caliskan-Islam et al. 2015, but with modifications to the feature set and modelling strategy for scalability and feature-selection robustness. We did not achieve 0.98 F1 achieved in this prior work, but managed a still reasonable 0.71 F1 under identical experimental conditions, and a 0.88 F1 given more data from the same set. \n \nSecond, we present an exploratory study focused on de-anonymizing programmers who have contributed to a repository, using other commits from the same repository as training data. We train random-forest classifiers using programmer data collected from 37 medium to large open-source repositories. Given a choice between active developers in a project, we were able to correctly determine authorship of a given function about 75% of the time, without the use of identifying meta-data or comments. We were also able to correctly validate a contributor as the author of a questioned function with 80\\% recall and 65\\% precision. This exploratory study provides empirical support for our approach. \n \nFinally, we present the results of a similar, but more difficult study wherein we attempt de-anonymize a repository in the same manner, but without using the target repository as training data. To do this, we gather as much training data as possible from the repository's contributors through the Github API. We evaluate our technique over 3 repositories: Bitcoin, Ethereum (crypto-currencies) and TrinityCore (a game engine). Our results in this experiment starkly contrast our results in the intra-repository study showing accuracies of 35% for Bitcoin, 22% for Ethereum, and 21% for TrinityCore which had candidate set sizes of 6, 5, and 7 respectively. \n \nOur results indicate that we can do somewhat better than random guessing, even under difficult experimental conditions, but they also indicate some fundamental issues with the state of the art of Code Stylometry. In this work we present our methodology, results, and some comments on past empirical studies, the difficulties we faced, and likely hurdles for future work in the area.
Rebecca S. Portnoff, Danny Yuxing Huang, Periwinkle Doerfler, Sadia Afroz ¡ 5 authors
Sites for online classified ads selling sex are widely used by human traffickers to support their pernicious business. The sheer quantity of ads makes manual exploration and analysis unscalable. In addition, discerning whether an ad is advertising a trafficked victim or an independent sex worker is a very difficult task. Very little concrete ground truth (i.e., ads definitively known to be posted by a trafficker) exists in this space. In this work, we develop tools and techniques that can be used separately and in conjunction to group sex ads by their true owner (and not the claimed author in the ad). Specifically, we develop a machine learning classifier that uses stylometry to distinguish between ads posted by the same vs. different authors with 90% TPR and 1% FPR. We also design a linking technique that takes advantage of leakages from the Bitcoin mempool, blockchain and sex ad site, to link a subset of sex ads to Bitcoin public wallets and transactions. Finally, we demonstrate via a 4-week proof of concept using Backpage as the sex ad site, how an analyst can use these automated approaches to potentially find human traffickers.
Today it is common that people are users of more than one social networking site, so more and more people have multiple accounts on the web. It is challenging to identify the anonymous users on the web. A technique based on the friend relationship is used to identify the same user on multiple social networking sites. It is based on the concept that friend relationships cannot be faked by others. Based on that FRUI (Friend Relationship User Identification) algorithm is developed to match users on multiple sites. It calculates the matching degree for all users with similar screen names on multiple sites and users with the highest matching degree will be considered as the identical users. The proposed method is extended to find the fraud identities in social media sites including the bitcoin concept.
Spam and Phishing Detection
Authorship Attribution and Profiling
Advanced Steganography and Watermarking Techniques