We live in an increasingly digital world, where Smart Cities have become a reality. One of the characteristics that make these cities smart is their ability to gather information and act upon it, improving their citizens lives. In this work, we present our system, Dignitas. A blockchain-based reputation system that allows citizens of a Smart City to assess the truthiness of information posted by other citizens. This assessment is based on a bet that reporters make, and all of those who agreed with him, that puts their gathered reputation at stake. This use of Reputation as a currency is a novel idea that allowed us to build an anonymous system. Using blockchain we were able to have multiple authorities, working with each other to make the system secure and thus avoiding centralized schemes. Our work was focused on developing our idea, a proof of concept, and testing the viability of our new solution.
We develop a new Data-Driven Phasic Word Identification (DDPWI) methodology to determine which words matter as the bitcoin pricing dynamic changes from one phase to another. With Google search volumes as a baseline, we find that Reddit submissions are both correlated with Google and have a comparable relationship with a variety of bitcoin metrics, using Spearman's rho. Reddit provides complete access to the text of submissions. Rather than associating sentiment with market activity, we describe the DDPWI method for finding specific 'price dynamic' words associated with changes in the bitcoin pricing pattern through 2017 and 2018. We assess the significance of these changes using Wilcoxon Rank-Sum Tests with Bonferroni corrections. These price dynamic words are used to pull out associated words in the submissions thereby providing the context to their use. For example, the price dynamic word 'ban', which became significantly higher in frequency as prices fell, occurred in the context of both government regulation and internet companies banning cryptocurrency adverts. This approach could be used more generally to look at social media and discussion forums at a granular level identifying specific words that impact the metric under investigation rather than overall sentiment.
Public software repositories such as GitHub make transparent the development history of an open source software system. Source code commits, discussions about new features and bugs, and code reviews are stored and carefully attributed to the appropriate developers. However, sometimes governments may seek to analyze these repositories, to identify citizens who contribute to projects they disapprove of, such as those involving cryptography or social media. While developers who seek anonymity may contribute under assumed identities, their body of public work may be characteristic enough to betray who they really are. The ability to contribute anonymously to public bodies of knowledge is extremely important to the future of technological and intellectual freedoms. Just as in security hacking, the only way to protect vulnerable individuals is by demonstrating the means and strength of available attacks so that those concerned may know of the need and develop the means to protect themselves. \n \nIn this work, we present a method to de-anonymize source code contributors based on the authors' intrinsic programming style. First, we present a partial replication study wherein we attempt to de-anonymize a large number of entries into the Google Code Jam competition. We base our approach on Caliskan-Islam et al. 2015, but with modifications to the feature set and modelling strategy for scalability and feature-selection robustness. We did not achieve 0.98 F1 achieved in this prior work, but managed a still reasonable 0.71 F1 under identical experimental conditions, and a 0.88 F1 given more data from the same set. \n \nSecond, we present an exploratory study focused on de-anonymizing programmers who have contributed to a repository, using other commits from the same repository as training data. We train random-forest classifiers using programmer data collected from 37 medium to large open-source repositories. Given a choice between active developers in a project, we were able to correctly determine authorship of a given function about 75% of the time, without the use of identifying meta-data or comments. We were also able to correctly validate a contributor as the author of a questioned function with 80\\% recall and 65\\% precision. This exploratory study provides empirical support for our approach. \n \nFinally, we present the results of a similar, but more difficult study wherein we attempt de-anonymize a repository in the same manner, but without using the target repository as training data. To do this, we gather as much training data as possible from the repository's contributors through the Github API. We evaluate our technique over 3 repositories: Bitcoin, Ethereum (crypto-currencies) and TrinityCore (a game engine). Our results in this experiment starkly contrast our results in the intra-repository study showing accuracies of 35% for Bitcoin, 22% for Ethereum, and 21% for TrinityCore which had candidate set sizes of 6, 5, and 7 respectively. \n \nOur results indicate that we can do somewhat better than random guessing, even under difficult experimental conditions, but they also indicate some fundamental issues with the state of the art of Code Stylometry. In this work we present our methodology, results, and some comments on past empirical studies, the difficulties we faced, and likely hurdles for future work in the area.
The rise of ubiquitous deepfakes, misinformation, disinformation, propaganda and post-truth, often referred to as fake news, raises concerns over the role of Internet and social media in modern democratic societies. Due to its rapid and widespread diffusion, digital deception has not only an individual or societal cost (e.g., to hamper the integrity of elections), but it can lead to significant economic losses (e.g., to affect stock market performance) or to risks to national security. Blockchain and other Distributed Ledger Technologies (DLTs) guarantee the provenance, authenticity and traceability of data by providing a transparent, immutable and verifiable record of transactions while creating a peer-to-peer secure platform for storing and exchanging information. This overview aims to explore the potential of DLTs and blockchain to combat digital deception, reviewing initiatives that are currently under development and identifying their main current challenges. Moreover, some recommendations are enumerated to guide future researchers on issues that will have to be tackled to face fake news, disinformation and deepfakes, as an integral part of strengthening the resilience against cyber-threats on today's online media.
Adnan Qayyum, Junaid Qadir, Muhammad Umar Janjua, Falak Sher
In recent years, `fake news' has become a global issue that raises unprecedented challenges for human society and democracy. This problem has arisen due to the emergence of various concomitant phenomena such as (1) the digitization of human life and the ease of disseminating news through social networking applications (such as Facebook and WhatsApp); (2) the availability of `big data' that allows customization of news feeds and the creation of polarized so-called `filter-bubbles'; and (3) the rapid progress made by generative machine learning (ML) and deep learning (DL) algorithms in creating realistic-looking yet fake digital content (such as text, images, and videos). There is a crucial need to combat the rampant rise of fake news and disinformation. In this paper, we propose a high-level overview of a blockchain-based framework for fake news prevention and highlight the various design issues and consideration of such a blockchain-based framework for tackling fake news.
Trust in a smart city is fundamental to its transparency, the participation of its people in governance, entrepreneurial initiatives, trade, commerce and hence the growth of its economy. A city gets smart by transforming itself to a digital city, and a digital city runs on data, analytics, internet of things, artificial intelligence and machine learning. This inevitable transformation creates the fundamental need for trust. How does data remain sacrosanct and verifiable? How do people trust institutions? How do institutions trust each other? How do devices trust each other? This article explores blockchain as an essential layer of trust in a smart city. It explains the technology by drawing real-life examples to the ‘memory game’ that operates in an ecosystem of ‘trust and consensus.’ The article provides further insight into institutions that can be governed on blockchain through ‘smart contracts’ in a sovereign and human independent manner. The use cases of blockchain have been corroborated with examples of successful blockchain implementation. The value for blockchain, in general, and smart cities, in particular, has been presented across four categories: (a) the network effect on trust on society, governments and industries; (b) empowering the individual and strengthening the economy; (c) the liquid economy and (d) the shareable economy. Given the current topology of technology innovations, there is no solution better than blockchain that embodies trust. It is a hope and expectation that this article will help smart city planners, developers, architects and thinkers implement blockchain as the embodiment of trust in smart cities that are increasingly becoming digital.
Die Unsicherheit über den intrinsischen Wert von Kryptowährungen und der nachgewiesene Einfluss der Aufmerksamkeit an dem Marktwert verschiedener Vermögenswerte haben uns veranlasst, den Einfluss der Aufmerksamkeit auf den Marktwert von Kryptowährungen zu untersuchen. Als Aufmerksamkeitsindikator haben wir das Volumen von Google-Suchen zu bestimmten der Suchwörter Suchbegriffe genutzt das Suchvolumen von Google-Suchmenge, die auf Schlüsselwörtern basieren, die sich auf unseren Satz von Kryptowährungen einer sehr genauen Granularität beziehen. Unter Verwendung von ARMA und VECM haben wir getestet, ob die Google-Suchmenge die Vorhersage für Kryptowährungspreisentwicklung im Zeitrahmen von 15 Minuten bis zu einem Tag verbessert. Anschließend haben wir den Handel mit dieser Out-of-Sample Prognose simuliert und kamen zu dem Schluss, dass im Fall von häufigem Handel ohne Gebühren, einfache, univariate, autoregressive Modelle besser Ergebnisse produzieren. Unter Vernachlässigung von Gebühren jedoch, verbessert sich durch die Einbeziehung der Variablen für das Google-Suchvolumen das Handelsergebnisse, insbesondere bei stündlichen und täglichen Frequenzen. Unter Verwendung solcher Frequenzen übertraf das Modell univariate Modelle sowie das Wachstum der zugrunde liegenden Vermögenswerte.
Mehrnoosh Mirtaheri, Sami Abu-El-Haija, Fred Morstatter, Greg Ver Steeg · 5 authors
Interest surrounding cryptocurrencies, digital or virtual currencies that are used as a medium for financial transactions, has grown tremendously in recent years. The anonymity surrounding these currencies makes investors particularly susceptible to fraud---such as ``pump and dump'' scams---where the goal is to artificially inflate the perceived worth of a currency, luring victims into investing before the fraudsters can sell their holdings. Because of the speed and relative anonymity offered by social platforms such as Twitter and Telegram, social media has become a preferred platform for scammers who wish to spread false hype about the cryptocurrency they are trying to pump. In this work we propose and evaluate a computational approach that can automatically identify pump and dump scams as they unfold by combining information across social media platforms. We also develop a multi-modal approach for predicting whether a particular pump attempt will succeed or not. Finally, we analyze the prevalence of bots in cryptocurrency related tweets, and observe a significant increase in bot activity during the pump attempts.
Rumors and misleading information detection and prevention still represent a big challenge against social network developers and researchers. Since newsworthy information propagation is a traditional behavior of most of the users in social media, then verifying information credibility and reliability is indeed a vital security requirement for social network platforms. Due to its immutability, security, tamper-proof and P2P design, Blockchain as a powerful technology can provide a magical solution to overcome this challenge. This Paper introduces a novel blockchain approach called Proof of Credibility (PoC) for detecting fake news and blocking its propagation in social networks. The functionality of the PoC protocol has been simulated on two datasets of newsworthy tweets collected from different news sources on Twitter. The results clarified a satisfying performance and efficiency of the proposed approach in detecting rumors and blocking its propagation.
Johannes Beck, Roberta Huang, David Lindner, Tian Guo · 7 authors
The ability to track and monitor relevant and important news in real-time is of crucial interest in multiple industrial sectors. In this work, we focus on the set of cryptocurrency news, which recently became of emerging interest to the general and financial audience. In order to track relevant news in real-time, we (i) match news from the web with tweets from social media, (ii) track their intraday tweet activity and (iii) explore different machine learning models for predicting the number of the article mentions on Twitter within the first 24 hours after its publication. We compare several machine learning models, such as linear extrapolation, linear and random forest autoregressive models, and a sequence-to-sequence neural network. We find that the random forest autoregressive model behaves comparably to more complex models in the majority of tasks.
In early 2018 prices peaked at USD 20,000 and, almost two years later, we still continue debating if cryptocurrencies can actually become a currency for the everyday life or not. From the economic point of view, and playing in the field of behavioral finance, this paper analyses the relation between prices and the search interest on Bitcoin since 2014. We questioned the forecasting ability of Google Trends for the behavior of price by performing linear and nonlinear dependency tests, and exploring performance of ARIMA and Neural Network models enhanced with this social sentiment indicator. Our analyses and models are founded upon a set of statistical properties common to financial returns that we establish for Bitcoin, Ethereum, Ripple and Litecoin.
Eaman Jahani, P. M. Krafft, Yoshihiko Suhara, Esteban Moro · 5 authors
Participants in cryptocurrency markets are in constant communication with each other about the latest coins and news releases. Do these conversations build hype through the contagiousness of excitement, help the community process information, or play some other role? Using a novel dataset from a major cryptocurrency forum, we conduct an exploratory study of the characteristics of online discussion around cryptocurrencies. Through a regression analysis, we find that coins with more information available and higher levels of technical innovation are associated with higher quality discussion. People who talk about "serious" coins tend to participate in discussion displaying signatures of collective intelligence and information processing, while people who talk about "less serious" coins tend to display signatures of hype and naïvety. Interviews with experienced forum members also confirm these quantitative findings. These results highlight the varied roles of discussion in the cryptocurrency ecosystem and suggest that discussion of serious coins may be oriented towards earnest, perhaps more accurate, attempts at discovering which coins are likely to succeed.
May 28, 2017·CEUR Workshop Proceedings, Vol-2104: Proceedings of the 14th International Conference on ICT in Education, Research and Industrial Applications. Integration, Harmonization and Knowledge Transfer. Volume II: Workshops
The paper investigates the Bitcoin exchange rate response to the dai- ly Twitter data. Sentiment score is computed for the number of obtained tweets. The prediction accuracy for the Bitcoin exchange rate employing the sentiment score reveals the influence of the Twitter social network on the news diffusion and target exchange rate volatility. We used the historical data on the Bitcoin exchange rate and the daily sentiment score of the pertinent tweets to forecast the direction of change for the Bitcoin. The results show better performance of the developed forecasting method with both historical data on the exchange rate and the sentiment score than using only the exchange rate data as an input.
The Internet pervades modern life, offering up opportunities to connect, inform and be informed. As the range and number of sources for information online explode, how people select and interpret information has become a pertinent area for study, not least in light of the prevalence of fake-news. People are well known to act upon information they believe to be trustworthy and where the decision to act incurs risk, an inability to accurately select and assess the credibility of information presents a challenge. Bitcoin, the nascent crypto-currency, presents a domain within which profound financial risk abounds. Even for those armed with experience and knowledge there are numerous challenges to assessing risk, especially as sources of Bitcoin information can be observed to be partisan and of questionable accuracy. Within the domain of bitcoin speculation, this thesis asks the central research question of: are people able to select and correctly evaluate information they might rely upon to make decisions? In addressing this research question, this thesis offers - through the application of a psychological model of informational trust to bitcoin speculators - two fundamental contributions: Firstly, that these users are able to identify relevant news without a reliance upon confirmation bias. Secondly, that a notable percentage of users are not evaluating the credibility of online news by expertly interpreting the fundamentals of information but, rather deferring their trust to either the source news website or a more broad trust of information on the Internet. For these users, chance or luck may mean that they are basing their decisions upon factually accurate news. But this is a position which makes them particularly vulnerable to fake-news where it is spread via sources which they might trust. This position of susceptibility provides evidence to support further security research of both the prevalence of, and counter-measures for fake-news.
Bitcoin is difficult to categorize and indeed has been associated with 112 different labels in the British media (e.g., “private money,” “commodity”) – most of which poorly describe bitcoin. Specifically, our analyses of 674 media articles, focusing on the relationship between labeling and categorization, identify classification inconsistencies at three levels: within clusters of labels, between labels and categories, and between category attributes. These inconsistencies hamper categorization based on attribute similarity, audience goals, and causal models, respectively. We identify four factors that nurture this categorical anarchy and conclude with a call for research on the socioeconomic revolution heralded by blockchain technology.
Digital currencies represent a new method for exchange and investment that differs strongly from any other fiat money seen throughout history. A digital currency makes it possible to perform all financial transactions without the intervention of a third party to act as an arbiter of verification; payments can be made between two people with degrees of anonymity, across continents, at any denomination, and without any transaction fees going to a central authority. The most successful example of this is Bitcoin, introduced in 2008, which has experienced a recent boom of popularity, media attention, and investment. With this surge of attention, we became interested in finding out how people both inside and outside the Bitcoin community perceive Bitcoin -- what do they think of it, how do they feel, and how knowledgeable they are. Towards this end, we conducted the first interview study (N = 20) with participants to discuss Bitcoin and other related financial topics. Some of our major findings include: not understanding how Bitcoin works is not a barrier for entry, although non-user participants claim it would be for them and that user participants are in a state of cognitive dissonance concerning the role of governments in the system. Our findings, overall, contribute to knowledge concerning Bitcoin and attitudes towards digital currencies in general.
Okay, kids, you can come back in; Uncle Mole is in a better mood now. If you're just joining us, you should know that I was in a very foul temper, because I'd only just found out that a prominent scientist, whose work I'd valued, had been exposed as a fake. His lovely work wasn't simply flawed, it was made up. My whole faith in science has been shaken, and I want to fix it. Instead, I've watched TV.I've just seen one of my favorite episodes of the old Twilight Zone. In `It's a Good Life' we meet a mind-reading, omnipotent monster who is terrorizing a small town. The monster is a six-year-old body played by Billy Muni, who would later gain quasi-immortality as the youngest member of the Robinson family on Lost in Space (as in, `Danger, Will Robinson!'). But in this Twilight Zone story by Jerome Bixby, he is a small boy who can do anything, and when he is displeased, he can transform townspeople into horrors, or make them disappear altogether, by `sending them to the cornfield'. So I imagined doing this to our scientific fraud, and I felt better. Maybe this wasn't the point of the story, but it made me feel better.The problem of scientific fraud isn't new, but it seems as though our efforts to eradicate it have not worked. We get tougher, but the fakers just get better at faking. Perhaps we need to further tighten security - remove your shoes and laptops prior to submission...There are two views of this problem, and the one we take will dictate what we should do. The `tip of the iceberg' position says that the fraud that has been exposed represents only the tiniest bit of a problem that is rotting science from the inside. Some have advocated a zero-tolerance policy, and have taken it on themselves to act, vigilante fashion, to publicize any discrepancies they find in publications, demanding satisfaction. The standard operating procedure here seems to be to contact the journal and the community, via emails for example, intimating that every questionable figure is evidence of fakery - mistakes cannot be tolerated. I know of one investigator who is being hounded to explain two identical images in a paper, which he asserts is a post-proof printing error (and was immediately corrected) but he can't prove it. But the vigilantes contend that zero tolerance demands that everyone subject themselves to a `trust no-one' process in the hope that we'll weed out the worst offenders.Don't get me wrong, I do think that there is a lot of fudging in the literature. My old Oxford English Dictionary defines fudge in this context as “to fit together or adjust in a clumsy, makeshift or dishonest manner”. (The most romantic etiology traces this to one Captain Fudge, c1664, a.k.a. Lying Fudge, who was probably a real person, although he may have made himself up.) In science, fudging data can be elimination of compelling results that don't fit the hypothesis (which may be for perfectly valid reasons or not) or adjusting the results, say, when molecular weight markers seem off. It occurs because researchers are under tremendous pressure to publish on a timetable - the need to publish any work that has used up time and resources, however questionable the conclusions. Journals promote this problem by demanding additional results that are conditions for publication, usually on even shorter timetables. Fudging seems inevitable. I am not forgiving it; I'm only saying why I think it happens. But when it goes too far it becomes fakery, and it is unforgivable. The tip of the iceberg view is that much of what we see is not simply fudged; it is faked.The alternative view, to which I subscribe, is that true fraud is exceedingly rare. Mistakes, misinter - pretations, and wishful thinking are more common - and problematic - but I think we can deal with them. But outright fabrication is rare enough to be news.Can I prove this second view? No. But I can demonstrate that it is a useful and profitable position to take. And the demonstration points to a route to the solution, not only for fraud, but also for errors and other problems.Unless you live in a cave and, for that matter, a cave without an internet connection, you know about eBay, the massively successful online auction system. Anyone can buy or sell anything on eBay (including, apparently, fabulously expensive grilled-cheese sandwiches) and can do so with a remarkable level of confidence. It is based on a seemingly naive, but ultimately profound, precept: most people are honest. This is backed up by a readily accessible rating system, where buyers and sellers provide feedback on their transactions, thereby exposing problems if and when they arise. The system is largely transparent: those who lie are quickly flamed, and anyone who gives inordinate numbers of negative comments is discredited. It is freewheeling, but for the most part it is wildly successful.Once, when science was conducted by an elite, feedback occurred in the literature and at meetings. This still happens, but in a manner that is difficult to assess unless one is in the center of the action (again, one of the elite). High-impact journals have no interest in publishing work that refutes other work, regardless of the rigor of the refutation, and the group psychology among researchers translates this into `high impact = true, low impact = less true'. Even when we think that everyone knows that a particular finding is flawed, one only has to take a stroll into a related but different venue, such as the department upstairs, to find that others who might be peripheral to the field can evince surprise at our suspicions. We need a feedback system that everyone can access.I propose that we take a cue from eBay. Link a system to PubMed, for example, by which we can identify a paper and offer feedback (“we repeated this finding, but couldn't reproduce that one” or “this result may be an artifact for the following reasons”). It must be transparent - commentators are registered and their identities known, and we can similarly access their other reviews. Vigilantes who only find fault will find their comments of less value than those from reviewers who are balanced in their views. And, of course, the authors will be able to respond to criticism if it is especially important. We will have a way to evaluate the experiences of the community, far beyond a paper's `impact', which is more likely to reflect the extent to which a finding is easy to mention. I think we'll gain confidence in the literature, we'll expose fudges, and we'll find very little fraud.For this to work, however, we need a fundamental change in the community of scientists. We need to realize that making mistakes is common and that honesty requires that errors be owned up to. We need to reward, not punish, those individuals who can say that they got it wrong. It happens all the time, and we pretend it doesn't, and, as a consequence, the fudging goes on. We can make it stop, but only if we take away the pressure not to admit to it.But what of the real monsters who are out there? The ones who simply make it up? Such allegations are serious and must be dealt with by informed investigation, as we do now, and evidence of genuine misconduct must come from close to home (fellow researchers with intimate knowledge of the lab and methods). We can deal with this, and only open such investigations when the evidence is overwhelming. But why do they do it, this outright fakery that is so antithetical to the entire process of scientific inquiry? I think I know, because I've just been watching the Twilight Zone.The monster in `It's a Good Life' can do anything, knows everything and, because he is only a child, has no goals but his own desires. He understands only that whatever he wants to happen, happens. When a bright, young, and very ambitious scientist begins his or her career, one of two things occur early on. They can chance on a set of ideas that happen to be correct, and their experiments flow effortlessly towards a happy conclusion. And if this is an important conclusion, rewards come quickly. Or, alternatively, they can be wrong, and they learn at this formative stage that no matter how wonderful an idea may be, and how much they need it to be true, it can still be wrong. This is an extremely important lesson that our first, lucky researcher may not learn, unless, of course, we teach them. The successful student guesses again, and again may be right - more rewards. By the time they come up against something that they cherish that turns out to be mistaken, they may have already become our monster-their idea, their need to be right, exceeds all other goals. It doesn't happen all at once - but someone who is always right begins to believe that they are special, and does not realize that luck is a major factor in all of this. So they make it right. They have made the leap to quasi-omnipotence. I once met a monster like this, and it was truly frightening.'It's a Good Life' was remade, years later, as a vignette in The Twilight Zone Movie. In the rewritten work, the ending was changed: a teacher takes on the task of educating the monster/child, who is desperate for guidance. And this, of course, is what we need to do with our most gifted, lucky young scientists. We have to teach them that ideas are frequently wrong, and this is fundamental to science. And we have to stop stressing that being right brings rewards, while being wrong brings despair. Let's stop giving awards for best poster, best thesis, best student. Science is a reward. We don't need more monsters. Let them be wrong sometimes.Otherwise, I'll want to send you to the cornfield.