Speaker anonymization protects against speaker identity inference, yet third parties cannot verify that released speech is authenticated and anonymized as predefined without revealing the original. We propose Verifiable Speaker Anonymization (VSA), a paradigm that enables public verification that a predefined anonymization has been applied while the original remains hidden. We instantiate this paradigm as ZK-VSA using zero-knowledge succinct non-interactive arguments of knowledge (ZK-SNARKs): we encode phase vocoder with time-scale modification (PV-TSM) as arithmetic constraints suitable for succinct proofs, complemented by SNARK-friendly phase handling, and integrate cryptographic commitments with digital signatures for authentication. We evaluate ZK-VSA on LibriSpeech, using automatic speech recognition (ASR) for intelligibility and automatic speaker verification (ASV) for anonymity. Our proof-constrained anonymization closely matches floating-point PV-TSM, while proofs add only a slight overhead and verify in milliseconds. These results demonstrate the practicality of VSA and open a path to proof-based guarantees for broader speech transformations.
Audio piracy detection is increasingly complex in decentralised distribution settings, where mainstream approaches fail to ensure robustness, verifiability, or computational efficiency. Conventional Digital Rights Management (DRM) systems mainly enforce licensed access, but once content is copied or redistributed outside their control they offer little protection. Classical fingerprinting approaches such as MFCC based hashes can detect near-exact duplicates, yet they often fail under signal edits like pitch shifting, time stretching or equalisation. Deep learning embeddings improve robustness but demand heavy computation and centralised resources, making them less suitable for edge or decentralised deployments. These limitations call for a solution that is both edit resilient and verifiable. We propose HashWave, a blockchain-integrated perceptual hashing framework that combines robust audio fingerprinting with tamper-proof verification. The system fuses MFCC, chroma and chroma CENS, CQT, spectral contrast, and lightweight tempo/energy cues, applying operation-aware weighting via [Formula: see text] and constrained DTW for time-scale edits. Evaluated across GTZAN, FMA-A Dataset for Music Analysis, and MUSAN (SLR17) with over twenty signal-processing transformations, HashWave achieves AUC 0.957 and TPR@1%FPR 0.952, outperforming MFCC-only baselines and approaching deep embeddings at lower CPU cost. The blockchain layer, built on Ethereum and IPFS, ensures decentralised hash storage, duplication control, and verifiable authorship with average upload and contract execution times of 0.017 s and 0.044 s. Together, these results establish HashWave as a practical, scalable, and secure framework for piracy detection across streaming, podcasting, and Web3 ecosystems.
Open access
Advanced Steganography and Watermarking Techniques
Recent years have seen great improvements in zero-knowledge proofs (ZKPs). Among them, zero-knowledge SNARKs are notable for their compact and efficiently-verifiable proofs, but suffer from high prover costs. Wu et al. (Usenix Security 2018) proposed to distribute the proving task across multiple machines, and achieved significant improvements in proving time. However, existing distributed ZKP systems still have quasi-linear prover cost, and may incur a communication cost that is linear in circuit size. In this paper, we introduce HyperPianist. Inspired by the state-of-the-art distributed ZKP system Pianist (Liu et al., S&P 2024) and the multivariate proof system HyperPlonk (Chen et al., EUROCRYPT 2023), we design a distributed multivariate polynomial interactive oracle proof (PIOP) system with a linear-time prover cost and logarithmic communication cost. Unlike Pianist, HyperPianist incurs no extra overhead in prover time or communication when applied to general (non-data-parallel) circuits. To instantiate the PIOP system, we adapt two additively-homomorphic multivariate polynomial commitment schemes, multivariate KZG (Papamanthou et al., TCC 2013) and Dory (Lee et al., TCC 2021), into the distributed setting, and get HyperPianistKand HyperPianistDrespectively. Both systems have linear prover complexity and logarithmic communication cost; furthermore, HyperPianistDrequires no trusted setup. We also propose HyperPianist+, incorporating an optimized lookup argument based on Lasso (Setty et al., EUROCRYPT 2024) with lower prover cost. Experiments demonstrate HyperPianistKand HyperPianistDachieve speedups of 63.1x and 40.2x over HyperPlonk with 32 distributed machines. Compared to Pianist, HyperPianistKcan be 2.9x and 4.6x as fast and HyperPianistDcan be 2.4x and 3.8x as fast, on vanilla gates and custom gates respectively. With layered circuits, HyperPianistKis up to 5.9x as fast on custom gates, and HyperPianistDachieves a 4.7x speedup.
In recent years, increased access to sophisticated tools in the domains of artificial intelligence (AI) and decentralized computation (in particular blockchain-enabled technologies) is leading to significant developments in the landscape of digital art. In the creative experiments of artists, designers, and technologists, such developments are evident, for example, in a focus on generative processes (e.g., AI-enabled text-to-image generation) and on the production of unique digital artifacts (such as blockchain-enabled non-fungible tokens, or NFTs). For now, the bulk of creative experimentation and theoretical reflection appears to have taken place in an occularcentric mode, and with a primary focus on visual, non-time-based artforms. Drawing on this existing discourse, in this chapter I will begin to explore some opportunities that emerging AI and blockchain technologies represent for new compositional practices, performance, collaboration, and distribution of music and sound-based aesthetic artifacts. Throughout this discussion, the underlying focus is on the shifting contours of creative agency effected by AI and blockchain technologies. With this focus in mind, key concerns include the following: How can creative agency be encoded in AI-augmented and blockchain-enabled musical objects? What are the implications of this ‘becoming-agential’ for questions related to authorship and ownership? It will not be possible to provide conclusive answers to these questions here. Instead, my aim in this chapter is to stake the relevance and importance of the questions raised by outlining underlying concerns and perspectives. In the following sections, this is done first by offering detailed contextualization, and subsequently by discussing an ongoing multimodal art project that is of great relevance to the concerns outlined above.
Blockchain and Web3 technologies were originally understood in a narrowly financial context, associated with cryptocurrencies such as Bitcoin. Non-fungible tokens, or NFTs, have put music (and other creative economy use cases such as art and gaming) center stage. Yet there remains a widespread misconception that speculative investments in cryptocurrencies and NFTs are the be-all and end-all of Web3. This chapter examines the potential impact of blockchain technology on the music industry from a more holistic perspective. Specifically, its contribution lies in proposing four lenses through which we can view Web3 music: financial, social, environmental and experiential. By examining the value of tokens – fungible as well as non-fungible – in terms that go beyond short-term price fluctuations, I hope to deepen our understanding of cultural value in Web3 and beyond.
In this work, we present SONAR, a web-based tool for multimodal exploration of Non-Fungible Token (NFT) inspiration networks. SONAR is conceived to support both creators and traders in the emerging Web3 by providing an interactive visualization of the inspiration-driven connections between NFTs, at both individual level and collection level. SONAR can hence be useful to identify new investment opportunities as well as anomalous inspirations. To demonstrate SONAR's capabilities, we present an application to the largest and most representative dataset concerning the NFT landscape to date, showing how our proposed tool can scale and ensure high-level user experience up to millions of edges.
Non-Fungible Tokens (NFTs) represent deeds of ownership, based on blockchain technologies and smart contracts, of unique crypto assets on digital art forms (e.g., artworks or collectibles). In the spotlight after skyrocketing in 2021, NFTs have attracted the attention of crypto enthusiasts and investors intent on placing promising investments in this profitable market. However, the NFT financial performance prediction has not been widely explored to date. In this work, we address the above problem based on the hypothesis that NFT images and their textual descriptions are essential proxies to predict the NFT selling prices. To this purpose, we propose MERLIN, a novel multimodal deep learning framework designed to train Transformer-based language and visual models, along with graph neural network models, on collections of NFTs' images and texts. A key aspect in MERLIN is its independence on financial features, as it exploits only the primary data a user interested in NFT trading would like to deal with, i.e., NFT images and textual descriptions. By learning dense representations of such data, a price-category classification task is performed by MERLIN models, which can also be tuned according to user preferences in the inference phase to mimic different risk-return investment profiles. Experimental evaluation on a publicly available dataset has shown that MERLIN models achieve significant performances according to several financial assessment criteria, fostering profitable investments, and also beating baseline machine-learning classifiers based on financial features.
Preventive measures to stop copyright infringement are yet to be implemented on current decentralized music-sharing platforms. There is no mechanism to reject modified audio before they go online, so some decentralized music platforms become places full of pirated audio files. To address this problem, a perceptual hash-based audio detection method for copyright protection in decentralized music sharing was proposed. Chromaprint, an open-source audio fingerprint program, generates a hash value to detect copyright infringement. To assess the perceptual hash technique’s robustness, Chromaprint generates a hash value from an audio file that can be modified with several signal processing attacks. The results of the detection system show that Chromaprint is very effective at spotting copyright infringement, with an average match rate of 92.64%. Deployed using a public Ethereum blockchain test network, the execution time from hashing to uploading to the IPFS distributed storage is only 616.3 ms.
Advanced Steganography and Watermarking Techniques
The purposes are to recognize and classify different music characteristics and strengthen the copyright protection system for original digital music in the big data era. Deep learning (DL) and blockchain technology are applied and researched herein. Based on CNN (Convolutional Neural Network), a music recognition method combined with hashing learning is proposed. The error generated when outputting the binary hash code is considered, and the semantic similarity of the hash code is ensured. Besides, the application of blockchain technology in the current intellectual property protection in original music is discussed. According to digital music property rights protection needs, the system is divided into modules, and its functions are designed. The system ensures its various functions by applying the application protocol designed in the Algor and network. In the experiments, the MagnaTagATune dataset is selected to verify the performance of the proposed CRNNH (Convolutional Recurrent Neural Network Hashing) algorithm. The algorithm shows the best music recognition performance under different bit numbers. When the number of connections is about 100, the QPS value of the blockchain-based music property rights protection system can be stabilized at about 20,000. At any number of threads, the system pressure will increase dramatically with the increase in the number of analog connections. The music recognition algorithm based on DL and hash method discussed is of great significance in improving the classification accuracy of music recognition. The application of blockchain technology in the copyright protection platform of original music works can protect the copyright of digital music and ensure the operation performance of the system.
Open access
Music and Audio Processing
Diverse Musicological Studies
Generative Adversarial Networks and Image Synthesis
Marcello Messina, Marcos Célio Filho, Carlos Mario Gómez Mejía, Damián Keller · 6 authors
We introduce IoMuSt — the Internet of Musical Stuff: a proposal to recalibrate the Internet of Musical Things in the light of the current reification of digital creative resources, epitomised by the Non-Fungible Tokens frenzy. As opposed to marketable “things”, “stuff ” is fluid, malleable, unfixable and pecuniarily irrelevant. Hence, stuff is good raw material for sustainable ubimus creative ecosystems
In recent years, the music industry has advanced a lot in technology and provided listeners with easy access to music via digital platforms. The internet has re-organized the value chain and accelerated processes with the evolution of streaming services. Associated issues include copyright protection and royalty distribution. By looking at the music business’s appropriation organization, it very well may be seen that the internet web-based stages have empowered buyers to access music easily, but however has introduced a level of mediation among specialists and the users, creating a wasteful eminence installment framework. The motivation behind this chapter is to identify blockchain applications that would draw in the disintermediation of the business, enabling artists to gain additional benefits from their music. This chapter details the impact of blockchain on the music business and various innovations to make this music appropriation framework safer and more reliable.
Blockchain, which started as a cryptocurrency bitcoin, has recently been applied to various fields such as finance, distribution, and public services beyond the category of cryptocurrency. In this paper, we propose a music distribution model to protect music copyrights and rights holders' rights based on blockchain and smart contract technology. By designing and implementing a blockchain-based music distribution model, it organizes music assets into blocks and distributes them among blockchain participating nodes, providing integrity, confidentiality and non-repudiation of assets, and a single point of blockchain advantage. Single Point of Failure (SPOF) problem can be minimized. Blockchain allows musicians to easily approve and manage their music copyright with distributed ledger technology. Rights holders can automatically and immediately receive royalties from the music industry, even if no broker is involved in the distribution process. By distributing and managing music using the proposed model, we can provide all transaction information and related tasks in the music market safely and transparently.
Music and Audio Processing
Digital Rights Management and Security
Advanced Steganography and Watermarking Techniques
To explore the automatic computer composition, investigate the copyright protection and management of digital music, and expand the application of deep learning and blockchain technologies in the generation of digital music works, piano composition was taken as a sample. First, through the elaboration of the neural network methods based on deep learning, the Recurrent Neural Network (RNN), Long-Short-Term Memory (LSTM), and Gated Recurrent Unit (GRU) networks were introduced, and the deep learning-based GRU-RNN automatic composition model was constructed. Second, the blockchain technology was analyzed and expressed, and the problems in the traditional copyright protection and management of digital music were analyzed. The three aspects, i.e., ownership, right of use, and right protection, were fully considered, and the blockchain technology was integrated into the copyright protection and management of digital music. Finally, the manual analysis evaluation and pause analysis were selected as the indicators to analyze and characterize the music composition quality of the GRU-RNN model, as well as analyzing the development of the digital music market integrated with blockchain technology. The results show that the GRU-RNN model shows satisfactory effects in manual analysis evaluation or in the pause analysis of the passage. The deep learning method has great potential for application in automatic computer composition of digital music; the integration of blockchain technology has played a promotive role in the expansion and popularization of the digital music market. However, in the meantime, it still faces some technical and policy challenges. The results have a positive effect on promoting the development and application of deep learning methods and blockchain technology in digital music.
[Associate Editor's note: In the September 2008 edition of Voice Research and Technology, I introduced a topic called Quantifying Tessitura in a Song. It was more of a teaser than a complete work. My good friend John Nix completed a real study, including dosimetry on a singer who performed a Mozart composition (different from Il mio tesoro intanto, which I used for my test case)-Ingo R. Titze.]INTRODUCTIONSelecting appropriate repertoire is a primary responsibility for a singing teacher. Astute repertoire selection can help address the technical and musical development of a singer, as well as advance a singer's career in the case of those who are working professionals. Traditionally, repertoire is assigned on an individual basis by carefully considering the singer's age, gender, technical challenges, personality, musicianship, and developmental level, then cross referencing that singer assessment with an evaluation of the vocal, musical, linguistic, and expressive challenges of potential repertoire.1 A number of print2 and web3 resources exist to assist teachers in choosing potential new literature.In the past, the selection process has depended upon the acquired ability of the singing teacher both to evaluate his or her students and to make judgments about vocal repertoire based upon personal knowledge and close examination of the literature. In recent years, however, the voice range profile (VRP) and voice dosimetry have begun to show promise as objective tools that could assist teachers with the selection process.Several articles have suggested ways in which the VRP could be used or improved upon for determining the voice classification of performers or for guiding repertoire choices. Emerich, Titze, Svec, Popolo, and Logan compared laboratory VRPs for eight professional actors with speech range profiles of the actors during a dramatic scene recorded in a laboratory and on stage in performance.4 At times, the actors exceeded their VRP thresholds when in either of the performance conditions. The study highlighted both the utility of the VRP when comparing laboratory values with actual performance voice usage and the shortcomings of the VRP in predicting what repertoire might exceed a performer's capabilities given the emotional content of live performance. Lamarche, Ternstrom, and Hertegârd also looked at the VRP as compared to the performance of repertoire, combining a commercial VRP program with a response button that subjects pressed during moments of difficult production.5 The subjects, all female professionally trained singers, performed three tasks, including the singing of their best aria, using performance-acceptable quality. Subjects also rated how well the button pressings reflected their experiences and usual vocal challenges; this rating included a viewing of a visual graph of the button responses. The investigators' coupling of objective data from the VRP tasks with (a) the singer's self-assessment captured mid-task and (b) the postperformance evaluation of the data and the responses shows significant promise as a means for enhancing singing teachers' guidance of their students. Lamarche, Ternstrom, and Pabon continued testing improvements of the VRP for singers by examining whether VRPs for singers should be limited to physiological measures (without considering performance quality elements), or whether performance abilities should also be measured in some fashion.6 After testing thirty female subjects, all professionally active singers, they concluded that examining the voice as it is used in performance has clinical importance, and that both types of VRPs should be used for assessing singers. Finally, Herbst, Duus, Jers, and Svec compared the maximum phonation frequency range (MPFR) of amateur choir singers (as gathered during a VRP) with the required pitch range (RPR) for their chosen voice part within a choir, based upon part ranges specified in a commonly used music reference. …