Trustworthy and Privacy-Preserving Perceptual Hashing with Zero-Knowledge Proofs for Client-Side Content Scanning
Abstract
In client-server applications such as copyright protection and content moderation, learning-based perceptual hashing compresses images into compact binary codes whose Hamming distances approximate perceptual similarity. Clients then transmit these codes to servers for comparison. However, this approach faces dual challenges: algorithmically, how to effectively balance robustness and discriminability while mitigating bit imbalance issues; protocol-wise, transmitting these hashes compromises client privacy through content inference and cross-platform user tracking. To address these challenges, we propose a trustworthy privacy-preserving framework that integrates deep hashing with zero-knowledge proofs. The framework comprises: (1) A robust deep hashing module that generates discriminative binary codes by optimizing a composite objective function composed of the Angular Triplet and quantization losses, while using a multi-scale strategy to correct bit imbalance. (2) A privacy-preserving similarity comparison protocol based on Sumcheck and Logarithmic Lookup, which enables clients to locally prove batch Hamming distance relationships against public dataset entries without disclosing their hash values. We conducted comprehensive evaluations to demonstrate the practicality and efficiency of our design compared to existing schemes. Source code is available at https://github.com/mengdehong/zkph.
Community
0 commentsNo discussion yet
Be the first to share a question or observation.