TrustX: An Explainable and Cryptographically Verifiable Deep Learning Framework for Multimodal Manipulation Detection
Abstract
In this study, a sophisticated model that combines deep learning, cryptographic verification, and explainable artificial intelligence (XAI) is presented to address multimodal manipulation risks in digital media. The proposed system uses a Hierarchical Multimodal Transformer (HMT) to model hierarchical relationships among facial movement, audio tone, and textual semantics. The Contrastive Cross-Modality Alignment (CCMA) mechanism improves the ability to distinguish authentic from doctored material by leveraging cross-modal contrastive learning. An XAI Forensic Analyser provides interpretability by using Grad-CAM++, temporal attention mapping, and saliency sequence visualisation to trace a transparent decision. Moreover, the Zero-Knowledge Cryptographic Verifier (ZKCV) is used to validate the model’s outputs with tamper-proof libsnark cryptographic hashing. The hybrid system takes multimodal CNN, WaveNet and BERT encoders’ embeddings and attains a detection accuracy of about 90 per cent and 92 per cent on benchmark data. This architecture provides a sustainable, explainable, and verifiable basis for multimedia authenticity, enabling a consistent, reliable multimodal forensic detection system.
Community
0 commentsNo discussion yet
Be the first to share a question or observation.