A Zero-Knowledge Framework for Verifiable Semantic Explanations : An Intrusion Detection Case Study
Abstract
Machine learning-based intrusion detection systems can identify malicious network activity, but their predictions and explanations are typically accepted without verifying that they were derived from the same input. This thesis develops a public-model/private-input zero-knowledge framework for certifying a prediction and its semantic explanation while keeping the processed network-flow features private. The framework is instantiated through a TON_IoT intrusion detection case study in which 104 processed features are mapped into five semantic groups. Logistic Regression is used as the proof-compatible public model, while XGBoost provides a stronger plaintext performance baseline. The main technical contribution is an implementation-backed proof relation that jointly verifies Logistic Regression inference and an ordered top-3 semantic explanation from the same private input. Under a fixed training-mean reference, semantic-group Exact SHAP for the linear score reduces to a direct group-wise weighted sum, enabling its implementation in a Circom circuit and verification using Groth16. The quantized relation achieves more than 99.99% prediction agreement with the floating-point model, while ordered top-3 explanation agreement is approximately 93.8%. Valid proofs are accepted, whereas incorrect predictions, malformed rankings, and out-of-range inputs are rejected. The results demonstrate the feasibility of cryptographically binding a prediction and a semantic explanation under private tabular inputs. The implemented relation remains limited to a public linear model, fixed semantic groups, and an approved reference vector, and does not provide arbitrary-model explanation verification, model confidentiality, or production-ready provenance.
Community
0 commentsNo discussion yet
Be the first to share a question or observation.