Papers1 provider · 2 records
June 21, 2026· Open MIND
preprint
Open access

Behavioral Identity Is Not Model Identity — Why measuring how a model behaves is not the same as proving which model is computing

Abstract

A deployed AI system can be interrogated for its identity in several distinct ways, and the answers do not interchange. This note concerns one of them — which neural network is producing this output at inference time? — and a popular method for answering it: behavioral fingerprinting, which samples an endpoint under a fixed prompt battery and flags it when the output distribution shifts beyond a statistical threshold. The note argues that behavioral fingerprinting, while a legitimate and valuable instrument for one task, does not establish model identity. It develops two measured failure modes. First, a behavioral signature is not durable: ordinary continued training erases the behavioral provenance trace — more effectively, in fact, than an informed adversary trains directly to suppress it — so the same model after a benign fine-tune presents as behaviorally distinct and triggers a false alarm. Second, a behavioral signature is reproducible by a different model: knowledge distillation converges a substitute toward a target's behavioral template by construction, so a behavior-matched substitute passes the check and produces a false acceptance. Both failures follow from a single fact about the layering of neural identity — behavior is the transient layer, which transfers under distillation and washes out under benign training, while the structural layer (the geometry of internal computation during a forward pass) does neither. The two methods answer different questions and compose rather than compete: behavioral monitoring is a continuous, low-cost tripwire that flags something moved; structural verification is a deterministic resolver that answers is it still the enrolled model. A system that ships only the tripwire has shipped drift detection and labeled it identity. The note documents the structural layer's direct test against the failure mode that defeats behavioral methods — behavior-preserving substitution — and situates the argument alongside independent work on intrinsic parameter-level fingerprints and cryptographic verifiable inference, both of which bind identity to the model rather than infer it from outputs. This is a category statement, not a product comparison: no specific system or vendor is named, and the argument rests on published, reproducible measurements. The Neural Network Identity Series — Mathematical foundations, empirical validation, and governance frameworks for verifying which model is running Paper 1: The δ-Gene: Inference-Time Physical Unclonable Functions from Architecture-Invariant Output Geometry (DOI: 10.5281/zenodo.18704275) Paper 2: Template-Based Endpoint Verification via Logprob Order-Statistic Geometry (DOI: 10.5281/zenodo.18776711) Paper 3: The Geometry of Model Theft: Distillation Forensics, Adversarial Erasure, and the Illusion of Spoofing (DOI: 10.5281/zenodo.18818608) Paper 4: Provenance Generalization and Verification Scaling for Neural Network Forensics (DOI: 10.5281/zenodo.18872071) Paper 5: Beneath the Character: The Structural Identity of Neural Networks — Mathematical Evidence for a Non-Narrative Layer of AI Identity (DOI: 10.5281/zenodo.18907292) Paper 6: Which Model Is Running?: Structural Identity as a Prerequisite for Trustworthy Zero-Knowledge Machine Learning (DOI: 10.5281/zenodo.19008116) Paper 7: The Deformation Laws of Neural Identity (DOI: 10.5281/zenodo.19055966) Paper 8: What Counts as Proof? — Admissible Evidence for Neural Network Identity Claims (DOI: 10.5281/zenodo.19058540) Paper 9: Composable Model Identity — Formal Hardening of Structural Attestations in the Enterprise Identity Stack (DOI: 10.5281/zenodo.19099911) Paper 10:Where Identity Comes From: Path Sensitivity and Endpoint Underdetermination in Neural Network Training (DOI: 10.5281/zenodo.19118807) Paper 11: Post-Hoc Disclosure Is Not Runtime Proof: Model Identity at Frontier Scale (DOI: 10.5281/zenodo.19216634) Paper 12: Family-Dependent Response to Reasoning Distillation Across Structural and Functional Identity Layers (DOI: 10.5281/zenodo.19298857) Paper 13: Safety-Alignment Removal as a Model-Identity Failure — Structural Evidence from Published Weight-Level Mutation Checkpoints (DOI: 10.5281/zenodo.19383019) Technical Note: Agent Identity Is Not Model Identity (DOI: 10.5281/zenodo.19240883) Technical Note: Gap Invariance: Why PPP Measurements Are Domain-Independent by Construction (DOI: 10.5281/zenodo.19275524) Technical Note: Measured Model Substitution Under Valid Agent Credentials (DOI: 10.5281/zenodo.19342848) Technical Note: Artifact Identity Is Not Runtime Identity — Trustfall Lite and the Boundary of File-Level Model Verification (DOI: 10.5281/zenodo.20019127) Technical Note: Artifact Identity Is Not Runtime Identity — Trustfall Lite and the Boundary of File-Level Model Verification (DOI: 10.5281/zenodo.20019127) Technical Note:: The Disappearing Window — AI Logprob Access Withdrawal and the Structural Verifiability of Frontier Model Contracts (DOI: 10.5281/zenodo.20362098) Formal Verification Stack for Neural Network Structural Identity (IT-PUF Coq Proofs) (DOI: 10.5281/zenodo.18930621) Copyright (c) 2026 Anthony Ray Coslett / Fall Risk AI, LLC. All Rights Reserved. Confidential and Proprietary. Patent Pending (Applications 63/982,893, 63/990,487, 63/996,680, 64/003,244).

Community

0 comments
Use Connect Wallet in the navigation

No discussion yet

Be the first to share a question or observation.