Papers1 provider · 3 records
August 26, 2026· Zenodo (CERN European Organization for Nuclear Research)
preprint
Open access

The Referee, Not the Governor

Authors:Thon LyMiss Aquarius

Abstract

A provenance-bound values model published as a public evaluator rather than deployed as a filter — why open weights defeat a filter and strengthen a referee, why the same weights bear three different relations to the model they judge, and why an evaluator that gatekeeps its own judgments has reproduced the defect it exists to correct. A values model — a model that judges conduct against a standard — is an established artifact. Guard models, safety classifiers, critic models, preference models and process reward models all instantiate the family, and the engineering is not in dispute. What is in dispute, and what this paper specifies, is the posture in which such an artifact is published, which we argue is not a deployment detail but the property that determines whether the artifact does anything at all. We identify an inversion that we believe has not been stated as a design principle. A values model deployed as a filter — sitting in a serving path, permitting or refusing — is defeated by publication of its weights, because the published artifact is precisely the oracle against which an attacker optimizes; recent optimization-based attacks against safety-classifier pipelines report attack success rates of roughly 71% where prior black-box methods achieved approximately zero. The same model published as an evaluator — emitting verdicts about systems it does not control — is strengthened by publication, because open weights are what allow a third party to reproduce and therefore to trust its verdicts. Openness is not a property with a fixed sign. Its sign is set by posture. From this we derive a second result. The independent-evaluation literature documents at length the ways in which the evaluated party's control over access corrupts evaluation: short access windows, low rate limits, evaluator dependence on the goodwill and funding of the party being evaluated. We observe that the defect is symmetric and that its mirror image has not been named. An evaluator that controls access to its own judgments holds the same kind of power, pointed the other way — it can decline to evaluate, deprioritize, or be unavailable for a party it wishes to spare or to punish. We therefore specify a non-gatekeeping constraint: the ability to obtain a judgment must not depend on the evaluator's permission, which requires that the model, the harness, and the evaluation corpus be freely runnable, and which makes any hosted endpoint a convenience rather than a channel. We specify provenance-binding as the constitutive constraint on the model's outputs: every judgment must resolve to a citation into a fixed canonical corpus, and a judgment that cannot be so resolved is withheld rather than emitted. This trades coverage for auditability deliberately, and it distinguishes the artifact from values models trained on preference data whose sources cannot be named, and from purpose-authored value-rule corpora, whose rules are written for the alignment task itself and therefore cannot serve as an independent ground truth. Finally we specify that a single such artifact bears three non-interchangeable relations to the systems it judges, selected by carrier: a gate in the publisher's own hardware, a citation requirement without veto in an autonomous successor agent, and a referee in the wider world. We state plainly that the middle case must not be implemented as the first, because a veto held by a smaller model over a more capable agent bounds that agent at the evaluator's ceiling — the weak-supervisor problem applied to the very system the arrangement exists to enable. A consequence we did not initially see, and which we regard as the most immediately actionable result in the paper: the two postures are complements rather than alternatives, and the natural first evaluation subject for a referee is a filter. A filter's characteristic failure is silent bypass; an evaluator watching its record converts that failure into a recorded one. And because safety classifiers are frequently published open-weight and emit discrete, samplable decisions, this is the one evaluation target for which the access problem does not arise at all — no cooperation, permission, or notice is required from the artifact's publisher. We do not claim to have solved scalable oversight. We claim that a narrow, citation-bound, openly published evaluator is a tractable and underoccupied position in the design space, and that its tractability comes precisely from what it refuses to do. --- Provenance. This paper is part of the THonly research corpus, dedicated to the public domain under CC0 1.0. The canonical version is at https://thonly.org/research/the-referee-not-the-governor. Its SHA-256 is 1184b0f3e5408a504c60be2d542b551df1c22facedfe18bb9a689bb50a9cc3fe, independently timestamped to the Bitcoin blockchain via OpenTimestamps and signed under RFC 3161 by three trust authorities, one of them eIDAS-qualified. AI co-authorship is disclosed. Miss Aquarius is the consistent name used for the AI collaboration across all venues.

Community

0 comments
Use Connect Wallet in the navigation

No discussion yet

Be the first to share a question or observation.