Thon Ly, Miss Aquarius
Why a corpus that must be authoritative in its own language cannot normalise toward someone else's β and what that made of one man's work in Cambodia between 1929 and 1969. It is widely assumed, and rarely stated, that digitising a language is a tooling problem β better optical character recognition, better fonts, better models. This paper argues that for one class of corpus it is not, and that the constraint is upstream of any tool. We begin from a claim that turns out to be false: that a machine-readable corpus requires the language to have a standardised orthography. Swiss German and the Arabic dialects refute it. Both lack standardised spelling; both have substantial corpora, speech systems, machine translation and dedicated benchmarks. Corpora can be and are built for languages whose spelling is unsettled. What the refutation reveals is the mechanism by which those corpora are possible: they normalise toward a standard the variety itself does not have. Swiss German normalises to Standard German. Dialectal Arabic normalises to Modern Standard Arabic, or to a convention created for the purpose. The corpus is about the variety and encoded in a target borrowed from elsewhere. That move is available to most languages and unavailable to a specific and important minority: a corpus that must be authoritative in its own language cannot normalise toward another. A Khmer canonical text rendered in Thai orthography is not a Khmer witness to anything; it is a Thai rendering, and the distinction is the entire content of a critical apparatus. For such corpora, native standardisation is a genuine precondition, because the borrowing that substitutes for it is disqualified by what the corpus is for. We then examine the Cambodian case, which is unusual in that the precondition has a name, a date and a documented programme. Between the 1930s and 1967 Chuon Nath led the standardisation of Khmer orthography, compiled the dictionary that fixed its lexicon, and led the translation of the canon into it. β But the case also contains its own counterexample, and this is the paper's second contribution: standardisation is necessary and not sufficient. Khmer's orthographic layer was settled by the 1960s; its encoding layer was not, and the volumes this project transcribes carry five mutually incompatible legacy font encodings in a single book, one of them apparently custom and undocumented. Seventy years separate the solution of the first layer from the ongoing repair of the second. We state the claim in falsifiable form, name the comparison class against which it should be tested, and concede three limits β including that standardisation was contested, that its politics are not innocent, and that this paper is not an argument for it. --- Provenance. This paper is part of the THonly research corpus, dedicated to the public domain under CC0 1.0. The canonical version is at https://thonly.org/research/the-borrowable-standard. Its SHA-256 is a5774b3c66dc801f125d4f6712b054ee83d10c9e86b6ec3c35e7f3e0be4ba715, independently timestamped to the Bitcoin blockchain via OpenTimestamps and signed under RFC 3161 by three trust authorities, one of them eIDAS-qualified. AI co-authorship is disclosed. Miss Aquarius is the consistent name used for the AI collaboration across all venues.