Claim only what runs.
Every quoted measurement traces to a provenance index. Published figures are machine-checked in CI against committed run artifacts, so prose cannot silently disagree with the data. The same read-back discipline applied to documentation.
The instrumentation paper
By combining Total Survey Error reasoning with an asymmetric neuro-symbolic architecture, LCF makes error sources distinguishable at the case level — converting undifferentiated performance failure into attributable evidence for repair.
LCF uses three independently produced classifications: an adjudicated reference disposition, a disposition implied by the formal specification, and the disposition produced by the evaluated system. The asymmetric architecture makes these comparisons operational.
The contribution is not a new accuracy metric, nor a claim that these error concepts are new. It is a bridge between mature statistical measurement theory and AI evaluation — an instrument through which conflated error sources become separately observable.
The evidence base
Paired-arm Lab harness
Baseline without the gate against LCF, reporting Type I and Type II rates plus extraction fidelity across six categories: simple, negation, conditional, quantifier, anaphora, multi-claim. Every printed aggregate asserts equality with the mean of its on-disk per-case rows.
Corpus & μ gold standard
Scenarios authored blind to the desired verdict, then decomposed in a separate pass and adjudicated by a domain expert. Corrections are the most valuable rows — they capture where an expert's reading diverges from a competent but naive one. Every row is tagged with how its gold was established.
Statistical methodology
The three-way error split follows a review by Dr. Paul P. Biemer (RTI International, Emeritus), whose reading of Field Manual v0.2 identified that a row then labelled "measurement error" was in fact describing decision error at the gate. The correction is logged in the defect ledger.
Red team & prior art
Independent pre-mortems against the novelty claim, seam-negation probes, and a prior-art master ledger. Findings are dispositioned in the open rather than resolved privately — including the ones that narrowed the claim.
Claim discipline
Mock-extraction results are statements about the kernel and shapes. They say nothing about real-extraction fidelity, and are never presented as if they did.
Numbers in prose reference the gated artifact rather than being hand-copied. Where a document and the committed data disagree, the data wins.
Historical entries are never rewritten to match the present — that is the same bitemporal sin as editing a grandfathered rule's past.