Not a model crisis.
A measurement crisis.
AI evaluation collapses specification, interpretation and decision error into a single performance statistic. LCF applies Total Survey Error reasoning so that a disagreement tells you not only that the system was wrong, but which part of it was.
The named hard problem
There is exactly one boundary in the pipeline where meaning becomes form. It is probabilistic by nature, and it is where the silent failures live: dropped negation, lost exception clauses, scope ambiguity, unresolved anaphora, flattened modality.
Determinism is real only downstream of it. So LCF makes the crossing explicit, requires every emitted triple to cite the span it came from, and records whatever the projection could not carry as typed residue.
No loss is silent.
"All production changes require design review, unless the change is a documented rollback of a release made within the previous 24 hours."
Three commitments
The gate validates the proposal, not the proposer's self-report.
Rules fire from structural conditions in the confirmed graph — an assertion and its negation triggering contradiction — never from a risk flag the extractor assigned to itself.
The human confirms the projection, never the raw triples.
Before validation, claims are back-translated into plain language and shown beside the original input. Measure twice, cut once. In testing, more than half of authored content defects changed no verdict — a structural gate alone cannot see them.
Every verdict recomputable from graph and shapes alone.
Each names shape, focus node, path, message and substrate version. No numeric coherence scores: v1 emits verdicts and named violations only, because scoring weights are banned until the Lab can justify them.
A conventional metric can report low agreement while leaving five incompatible explanations unresolved.
The model misunderstood the input. The evaluator implemented the rules incorrectly. The rules omitted a requirement. The reference classification was fallible. The decision threshold produced the wrong disposition.
Each requires a different remedy. Retraining a model cannot correct an incomplete specification; expanding a rule set cannot repair a parsing defect. An evaluation system that cannot separate these failures risks optimizing the wrong component.
Layered essence
Status is claimed only for what runs. Everything else is marked as specified and not yet built.
Substrate Loader
Loads a versioned domain ontology and its SHACL shapes verbatim; never mutates them. Swapping the domain module changes what is enforced with zero kernel changes.
SPECIFIEDExtraction
Probabilistic and swappable behind one interface. Every triple carries its source span; residue is mandatory. Strict JSON schema output with retry-on-violation — no free-text parsing anywhere.
IMPLEMENTEDRead-Back Gate
The reconciliation station. Extracted claims are back-translated into plain language beside the original input, and a human confirms the two say the same thing — or corrects.
SPECIFIEDValidator & Gate
The deterministic kernel. SHACL and named rules evaluated over the confirmed graph, emitting ACCEPT, MODIFY or REJECT. Any verdict path fixtures cannot reach is a build failure.
SPECIFIEDThe Lab
The evidence arm. A paired-arm harness — baseline without the gate against LCF — over labeled scenarios, reporting Type I and Type II rates and extraction fidelity per category.
IMPLEMENTEDDrift Instrumentation
The steward's ledger: last-reconciliation date, human-override rate, residue-type trends. Prevention is unprovable; decoherence is measurable. Waits until the Lab exists.
POST-V1Where it sits
Goes text to text. Its failure mode is the wrong or missing source. It can retrieve exactly the right document and still drop a negation in its answer — because it never converts the source into a checkable structured claim.
Deliverable: a fluent answer.
Sits underneath the submission-to-decision path, not the question-answering path. It makes the conversion explicit, surfaces residue, gates it with human read-back, and validates the confirmed graph.
Deliverable: a reconstructable, auditable verdict.
These are different failures with different fixes. Correcting one shifts the composition of the other: fix retrieval and the residual error concentrates into interpretation — which is precisely LCF's territory.
Reviewers and red-teamers wanted.
LCF is in pilot. We are looking for domain experts to adjudicate gold-standard decompositions, and adversarial readers to attack the shapes before they enter the substrate. Corrections are the most valuable contribution.
Not a better guess.
A better instrument.