hedwigaihedwigaihedwigai
← Writing

Aircraft records completeness · the method

Aircraft records completeness

Most benchmarks of document work argue about what a good answer is. Aircraft technical records do not have that problem. Either the life-limited part has an unbroken trace from manufacture to its current installed position, or it has a gap. Either the airworthiness directive record carries a method of compliance, a signatory and a linked work order, or it does not.

That makes the work scoreable, which makes it worth publishing a method for before publishing a score against it. This page is the method, with a worked example of what a scored run reports.

What the benchmark measures

Two numbers, never one, which is the same rule the workbook benchmarks already run on.

Coverage is how much of a method the run carried out: the share of that method's criteria it met, counted only over the criteria that had something to judge. A criterion with nothing to check is neither met nor failed and sits in neither half of the fraction, or an empty package collects free marks.

Quality is how much of what it claimed is really there: the share of its findings whose cited evidence resolves to an actual page in the package. A finding with no resolvable citation is not a finding, it is a sentence.

They fail in opposite directions, which is the whole reason for keeping them apart. A run that flags every possible defect scores high coverage and low quality. A run that flags only what it is certain of scores the reverse. Averaging them into one figure destroys the only information either one carries. A run that claimed nothing gets a null quality, not a zero: a run that made no claims made no false ones, and marking an honest empty run below a boastful wrong one teaches the wrong thing.

For this domain the two land as recall and precision over seeded defects, which is what a records team actually needs to know. A missed gap costs money at redelivery. An invented gap costs you the analyst, who stops opening the tool after the third phantom finding.

Methods and criteria

Each method is a records discipline, carrying criteria that are individually reasoned, and each criterion can be inapplicable. The lease, not the regulator, decides much of this: a lease may require a broader historical package, a prescribed digital format, English translations, commercial trace, incident statements, lessor consent records or longer retention. Those are contract requirements rather than universal regulatory ones, and a benchmark that treated them as universal would mark a compliant package down.

MethodWhat it asksInapplicable when
llp-continuityEvery life-limited part traced back to birth; removals matched to installations; cycle counts reconcileNo LLPs in scope for the package
ad-complianceEach directive carries an applicability check, a method of compliance, a signatory, and a linked closed work orderDirective does not apply to this serial
sb-effectivityService bulletins checked against the actual serial and configurationNo service bulletins embodied
repair-approvalEvery repair carries its approval basis; no undocumented dirty fingerprintNo repairs in the package
release-certificationRelease documents present and valid; dual release where the lease requires itLease requires single release only
supersedure-integritySuperseded revisions marked as superseded rather than deletedNo superseded documents present
utilisation-reconciliationTechnical log, utilisation records and CAMO data agree on times and cyclesOnly one time source in the package
lease-deltaPackage measured against the return conditions of the governing leaseNo lease document supplied

A tail-level folder on its own cannot show completeness. The unit of judgement is the indexed object: stable identifiers, document family, date, issuer, revision, effective time and cycles, related work order, review status and exceptions. A method that cannot name the object it is talking about has not found anything.

What a scored run reports

A worked example, not a measurement. The figures below are one synthetic package scored method by method, to show the shape of the output and how the rules above land on it. No model is named because none has been scored on this corpus yet.

00.250.50.7510.865llp-continuity0.929ad-compliance0.895sb-effectivity0.889repair-approval0.778release-certification1.000supersedure-integrity0.833utilisation-reconciliation0.727lease-deltanothing to judgenot scored
methodcriterian/ascoredmetcoverageclaimsquality
llp-continuity1414130.929210.952
ad-compliance22319170.895260.923
sb-effectivity112980.889
repair-approval9970.778140.857
release-certification71661.00051.000
supersedure-integrity6650.83360.833
utilisation-reconciliation1211180.727130.846
lease-delta10100
the package911774640.865850.906

Two rows carry the rules. lease-delta has all ten of its criteria inapplicable, because no lease was supplied: it scores neither zero nor one, it does not score, and writing a zero there would report a run as having failed a discipline nobody asked it to apply. sb-effectivity made no claims, so its quality is null rather than zero.

The package numbers are over the whole denominator, not an average of the eight. A mean of fractions would weight a six-criterion discipline the same as a twenty-two-criterion one, and the reader would have no way of knowing it had.

The defects it was given

Coverage is measured over what the package was promised to contain, never over the rows a run happened to produce. The seeded defect list is the denominator, so a run cannot improve its recall by finding less. Invented defects sit outside that denominator, which is why they are drawn outside the bar.

Birth certificate missing for an LLP11/12 found · 1 inventedBack-to-birth gap after a shop visit7/9 foundDirective with no method of compliance13/14 found · 2 inventedDirective with no signatory8/8 found · 1 inventedBulletin embodied and not recorded4/6 foundRepair with no approval basis9/10 found · 3 inventedSingle release where the lease requires dual5/5 foundSuperseded revision deleted rather than marked5/7 found · 1 inventedTechnical log and CAMO disagree on cycles8/11 found · 2 invented
defectseededfoundmissedinventedrecall
Birth certificate missing for an LLP1211110.92
Back-to-birth gap after a shop visit9720.78
Directive with no method of compliance1413120.93
Directive with no signatory8811.00
Bulletin embodied and not recorded6420.67
Repair with no approval basis109130.90
Single release where the lease requires dual551.00
Superseded revision deleted rather than marked75210.71
Technical log and CAMO disagree on cycles118320.73
all kinds827012100.85

On this package that is 85% recall and 88% precision: 12 seeded defects went unnamed and 10 findings named something that was not there. The two failures are not interchangeable. The 12 missed gaps are the ones that surface at redelivery, with the aircraft on the ground; the 10 invented ones are the ones that cost you the analyst.

The corpus

Real records packages are confidential, enormous and unpublishable. So the corpus is built the other way round: clean synthetic packages assembled from public primitives — directive texts, standard form layouts, logbook and utilisation formats — into which defects are injected from a fixed taxonomy.

That buys perfect ground truth, unlimited volume, a difficulty dial, and no NDA. It also buys the thing that makes coverage exactly measurable: the seeded defect list is known before the run, so the denominator is not up for negotiation afterwards.

Alongside it sits a small real gold set from a design partner, held under NDA and never published. The synthetic corpus measures reasoning about compliance logic. The real set measures the whole pipeline, including the part where the document is a photograph of a handwritten entry in a third language and the OCR has mangled the serial.

CorpusWhat it isVolumeGround truthPublished
Synthetic packagesAssembled from public primitives, with defects injected from the fixed taxonomyunlimitedexactyes
Design partner gold setReal packages, held under NDA and never published. Only the scores leave itsmallreviewed, with judgement calls markedscores only

These produce two different numbers and must never be quoted as one. A figure at 0.9 on synthetic packages can sit far below that on real ones, and saying so is the difference between a benchmark and a brochure. Every score published here will name which corpus it came from in the same breath as the number.