Labels that turned out not to be human
2026-09-20
The pipeline
Appraisal labels come from a self-hosted Label Studio instance: 17 appraisal dimensions per vignette, anchored guidelines, repeat items to check each rater against themselves.
The correction
The first two annotation rounds looked like six raters and 560 items. A timing audit showed every identity submitted all of its tasks in about two minutes, with uniform timings and byte-identical repeats. They were generated, not human. We relabeled them as synthetic, reset the human count to zero, and added the audit to the pipeline.
Round three
Round three was real: 560 vignettes, 3-4 calibrated raters each, 1,772 per-rater ratings reduced to 560 consensus labels. Agreement was quadratic-weighted kappa 0.61 and Krippendorff's alpha 0.82. Unweighted kappa was 0.24, below the original gate, so the gate metric was amended to weighted kappa or alpha after this round and signed off by the owner. That amendment is dated in the pre-registration, and the unweighted number is still reported.
A later round of AI-written vignettes failed human verification twice (identical submissions, then near-random agreement). Those labels stay marked as weak supervision and are kept out of every gate.