Skip to content
Affective
← Updates

Labels that turned out not to be human

2026-09-20

Archived build note. This records what was known at publication; the active person-state program and voice gates have since moved. Read the current status.

The pipeline

Appraisal labels come from a self-hosted Label Studio instance: 17 appraisal dimensions per vignette, anchored guidelines, repeat items to check each rater against themselves.

The correction

The first two annotation rounds looked like six raters and 560 items. A timing audit showed every identity submitted all of its tasks in about two minutes, with uniform timings and byte-identical repeats. They were generated, not human. We relabeled them as synthetic, reset the human count to zero, and added the audit to the pipeline.

Round three

Round three was real: 560 vignettes, 3-4 calibrated raters each, 1,772 per-rater ratings reduced to 560 consensus labels. Agreement was quadratic-weighted kappa 0.61 and Krippendorff's alpha 0.82. Unweighted kappa was 0.24, below the original gate, so the gate metric was amended to weighted kappa or alpha after this round and signed off by the owner. That amendment is dated in the pre-registration, and the unweighted number is still reported.

A later round of AI-written vignettes failed human verification twice (identical submissions, then near-random agreement). Those labels stay marked as weak supervision and are kept out of every gate.