Skip to content
Affective
← Updates

The model had seen the test

2026-09-24

Archived build note. This records what was known at publication; the active person-state program and voice gates have since moved. Read the current status.

The first leak

On 09-20 we found the training config still pointed at the original author-labeled corpus, which included 113 validation and 114 test events. The fix was a merged, leakage-free corpus with the splits enforced at build time.

The second leak

On 09-24, auditing the v3 corpus builder, we found it had trained on all 84 test items. Every v3 number measured before that point was contaminated. The builder now calls a hard assertion that no held-out item is present, and a run that violates it fails instead of training.

The same audit found the H5 evaluation was reading the wrong head at the wrong layer, and that the dynamics model was trained on vignette pairs that had nothing to do with each other. All three are fixed, and the corrections are recorded in the pre-registration's amendment log with their dates.

Why this is a post

A leak inflates results in exactly the direction you hope for. That is why it has to be written down where people will see it, not quietly patched before the next run.