Skip to content
Affective
← Updates

A 2,300x activation and a NaN

2026-09-24

Archived build note. This records what was known at publication; the active person-state program and voice gates have since moved. Read the current status.

Symptom

The second full run trained on every open corpus at once. The dynamics model, which predicts how emotional state changes from turn to turn, diverged to NaN on every seed.

Cause

Qwen, like many transformers, has massive activations: a few residual-stream dimensions that sit orders of magnitude above the rest. Pooled over a short utterance, some reached magnitudes around 2,300. We were feeding those unnormalized into the dynamics and policy heads.

Fix

LayerNorm on the pooled input to both heads, gradient clipping, and skipping any step with a non-finite loss. After the fix, dynamics trained cleanly on real data. A later fair-baseline check showed a same-shape MLP without the EOS state does about as well, which is its own entry in the pre-registration.