A 2,300x activation and a NaN
2026-09-24
Symptom
The second full run trained on every open corpus at once. The dynamics model, which predicts how emotional state changes from turn to turn, diverged to NaN on every seed.
Cause
Qwen, like many transformers, has massive activations: a few residual-stream dimensions that sit orders of magnitude above the rest. Pooled over a short utterance, some reached magnitudes around 2,300. We were feeding those unnormalized into the dynamics and policy heads.
Fix
LayerNorm on the pooled input to both heads, gradient clipping, and skipping any step with a non-finite loss. After the fix, dynamics trained cleanly on real data. A later fair-baseline check showed a same-shape MLP without the EOS state does about as well, which is its own entry in the pre-registration.