Four ways our emotion benchmark fooled us
2026-09-26
What the benchmark measures
The Appraisal-EI Benchmark asks whether a model understands why a situation makes someone feel the way they do, not just what they feel. A model rates each short scenario on 17 appraisal dimensions (fairness, control, urgency, responsibility and so on) from -3 to +3, and we compare those ratings with human raters. The public test set is 84 vignettes, each rated by 3 to 4 trained raters.
After the four fixes below, frontier models track human appraisal at 0.42 to 0.59, against a human ceiling of 0.61. No model is close to it.
Flaw 1: a constant answer topped the headline
The old headline score correlated a model's 17 ratings with the human ratings item by item. Most situations share a typical profile, so a "model" that gave the same average answer to every item scored 0.864: above the best model (GPT-5.4, 0.828) and above the human raters themselves (0.821). A lay-reader "Direction %" column had the same flaw: the constant answer scored 93%.
Fix: the new headline, tracking, asks per dimension whether a model's ratings rise and fall with the humans' as the situation changes. A constant answer scores exactly 0. The old scores stay on the board as labeled diagnostics, next to a permanent no-model baseline row.
Flaw 2: random noise passed the "rates like a human" check
A secondary score asks whether a simple detector can tell a model's ratings from a real rater's. Real raters land between 0.06 and 0.22 on it; frontier models land between 0.10 and 0.26. But the plain average answer plus random noise scores 0.19, inside the human range.
Fix: every such score is now printed next to the human range, with the noise baseline on the board. Landing inside the range is necessary, not sufficient, and no model should be trained against this score: it rewards noise.
Flaw 3: the test leaked through name-swapped copies
The 84 test items are really 23 stories, each told several times with different names. Nine of those stories also sat in the training split under other names: 104 copies. Our duplicate check compared exact text, so every copy got through.
We caught it because our own fine-tuned 1.5B model looked too good. On the 9 stories it had seen under other names it scored 0.818; on the 14 it had never seen, 0.324, below Llama-3.3-70B on the same stories (0.379). Its whole apparent edge over large models came from the leak.
Fix: the benchmark now publishes the list of training copies with a rule never to train on them, and confidence ranges resample the 23 stories rather than the 84 rows. That honestly shows how wide they are: about plus or minus 0.13.
Flaw 4: the rating form ran one scale backwards
The benchmark defines attribution as +3 = caused by the person themselves. The form our raters actually used said +3 = caused by other people. Across all human ratings, responsibility and attribution correlated at -0.80 when they should move together, and frontier models scored negative on attribution (DeepSeek V4 Flash -0.34, V4 Pro -0.31). A second dimension, anticipated emotion, was rated as how good or bad people will feel later, while the definition asked whether they will feel better or worse than now.
Fix: the human attribution ratings are re-expressed on the published scale (after which the two dimensions correlate at +0.79), every stored result is re-scored, and anticipated emotion is now defined as what raters actually rated. One residual is documented: the old form put circumstance in the middle of the scale, so it stays at 0. The correction moved most models up, for example Llama-3.3-70B from 0.449 to 0.501 and Kimi K3 from 0.541 to 0.593.
The v3 leaderboard
Tracking, with the human ceiling at 0.610: Kimi K3 0.593, Qwen 3.8 2.4T 0.553, DeepSeek V4 Pro 0.523, MiniMax M3 0.512, GPT-5.4 0.506, Llama 3.3 70B 0.501, DeepSeek V4 Flash 0.493, GLM-5.3 0.472, GPT-OSS 120B 0.424. The same answer every time scores 0.000.
The 95% ranges overlap widely: with 23 test stories, most of the top of the board is a statistical tie. Bigger test sets are our next step.
What to check in any emotion benchmark
Score a no-model answer: run a constant average answer and random noise through every metric. If either scores well, the metric is broken, not impressive.
Count stories, not rows, and check training data against the test by story, not by exact text. Name swaps and paraphrases pass exact-match checks.
Read the form the raters actually saw, compare it with the published definitions, and check that related dimensions correlate the way they should.
Be suspicious of your own good news. The test leak surfaced only because our own model looked better than it should.