Skip to content
Affective
← Updates

A benchmark before a model

2026-09-20

Archived build note. This records what was known at publication; the active person-state program and voice gates have since moved. Read the current status.

What it is

The Appraisal-EI Benchmark scores how closely a model's appraisal of a situation matches human raters, across 17 dimensions. It launched as v1.0 on 09-18 with a scorer, an eval runner, and a public leaderboard.

The freeze

On 09-20 v2.0 froze the protocol: a human-gold held-out set, a human ceiling row, and a leaderboard of nine measured frontier and open models. Every score on it is measured against human consensus, and the human baseline is drawn on every chart so no model is read without it.

Freezing it before our own model is ready matters. A benchmark you can still edit once you have seen your own number is not a benchmark.