Updates
Research notes and dated entries from the build. Every result is reported with how it was measured, and nothing is a roadmap item dressed as progress.
Where the person-state and voice programs stand
The current POC uses gated longitudinal comparisons, while a separate voice repository builds uncertain speech observations.
Four ways our emotion benchmark fooled us
We audited our own Appraisal-EI Benchmark and found four flaws that distorted every score. Version 3.0.0 fixes all four.
Introducing Affective
Same company, same model, same bet. New name across every repo in one day.
The model had seen the test
Our training corpus contained every item of the held-out test set. Twice, in two different ways.
A 2,300x activation and a NaN
The dynamics model went NaN on every seed. The cause was a handful of enormous activations inside the base model.
Labels that turned out not to be human
Six raters, 560 vignettes, and a provenance correction we had to make on ourselves.
A benchmark before a model
We built the ruler first, froze it, and measured other people's models on it.
This site shipped, and immediately broke dark mode for everyone
The first Tailwind surface to render in dark mode exposed a kit bug that three reviews missed.
One meter to rule them
Two byte-identical meter implementations is the drift the kit exists to prevent.
Skeletons, status pages, and the pulse that wasn't
Loading states, error pages, and twelve rejected logos.
Emotion is a surface pattern in every model you have used
Why frontier models miss crisis signals, forget how you felt, and flatter you when they should not. The failure is structural.
The Emotional Operating System
Four components, injected into the residual stream.
The bet, stated falsifiably
One falsifiable claim carries the whole company.
What we are honest about
The objections we take seriously, stated by us before our critics.
The design system exists before the product does
Foundations, components, and the chat pattern in one day.
Pre-registration before experiments
The core claim is falsifiable, and the thresholds are fixed before any experiment runs.