Skip to content
Affective

Evidence snapshot · October 11, 2026

What has been tested, and what has not.

Affective is testing whether a structured, inspectable person-state layer adds value beyond raw history, summaries, learned memory, and larger-model controls. The program uses sealed tests, matched comparisons, uncertainty, cost, and stop rules.

01

Active work

M0 Arena

Active

Registered loaders, sealed-test controls, a comparison harness, Sim-C, and a synthetic-tested diary pipeline exist. The full matched tuning was still running at the latest ledger entry. A 1% pilot checked the runner and cost; it was not a performance result.

M2 Text perception

Active early start

Event extraction and witness annotation tooling are in place. There are no human witness gold labels or event-label results yet.

Later milestones

Locked

Person-state architecture, forecast, decisions, efficiency, safety, and external validation still have their own gates.

02

Results with boundaries

Each finding below is paired with what it supports and what it does not. Synthetic, development, and real-person evidence are kept apart.

Deterministic ledger vs 70B full-history reader (next-day stress forecasting)

What it supports

The ledger won on a real-data forecasting task.

Limit

One task, small cohort. Does not show general decision value.
Typed state vs raw history, 14B reader, matched token budget

What it supports

Compression helped that reader under that budget.

Limit

An order-free count summary tied it. The cheap summary stays a required control.
Real longitudinal tests (person, time, event order)

What it supports

Structure is worth testing against history.

Limit

Recurrence did not prove necessary. The ledger only tied the person mean on Corona. A generic GRU won on PSPS.
Sim-C synthetic discrimination check (2026-10-10)

What it supports

The arena detects a planted exact-state advantage.

Limit

Synthetic only. The prompt-only 14B reader was weak, so this is not a scale victory.
Perspective boundary

What it supports

Assistant-only information cannot update a person's state.

Limit

Extracting correct events and proving better decisions on real people remain open.

No M1 thesis verdict and no real-person decision-value verdict has been issued. The full M1 mode and the real-person verdict remain open.

03

Milestone tracker

Task counts show engineering work remaining, not scientific progress. A milestone unlocks only when its predecessor passes its own gate.

M0

Milestone

Arena: locked longitudinal benchmark

Status

Active

Tasks closed

23 of 37
M1

Milestone

Kill test: oracle state vs scale

Status

Locked

Tasks closed

0 of 16
M2

Milestone

Perception: events and witness from text

Status

Active (early start)

Tasks closed

8 of 34
M3

Milestone

Representation: compositional editable state

Status

Locked

Tasks closed

0 of 18
M4

Milestone

Prediction: state forecasts change

Status

Locked

Tasks closed

0 of 17
M5

Milestone

Decision: right action on real people

Status

Locked

Tasks closed

3 of 24
M6

Milestone

Causation and personalization

Status

Locked

Tasks closed

2 of 23
M7

Milestone

Efficiency moat: small + state beats large

Status

Locked

Tasks closed

1 of 19
M8

Milestone

Data moat: consented flywheel

Status

Locked

Tasks closed

1 of 23
M9

Milestone

Primitive: portable and safe

Status

Locked

Tasks closed

2 of 25
M10

Milestone

Credibility: external validation

Status

Locked

Tasks closed

0 of 19
Final

Milestone

Affective Model v1: our own model

Status

Locked

Tasks closed

0 of 23
04

Human data and voice

The diary ethics and operations tier has nine open items. No participant contact or real diary enrollment is cleared. Voice has a separate offline observation path and an initial TTS adapter, but no Core integration or production voice service. Audio has not yet shown a validated improvement in person-state prediction or decisions. Read the voice research track.

05

How we word claims

Every number is paired with its dataset, population, comparator, and status. This is the wording rule we apply to ourselves.

We are testing an auditable person-state layer against strong longitudinal controls.

Not until the evidence exists

Our state improves real-person decisions.
Our synthetic Sim-C arena detected a planted exact-state effect.

Not until the evidence exists

A small state model beats large models.
Our voice work produces offline transcript and acoustic observations with provenance and uncertainty.

Not until the evidence exists

We infer how a person feels from their voice.
We plan transcript-based agent evaluations as the first product.

Not until the evidence exists

Evaluations or Runtime are live.

Want the full ledgers?

This page is a dated summary of our working ledgers, which live in private repositories. If you are a researcher, investor, or design partner and want run IDs, conditions, and intervals, ask at founder@affective-llc.site.