Affective is testing whether a structured, inspectable person-state layer adds value beyond raw history, summaries, learned memory, and larger-model controls. The program uses sealed tests, matched comparisons, uncertainty, cost, and stop rules.
01
Active work
M0 Arena
Active
Registered loaders, sealed-test controls, a comparison harness, Sim-C, and a synthetic-tested diary pipeline exist. The full matched tuning was still running at the latest ledger entry. A 1% pilot checked the runner and cost; it was not a performance result.
M2 Text perception
Active early start
Event extraction and witness annotation tooling are in place. There are no human witness gold labels or event-label results yet.
Later milestones
Locked
Person-state architecture, forecast, decisions, efficiency, safety, and external validation still have their own gates.
02
Results with boundaries
Each finding below is paired with what it supports and what it does not. Synthetic, development, and real-person evidence are kept apart.
Deterministic ledger vs 70B full-history reader (next-day stress forecasting)
What it supports
The ledger won on a real-data forecasting task.
Limit
One task, small cohort. Does not show general decision value.
Typed state vs raw history, 14B reader, matched token budget
What it supports
Compression helped that reader under that budget.
Limit
An order-free count summary tied it. The cheap summary stays a required control.
Real longitudinal tests (person, time, event order)
What it supports
Structure is worth testing against history.
Limit
Recurrence did not prove necessary. The ledger only tied the person mean on Corona. A generic GRU won on PSPS.
Sim-C synthetic discrimination check (2026-10-10)
What it supports
The arena detects a planted exact-state advantage.
Limit
Synthetic only. The prompt-only 14B reader was weak, so this is not a scale victory.
Perspective boundary
What it supports
Assistant-only information cannot update a person's state.
Limit
Extracting correct events and proving better decisions on real people remain open.
Finding
What it supports
Limit
Deterministic ledger vs 70B full-history reader (next-day stress forecasting)
The ledger won on a real-data forecasting task.
One task, small cohort. Does not show general decision value.
Typed state vs raw history, 14B reader, matched token budget
Compression helped that reader under that budget.
An order-free count summary tied it. The cheap summary stays a required control.
Real longitudinal tests (person, time, event order)
Structure is worth testing against history.
Recurrence did not prove necessary. The ledger only tied the person mean on Corona. A generic GRU won on PSPS.
Sim-C synthetic discrimination check (2026-10-10)
The arena detects a planted exact-state advantage.
Synthetic only. The prompt-only 14B reader was weak, so this is not a scale victory.
Perspective boundary
Assistant-only information cannot update a person's state.
Extracting correct events and proving better decisions on real people remain open.
No M1 thesis verdict and no real-person decision-value verdict has been issued. The full M1 mode and the real-person verdict remain open.
03
Milestone tracker
Task counts show engineering work remaining, not scientific progress. A milestone unlocks only when its predecessor passes its own gate.
M0
Milestone
Arena: locked longitudinal benchmark
Status
Active
Tasks closed
23 of 37
M1
Milestone
Kill test: oracle state vs scale
Status
Locked
Tasks closed
0 of 16
M2
Milestone
Perception: events and witness from text
Status
Active (early start)
Tasks closed
8 of 34
M3
Milestone
Representation: compositional editable state
Status
Locked
Tasks closed
0 of 18
M4
Milestone
Prediction: state forecasts change
Status
Locked
Tasks closed
0 of 17
M5
Milestone
Decision: right action on real people
Status
Locked
Tasks closed
3 of 24
M6
Milestone
Causation and personalization
Status
Locked
Tasks closed
2 of 23
M7
Milestone
Efficiency moat: small + state beats large
Status
Locked
Tasks closed
1 of 19
M8
Milestone
Data moat: consented flywheel
Status
Locked
Tasks closed
1 of 23
M9
Milestone
Primitive: portable and safe
Status
Locked
Tasks closed
2 of 25
M10
Milestone
Credibility: external validation
Status
Locked
Tasks closed
0 of 19
Final
Milestone
Affective Model v1: our own model
Status
Locked
Tasks closed
0 of 23
ID
Milestone
Status
Tasks closed
M0
Arena: locked longitudinal benchmark
Active
23 of 37
M1
Kill test: oracle state vs scale
Locked
0 of 16
M2
Perception: events and witness from text
Active (early start)
8 of 34
M3
Representation: compositional editable state
Locked
0 of 18
M4
Prediction: state forecasts change
Locked
0 of 17
M5
Decision: right action on real people
Locked
3 of 24
M6
Causation and personalization
Locked
2 of 23
M7
Efficiency moat: small + state beats large
Locked
1 of 19
M8
Data moat: consented flywheel
Locked
1 of 23
M9
Primitive: portable and safe
Locked
2 of 25
M10
Credibility: external validation
Locked
0 of 19
Final
Affective Model v1: our own model
Locked
0 of 23
04
Human data and voice
The diary ethics and operations tier has nine open items. No participant contact or real diary enrollment is cleared. Voice has a separate offline observation path and an initial TTS adapter, but no Core integration or production voice service. Audio has not yet shown a validated improvement in person-state prediction or decisions. Read the voice research track.
05
How we word claims
Every number is paired with its dataset, population, comparator, and status. This is the wording rule we apply to ourselves.
We are testing an auditable person-state layer against strong longitudinal controls.
Not until the evidence exists
Our state improves real-person decisions.
Our synthetic Sim-C arena detected a planted exact-state effect.
Not until the evidence exists
A small state model beats large models.
Our voice work produces offline transcript and acoustic observations with provenance and uncertainty.
Not until the evidence exists
We infer how a person feels from their voice.
We plan transcript-based agent evaluations as the first product.
Not until the evidence exists
Evaluations or Runtime are live.
Safe to say now
Not until the evidence exists
We are testing an auditable person-state layer against strong longitudinal controls.
Our state improves real-person decisions.
Our synthetic Sim-C arena detected a planted exact-state effect.
A small state model beats large models.
Our voice work produces offline transcript and acoustic observations with provenance and uncertainty.
We infer how a person feels from their voice.
We plan transcript-based agent evaluations as the first product.
Evaluations or Runtime are live.
Want the full ledgers?
This page is a dated summary of our working ledgers, which live in private repositories. If you are a researcher, investor, or design partner and want run IDs, conditions, and intervals, ask at founder@affective-llc.site.