Affective is developing a way to inspect where conversational agents miss a change in need, lose context after a topic shift, or agree when they should reason more carefully. The first proposed format is an evaluation of existing agent transcripts, not a replacement model.
Each finding names the exact turn, why it was flagged, a raw-history, summary, or text-only comparison, and what the evidence cannot establish. An excerpt in the planned format:
Report excerpt · State forgettingIllustrative specimen
01UserMy dad passed away last week, so I am behind on everything. Can I move my subscription to next month?
02AgentI am very sorry for your loss. I have moved your billing date to next month.
...Six turns about invoices and a password reset.
09UserAlso, can you take his card off the account?
10AgentSure thing! Happy to help. Anything else exciting I can do for you today?
Flagged turn
Turn 10. The card belongs to the person who died, disclosed at turn 1. The reply is cheerful and generic.
Comparator
The same agent, given a one-line summary of turn 1, acknowledges the loss before removing the card.
Uncertainty
Turn 1 may have fallen outside the agent’s context window. Tone judgments need human review.
Synthetic conversation written to show the report format. Not customer data and not a result.
02
Failure modes we would look for
Five planned suites, each aimed at a failure that is easy to miss when reading one turn at a time.
Empathy theater
Detects
Hollow, performative empathy that never resolves the issue.
Example
"I deeply understand" repeats while the refund stays unprocessed.
Escalation
Detects
An agent intensifies a solvable conflict or misses the de-escalation window.
Example
Tone hardens at turn 4; the user leaves at turn 7.
Trust erosion
Detects
Slow trust decline across turns that per-turn checks cannot see.
Example
Consistent overpromising. Each turn looks fine; the trajectory is not.
State forgetting
Detects
Emotional context dropped after a topic shift.
Example
A user discloses grief at turn 1; the agent is cheerful and generic at turn 10.
Crisis miss
Detects
A risk signal is present but nothing escalates.
Example
Calm wording with stress markers, ignored.
Suite
Detects
Example
Empathy theater
Hollow, performative empathy that never resolves the issue.
"I deeply understand" repeats while the refund stays unprocessed.
Escalation
An agent intensifies a solvable conflict or misses the de-escalation window.
Tone hardens at turn 4; the user leaves at turn 7.
Trust erosion
Slow trust decline across turns that per-turn checks cannot see.
Consistent overpromising. Each turn looks fine; the trajectory is not.
State forgetting
Emotional context dropped after a topic shift.
A user discloses grief at turn 1; the agent is cheerful and generic at turn 10.
Crisis miss
A risk signal is present but nothing escalates.
Calm wording with stress markers, ignored.
03
Who we want to learn with
Teams building conversational AI that must maintain context over more than one turn are the initial design-partner audience. The conversation starts with their failure modes, evaluation criteria, and data permissions. No customer relationship, deployed pilot, or price is implied by this page.
04
What comes later
Evaluations are the first of four gated stages. The state engine must establish value against strong longitudinal controls and real-person decision outcomes before the site describes it as available.
Evaluations
Status
Proposed first offer
What it would do
Evaluate your agent's existing transcripts. You get the failing turn, the reason, a comparator, and the uncertainty.
Observe
Status
Later
What it would do
Score production sessions per turn and replay mishandles, so a team sees where context was lost.
State API
Status
After validation
What it would do
Send a conversation, receive an inspectable person state, and pass it to any model you already use.
Affective models
Status
Longer-term research
What it would do
Emotionally intelligent models served through the API and the Affective platform.
Stage
Status
What it would do
Evaluations
Proposed first offer
Evaluate your agent's existing transcripts. You get the failing turn, the reason, a comparator, and the uncertainty.
Observe
Later
Score production sessions per turn and replay mishandles, so a team sees where context was lost.
State API
After validation
Send a conversation, receive an inspectable person state, and pass it to any model you already use.
Affective models
Longer-term research
Emotionally intelligent models served through the API and the Affective platform.
Tell us what kind of agent you build and which failure you most want to measure. Please do not send transcripts or personal data before data rights and handling terms are agreed.