Skip to content
Affective
← Research

Research track · Voice · October 11, 2026

Speech observations, with uncertainty.

Affective Voice handles speech input and eventual authorized audio output around Affective Core. It does not decide a person’s state or choose a response. A pitch, pause, transcript, or speech-emotion estimate is evidence about an utterance, not a diagnosis.

01

What is built, and its limit

Offline perception

Implemented by Oct 11

Decode, voice activity detection, ASR, prosody, uncertainty, provenance, a consent-gated personal baseline, degraded-input handling, and a versioned SpeechObservation. Thirteen of thirteen first-milestone work packages are complete.

Remaining boundary

Detailed acoustic spans depend on a human annotation round. No real Core is connected yet.
Measured audio checks

Implemented by Oct 11

A 24-clip clean-speech set gave a pooled voice-activity boundary p95 of 116 ms against a 250 ms threshold. A separate fixture baseline measured 3.87% word error rate.

Remaining boundary

Scoped offline fixture results, not noisy, live, multi-speaker, or emotion-accuracy claims.
Speech emotion estimate

Implemented by Oct 11

An interface, a stub that can abstain, and research adapters exist.

Remaining boundary

No trained, rights-approved production model and no demonstrated within-person advantage.
Speech output

Implemented by Oct 11

A synthesis framework and a first Chatterbox Turbo adapter, with native 24 kHz output, seeded deterministic audio, and watermark detection checked on CPU.

Remaining boundary

Not production-approved. CPU real-time factor was about 12. Listening tests and voice-profile authorization remain open.
Data and consent

Implemented by Oct 11

Rights registry, consent ledger, encrypted collection store, label schema, and an automated annotation dry run.

Remaining boundary

The human three-rater dry run is open. No real-user collection is authorized.
02

Core owns the decision

Authorized audio → uncertain SpeechObservation → Affective Core → reply text and SpeechDirective → authorized synthesis

Core owns appraisal, person state, and response policy. Voice owns no emotional-state ledger. Assistant-only audio or playback must never silently become evidence about the person.

03

Design decisions

Core input

Current position

Core receives utterance summaries plus transcript for the MVP. Raw audio and embeddings do not travel on the general event bus.
Voice output

Current position

Off-the-shelf speech synthesis executes capability-checked directives from Core. Unsupported controls are reported as unsupported.
Sequence

Current position

Offline evidence and authorized audio rendering meet Core in a reference service first. Streaming comes after the MVP.
Personalized speech

Current position

A later study tests whether consented within-person baselines add information beyond text and history. A voice-only result would not establish a person-state result.
Rights and safety

Current position

Consent to analyze speech does not grant voice cloning or training rights. No medical, deception, personality, or private-emotion diagnosis follows from these measurements.
04

What remains gated

Human listening evaluation, voice-profile authorization, directive mapping, Core integration, streaming, production model approval, and real-user collection are not complete. Future work may test consented within-person vocal change and full-duplex conversation, but neither is a released capability or evidence that audio improves person-state decisions.

To judge value, text-only, transcript plus acoustics, and any validated speech hypothesis will be compared on the same people and outcome, with speaker-disjoint splits and calibration. A demo or a listening preference score cannot stand in for that result.

Want the full voice ledger?

Work packages, gates, and evidence live in a private repository. Ask at founder@affective-llc.site. See also the overall evidence map.