Affective Voice handles speech input and eventual authorized audio output around Affective Core. It does not decide a person’s state or choose a response. A pitch, pause, transcript, or speech-emotion estimate is evidence about an utterance, not a diagnosis.
01
What is built, and its limit
Offline perception
Implemented by Oct 11
Decode, voice activity detection, ASR, prosody, uncertainty, provenance, a consent-gated personal baseline, degraded-input handling, and a versioned SpeechObservation. Thirteen of thirteen first-milestone work packages are complete.
Remaining boundary
Detailed acoustic spans depend on a human annotation round. No real Core is connected yet.
Measured audio checks
Implemented by Oct 11
A 24-clip clean-speech set gave a pooled voice-activity boundary p95 of 116 ms against a 250 ms threshold. A separate fixture baseline measured 3.87% word error rate.
Remaining boundary
Scoped offline fixture results, not noisy, live, multi-speaker, or emotion-accuracy claims.
Speech emotion estimate
Implemented by Oct 11
An interface, a stub that can abstain, and research adapters exist.
Remaining boundary
No trained, rights-approved production model and no demonstrated within-person advantage.
Speech output
Implemented by Oct 11
A synthesis framework and a first Chatterbox Turbo adapter, with native 24 kHz output, seeded deterministic audio, and watermark detection checked on CPU.
Remaining boundary
Not production-approved. CPU real-time factor was about 12. Listening tests and voice-profile authorization remain open.
Data and consent
Implemented by Oct 11
Rights registry, consent ledger, encrypted collection store, label schema, and an automated annotation dry run.
Remaining boundary
The human three-rater dry run is open. No real-user collection is authorized.
Area
Implemented by Oct 11
Remaining boundary
Offline perception
Decode, voice activity detection, ASR, prosody, uncertainty, provenance, a consent-gated personal baseline, degraded-input handling, and a versioned SpeechObservation. Thirteen of thirteen first-milestone work packages are complete.
Detailed acoustic spans depend on a human annotation round. No real Core is connected yet.
Measured audio checks
A 24-clip clean-speech set gave a pooled voice-activity boundary p95 of 116 ms against a 250 ms threshold. A separate fixture baseline measured 3.87% word error rate.
Scoped offline fixture results, not noisy, live, multi-speaker, or emotion-accuracy claims.
Speech emotion estimate
An interface, a stub that can abstain, and research adapters exist.
No trained, rights-approved production model and no demonstrated within-person advantage.
Speech output
A synthesis framework and a first Chatterbox Turbo adapter, with native 24 kHz output, seeded deterministic audio, and watermark detection checked on CPU.
Not production-approved. CPU real-time factor was about 12. Listening tests and voice-profile authorization remain open.
Data and consent
Rights registry, consent ledger, encrypted collection store, label schema, and an automated annotation dry run.
The human three-rater dry run is open. No real-user collection is authorized.
02
Core owns the decision
Authorized audio → uncertain SpeechObservation → Affective Core → reply text and SpeechDirective → authorized synthesis
Core owns appraisal, person state, and response policy. Voice owns no emotional-state ledger. Assistant-only audio or playback must never silently become evidence about the person.
03
Design decisions
Core input
Current position
Core receives utterance summaries plus transcript for the MVP. Raw audio and embeddings do not travel on the general event bus.
Voice output
Current position
Off-the-shelf speech synthesis executes capability-checked directives from Core. Unsupported controls are reported as unsupported.
Sequence
Current position
Offline evidence and authorized audio rendering meet Core in a reference service first. Streaming comes after the MVP.
Personalized speech
Current position
A later study tests whether consented within-person baselines add information beyond text and history. A voice-only result would not establish a person-state result.
Rights and safety
Current position
Consent to analyze speech does not grant voice cloning or training rights. No medical, deception, personality, or private-emotion diagnosis follows from these measurements.
Decision
Current position
Core input
Core receives utterance summaries plus transcript for the MVP. Raw audio and embeddings do not travel on the general event bus.
Voice output
Off-the-shelf speech synthesis executes capability-checked directives from Core. Unsupported controls are reported as unsupported.
Sequence
Offline evidence and authorized audio rendering meet Core in a reference service first. Streaming comes after the MVP.
Personalized speech
A later study tests whether consented within-person baselines add information beyond text and history. A voice-only result would not establish a person-state result.
Rights and safety
Consent to analyze speech does not grant voice cloning or training rights. No medical, deception, personality, or private-emotion diagnosis follows from these measurements.
04
What remains gated
Human listening evaluation, voice-profile authorization, directive mapping, Core integration, streaming, production model approval, and real-user collection are not complete. Future work may test consented within-person vocal change and full-duplex conversation, but neither is a released capability or evidence that audio improves person-state decisions.
To judge value, text-only, transcript plus acoustics, and any validated speech hypothesis will be compared on the same people and outcome, with speaker-disjoint splits and calibration. A demo or a listening preference score cannot stand in for that result.