Voice feedback (text-to-speech)
Spoken rep feedback for live tracking: when a repetition is recorded, the app speaks that rep's first focus-metric value aloud in the app's language — e.g. "0.45 meters per second". It serves velocity-based-training style coaching, where the athlete hears their bar speed without looking at the screen.
It works in both tracking modes: active training (ExerciseDetailView) and
feedback training (the sensor-only, no-plan flow — see
tracking app features §2.8).
What gets spoken
The value is read from the same source as the per-rep bars, so spoken always matches shown:
- The metric is
useFocusMetric(exerciseGroupID)— the first selected focus metric (seeapps/tracking/src/app/workouts/today/focus-metric.tsx). - The per-rep value is
buildRepValueMap(rep, loadMass)[metric.key]— the identical map the live overlay/bars use (rep-feedback-shared.tsx). - The number text is
useUnits().formatMetric(value, metric).text(locale- and unit-system-aware). - The unit is spoken as a long word ("meters per second"), not the display
symbol ("M/S") — a synthesizer reads "m/s" as "m slash s". The spoken unit is
localized through
t(...)inrep-voice-utterance.ts(thelocalizepattern: explicit literalt("…")calls keyed by the English unit name frompackages/core/src/units.ts).
Local language
The voice language follows the active app locale (useTranslation().locale).
resolveVoiceLang (in @enode/core/text-to-speech) normalizes the locale to a
BCP-47 tag the on-device engine actually supports — it tries the exact tag, then
any installed tag sharing the base language ("de" → "de-DE"), and caches the
result. When no voice is installed for the language, nothing is spoken (the
feature degrades silently rather than reading in the wrong language).
Platforms
One common implementation via
@capacitor-community/text-to-speech:
- iOS →
AVSpeechSynthesizer - Android →
android.speech.tts.TextToSpeech - Web → Web Speech API (best-effort bonus; gated on
speechSynthesis)
The plugin module touches window at evaluation time, so
@enode/core/text-to-speech loads it lazily via dynamic import() — importing
it eagerly breaks Next.js static prerender (window is not defined).
speakValue uses QueueStrategy.Flush, so a burst of reps interrupts the
previous utterance instead of queueing and lagging behind the athlete.
Rate & latency
Speech rate is set per platform (VOICE_RATE_IOS / VOICE_RATE_DEFAULT in
@enode/core/text-to-speech) because the plugin scales rate differently: iOS
compresses it (0.1*rate + 0.4 on AVSpeech's 0–1 scale, max 1.0 — so the
current 2.5 lands at ~0.65 for a brisk callout), while Android applies it
straight to setSpeechRate (1.25, where 1.0 = normal). Both are tunable from
device testing.
prepareVoice(locale) removes the first-rep cold start: the dynamic import, the
supported-languages probe, and the engine spin-up + voice load would otherwise
all land on rep 1. The tracking hook calls it the moment voice is switched on
(well before the first rep) — it loads the plugin, primes the language cache,
and runs a silent (volume 0) warm-up utterance. Spoken reps are also kept short
(the unit only on the first rep of a set) to minimise per-rep speech time.
iOS audio session
IOS_AUDIO_CATEGORY in @enode/core/text-to-speech selects the
AVAudioSession category. The default "ambient" mixes with other audio and
respects the hardware silent switch (a muted phone stays silent);
"playback" is always audible but interrupts the athlete's music. This is the
single knob to tune from device testing. Voice also shares the audio session
with video recording (@enode/core/video, AVFoundation) — verify the two don't
fight when recording a set.
Voice owner — one station speaks
Several stations speaking over each other is meaningless, so exactly one
station owns the voice at a time. Ownership lives in
apps/tracking/src/app/workouts/today/voice-station.ts.
Every view able to speak registers itself with that module while mounted — live
training columns keyed by their station id, feedback-training flows keyed by
feedback:<sensorStationId> (the feedback: prefix is what keeps the two id
spaces apart). The set of candidates is therefore assembled in the store rather
than prop-drilled down from page.tsx. Two rules then decide the owner:
- Some active station always owns the voice, so the common single-station case never needs a tap and closing the owner hands off rather than falling silent.
- An owner that is still active keeps the voice when other stations open or close, so an explicit choice is never overridden.
Ownership is deliberately in-memory and not persisted: station ids are
ephemeral crypto.randomUUID() values minted per session, so a stored owner id
would never match anything on the next launch.
UI & persistence
The on/off switch is the "Voice feedback" row in the Local settings sheet
(local-settings-sheet.tsx, opened from the station sidebar) — device-local,
governing both tracking modes. The active tracking view's "⋯" options menu
(ExerciseDetailView.tsx) carries a second copy of the same toggle; both write
one store, so they can't disagree.
The speaker symbol is a station picker, not a power button. It appears in
both tracking views only while voice is switched on, showing SpeakerWaveIcon
(lit) on the owning station and SpeakerXMarkIcon (muted) elsewhere. Tapping a
muted speaker moves the voice to that station; tapping the lit one does nothing.
State is the voiceFeedback flag in @enode/core/training/flags, persisted in
localStorage (enode.voiceFeedback, default off). The whole feature sits
behind the reversible VOICE_FEEDBACK constant exported from
use-rep-voice.ts, shared by both speaking views.
Feedback training: no load
Feedback training runs without a workout, so there is no load to scale by and
useRepVoice is called with loadMass: null. buildRepValueMap then omits
loading-factor metrics (force, power, …) instead of leaving a per-kilogram value
in place — an unscaled number would be wrong, not merely unlabelled. The
feedback view only ever picks a load-independent focus metric (the first such
entry of the coach's selection), so nothing spoken there is a load-dependent
number, and spoken matches shown. A note under the mode name tells the coach
that those metrics are measured in a workout.