Skip to main content

Voice feedback (text-to-speech)

Spoken rep feedback for live tracking: when a repetition is recorded, the app speaks that rep's first focus-metric value aloud in the app's language — e.g. "0.45 meters per second". It serves velocity-based-training style coaching, where the athlete hears their bar speed without looking at the screen.

It works in both tracking modes: active training (ExerciseDetailView) and feedback training (the sensor-only, no-plan flow — see tracking app features §2.8).

What gets spoken​

The value is read from the same source as the per-rep bars, so spoken always matches shown:

  • The metric is useFocusMetric(exerciseGroupID) — the first selected focus metric (see apps/tracking/src/app/workouts/today/focus-metric.tsx).
  • The per-rep value is buildRepValueMap(rep, loadMass)[metric.key] — the identical map the live overlay/bars use (rep-feedback-shared.tsx).
  • The number text is useUnits().formatMetric(value, metric).text (locale- and unit-system-aware).
  • The unit is spoken as a long word ("meters per second"), not the display symbol ("M/S") — a synthesizer reads "m/s" as "m slash s". The spoken unit is localized through t(...) in rep-voice-utterance.ts (the localize pattern: explicit literal t("…") calls keyed by the English unit name from packages/core/src/units.ts).

Local language​

The voice language follows the active app locale (useTranslation().locale). resolveVoiceLang (in @enode/core/text-to-speech) normalizes the locale to a BCP-47 tag the on-device engine actually supports — it tries the exact tag, then any installed tag sharing the base language ("de" → "de-DE"), and caches the result. When no voice is installed for the language, nothing is spoken (the feature degrades silently rather than reading in the wrong language).

Platforms​

One common implementation via @capacitor-community/text-to-speech:

  • iOS → AVSpeechSynthesizer
  • Android → android.speech.tts.TextToSpeech
  • Web → Web Speech API (best-effort bonus; gated on speechSynthesis)

The plugin module touches window at evaluation time, so @enode/core/text-to-speech loads it lazily via dynamic import() — importing it eagerly breaks Next.js static prerender (window is not defined).

speakValue uses QueueStrategy.Flush, so a burst of reps interrupts the previous utterance instead of queueing and lagging behind the athlete.

Rate & latency​

Speech rate is set per platform (VOICE_RATE_IOS / VOICE_RATE_DEFAULT in @enode/core/text-to-speech) because the plugin scales rate differently: iOS compresses it (0.1*rate + 0.4 on AVSpeech's 0–1 scale, max 1.0 — so the current 2.5 lands at ~0.65 for a brisk callout), while Android applies it straight to setSpeechRate (1.25, where 1.0 = normal). Both are tunable from device testing.

prepareVoice(locale) removes the first-rep cold start: the dynamic import, the supported-languages probe, and the engine spin-up + voice load would otherwise all land on rep 1. The tracking hook calls it the moment voice is switched on (well before the first rep) — it loads the plugin, primes the language cache, and runs a silent (volume 0) warm-up utterance. Spoken reps are also kept short (the unit only on the first rep of a set) to minimise per-rep speech time.

iOS audio session​

IOS_AUDIO_CATEGORY in @enode/core/text-to-speech selects the AVAudioSession category. The default "ambient" mixes with other audio and respects the hardware silent switch (a muted phone stays silent); "playback" is always audible but interrupts the athlete's music. This is the single knob to tune from device testing. Voice also shares the audio session with video recording (@enode/core/video, AVFoundation) — verify the two don't fight when recording a set.

Voice owner — one station speaks​

Several stations speaking over each other is meaningless, so exactly one station owns the voice at a time. Ownership lives in apps/tracking/src/app/workouts/today/voice-station.ts.

Every view able to speak registers itself with that module while mounted — live training columns keyed by their station id, feedback-training flows keyed by feedback:<sensorStationId> (the feedback: prefix is what keeps the two id spaces apart). The set of candidates is therefore assembled in the store rather than prop-drilled down from page.tsx. Two rules then decide the owner:

  • Some active station always owns the voice, so the common single-station case never needs a tap and closing the owner hands off rather than falling silent.
  • An owner that is still active keeps the voice when other stations open or close, so an explicit choice is never overridden.

Ownership is deliberately in-memory and not persisted: station ids are ephemeral crypto.randomUUID() values minted per session, so a stored owner id would never match anything on the next launch.

UI & persistence​

The on/off switch is the "Voice feedback" row in the Local settings sheet (local-settings-sheet.tsx, opened from the station sidebar) — device-local, governing both tracking modes. The active tracking view's "⋯" options menu (ExerciseDetailView.tsx) carries a second copy of the same toggle; both write one store, so they can't disagree.

The speaker symbol is a station picker, not a power button. It appears in both tracking views only while voice is switched on, showing SpeakerWaveIcon (lit) on the owning station and SpeakerXMarkIcon (muted) elsewhere. Tapping a muted speaker moves the voice to that station; tapping the lit one does nothing.

State is the voiceFeedback flag in @enode/core/training/flags, persisted in localStorage (enode.voiceFeedback, default off). The whole feature sits behind the reversible VOICE_FEEDBACK constant exported from use-rep-voice.ts, shared by both speaking views.

Feedback training: no load​

Feedback training runs without a workout, so there is no load to scale by and useRepVoice is called with loadMass: null. buildRepValueMap then omits loading-factor metrics (force, power, …) instead of leaving a per-kilogram value in place — an unscaled number would be wrong, not merely unlabelled. The feedback view only ever picks a load-independent focus metric (the first such entry of the coach's selection), so nothing spoken there is a load-dependent number, and spoken matches shown. A note under the mode name tells the coach that those metrics are measured in a workout.