Error handling
The app routes every async failure through one error engine in
@enode/core. The UI reacts to a disposition, never to a raw status code.
This page is the mental model plus a cookbook so new code (and AI agents) extend
the system by editing tables and seams, not by inventing patterns. See
ADR 0002 for the rationale.
This page covers failures. For when to proactively confirm an outcome — success toasts, and how failures map onto toast vs inline vs dialog — see user feedback (toasts & confirmation).
Where the code lives
Error handling spans the three layers by design — pure logic in @enode/core,
presentational surfaces in @enode/ui, app wiring in @enode/tracking:
| Layer | Location | What's there |
|---|---|---|
Logic (@enode/core) | errors/ | the engine: classify, status-table, action-table, action-resolver, error types, the report/handler/resolver seams |
diagnostics/, connectivity/, store-kit/ | breadcrumbs + crash capture + the foreign-DOM guard, the online/offline store, the load-status (RequireData) backbone | |
api/client.ts | apiRequest (throws + classifies + resumes redirects) and rawApiRequest | |
Surfaces (@enode/ui) | feedback/ | InlineAlert, Toaster/toast, BlockingErrorDialog, StalenessHint, the support-report sheet, RequireData |
App wiring (@enode/tracking) | app/error-handling/ | ErrorEngineInit (registers the seams at startup), use-error-presenter (toast/inline/dialog), OfflineBanner, and redirects/ (the 310–313 resolution dialogs + stores) |
app/error.tsx, app/global-error.tsx | the route + root error boundaries (Next requires these at the app root, so they stay there) | |
app/debug/errors/ | the dev screen that triggers every state by hand |
The portal (@enode/portal) mirrors the app-wiring layer with its own slim
app/error-handling/ErrorEngineInit.tsx (no SSE escalation) and its own
app/error.tsx / app/global-error.tsx boundaries — the segment boundary
renders through the shared FullStopState.
The four dispositions
A pure classifyError(error) (packages/core/src/errors/classify.ts) maps any
thrown value to one ErrorDisposition:
| Disposition | When | Surface |
|---|---|---|
offline | offline / DNS / timeout / 408 | ambient OfflineBanner + toast, auto-retry; never tears down |
graceful | 400/403/404/409/422/429, 5xx, unknown | InlineAlert (forms) or toast + retry; screen stays |
redirect | 401 → re-login; 310/311/312 → a resolution dialog | a guard navigates (401) or a dialog resolves + the request resumes (310–312) |
blocking | unrecoverable / escalated critical write | BlockingErrorDialog / error boundary; always shows userMessage |
Two ambient signals are not dispositions: connectivity (online/offline) and per-store staleness — see offline & sync. They escalate to a real error only when a view has no usable data at all.
Policy lives in tables
status-table.ts— HTTP status →{ disposition, retryable }. The single source of HTTP policy. Change a row to change behaviour; tests read the table.action-table.ts— HTTP status → aRequiredAction(e.g. 401 →reauth). Checked before the status table.
The error objects carry facts (status, reason, userMessage, kind);
the classifier derives the disposition. The backend userMessage is the only
string shown to users — reason is developer-facing and never displayed.
Resumable redirects (310/311/312)
The backend's Enode-specific redirects are an interruption, not a failure: the
request pauses, a resolution dialog appears, and once resolved the original
request resumes (mirrors the iOS ENDialog flow — these mostly fire at login).
apiRequest(api/client.ts): whenactionForError(status)returns a resumable action (310 →resolve-device-limit, 311 →complete-onboarding, 312 →resolve-invitations), itawaits the resolver seamresolveRequiredAction(...)(errors/action-resolver.ts) and, on"resolved", retries the same request (execute(path, options, attempt+1), capped atMAX_ACTION_RETRIES)."cancelled"/ no resolver → throwsActionRequiredError. 401 is unchanged (destructive logout viaAuthGuard); 313 stays unmapped (the definition-editing layer owns it, which the web app doesn't have yet).- UI:
app/error-handling/redirects/ActionResolverInit.tsxregisters the resolver →action-dialog-store.ts(queues/coalesces) →ActionDialogHost.tsxrenders the dialog for the kind: DeviceLimitDialog (remove a session viaremoveDeviceSession+ the body'ssessionRemoveToken), InvitationsDialog (confirmInvitations; every invitation is sent back with an explicitacceptedflag — unselected meansaccepted: false, not omission), OnboardingStubDialog (311 placeholder — no onboarding screen yet, so it cancels). Resolving settles the store promise, which unblocksapiRequest. The invitations sheet refuses backdrop/Escape/Back, so a 312 can only be resolved by answering it. Resolution must stay synchronous with the store — awaiting a request beforeresolveActiondeadlocks, because that request's own redirect coalesces onto the entry still at the head of the queue. Post-resolution access checks live in the auth flow instead (see auth.md). - New
apiRequestoptionbearerTokensends a one-off token (thesessionRemoveToken) instead of the stored session token. - Try it:
/debug/errors→ fire 310/311/312 → "Resolve in app UI" opens the real dialog populated from the response body (no live login needed).
313 (shared-definition conflict) is a save-call-site concern, not a resume
Unlike 310–312, a 313 resolution doesn't replay the same request — the user's
choice changes which endpoint the save hits (iOS WorkoutModifyModel). So
313 is resolved where a schedule is saved, not in apiRequest, and actionForError
leaves it null. The pieces are in place for when the schedule-save flow lands:
scheduleEndpoint(mode, resolution)(api/workouts.ts, unit-tested) — the endpoint table: base path first, then*_definition_update(edit shared) or*_definition_copy(private fork).requestSharedConflictResolution({ canEditShared })(app/error-handling/redirects/shared-conflict-store.ts) +SharedDefinitionConflictDialog+SharedConflictHost— the choice dialog ("Edit for everyone" is write-gated / "Save a private copy" / Cancel).
The call-site pattern a future saveSchedule uses:
async function saveScheduleWithConflictHandling(body, mode, canEditShared) {
try {
return await apiRequest(scheduleEndpoint(mode, null), {
method: "POST",
body,
});
} catch (e) {
if (!(e instanceof ApiError) || e.status !== 313) throw e; // not a conflict
const resolution = await requestSharedConflictResolution({ canEditShared });
if (!resolution) return null; // user cancelled
return apiRequest(scheduleEndpoint(mode, resolution), {
method: "POST",
body,
});
}
}
Status: the web app has no schedule-save-to-server flow yet (
api/workouts.tsis read-only;WorkoutEditingFlowedits locally and only starts a session). So 313 has no live trigger — the dialog + endpoint table are ready, and/debug/errors→ 313 → "Resolve conflict" demos the choice → endpoint mapping. (403 definitionAccessDenied is just a graceful error — no special handling.)
"Exercise already exists" is a result, not an error
POST /exercises rejects a composition (base + equipment + the exact variation
set) the user already owns with 400 + reason: "Exercise already exists".
That is a normal thing for a user to do and each caller resolves it differently,
so — like 313 — it never reaches the presenter: createExercise
(api/exercises.ts) converts it into a typed ExerciseCreateResult
({ status: "ok", exercise } | { status: "duplicate" }) and rethrows everything
else. The exercise-creation wizard branches on the result, recovers the exercise
it collided with via findExerciseByComposition (exercises/duplicate.ts — the
rejection identifies nothing), and hands it to its host through onDuplicate:
the tracking picker selects it, the portal catalogue opens its detail drawer.
Detection lives in one predicate, isExerciseAlreadyExistsError
(exercises/duplicate.ts), because the reason string is the only signal the
response carries today:
export function isExerciseAlreadyExistsError(err: unknown): boolean {
if (!(err instanceof ApiError) || err.status !== 400) return false;
return err.reason?.trim().toLowerCase() === "exercise already exists";
}
The status alone is not enough — a plain 400 also covers ordinary validation failures. And since the backend localizes
reason, a non-English client falls through to the generic graceful path (a toast carrying the backend's own message). When the backend adds a machine-readable error code to the response object, match that here instead: this predicate is the only place to change.
Cookbook
Every extension point is a table, registry, or seam — no new patterns required.
- Handle a failure in a component: in a screen,
const { present } = useErrorPresenter()thentry { … } catch (e) { present(e) }(toast auto- chosen);present(e, { inline: true })for a form-levelInlineAlert; or let it throw to hitapp/error.tsx. Outside a screen,presentError(e, t)or the corehandleError(e). - Change how a status code behaves: edit one row in
status-table.ts. - Add a required-action / redirect: add a
RequiredActionvariant (errors/required-action.ts), anaction-table.tsrow, anaction-routesentry, and a resolution screen. Nothing else. - Record context for support:
addBreadcrumb({ cat, msg, data? })at the meaningful moment (@enode/core/diagnostics/breadcrumbs). - Send to support:
openSupportReport(error?)from any surface, orreportToSupport(buildSupportReport(origin, { error }))directly. - Swap telemetry / transport:
setErrorReporter(fn)/setSupportTransport(fn)once at startup (done inapp/error-handling/ErrorEngineInit.tsx).
Surfaces (@enode/ui/feedback)
InlineAlert/ErrorAlert— graceful inline message (forms), variants error/warning/offline/info, optionalonRetry/onReport/dismiss.toast+Toaster— transient graceful/offline messages, callable from anywhere. For an application error (4xx/5xx) both Try again (primary) and Report (secondary) are offered; a transientofflinefailure drops Report. Thepresent/presentErrorhelpers wire this viatoastActions(...).BlockingErrorDialog— full-stop modal, always showsuserMessage+ “Contact support”. NOT for 401 (the redirect handles that).FullStopState— the whole-screen full-stop that replaces the content area (vs the modal above): a haloed emblem, a faint brand mark, one primary action + a quiet support link. Presentational (caller passest(...)strings); fills its parent's height, so place it in a height-bearing region.RequireDatarenders it forerrorVariant="block".StalenessHint— subtle “couldn’t refresh” pill over cached data.SupportReportSheet— one-tap report form (opened viaopenSupportReport). Besides Send (POST to the token-protected backend relay viaapiRequest), it always offers Copy report — the same JSON payload onto the clipboard. Copy is the delivery channel when Send structurally cannot work: pre-login (no token for the relay) or with the API unreachable — exactly the situations a report is most needed in. The user pastes it into any support conversation.
The login screen reports too
Auth failures used to bypass the engine entirely (no telemetry, error object
dropped at the reducer). Now authErrorInfoFor carries the original error as
AuthErrorInfo.cause, the auth steps call reportError(err, { source: "auth" }),
and the login ErrorAlert passes onReport={() => openSupportReport(error.cause)}.
The auth alert also shows userMessage ?? reason (the backend puts its
localized message in reason — same rule as backendMessage), so a real
backend rejection is no longer rendered as the generic copy.
What a report carries for a transport failure
DiagnosticsSnapshot.error.causeChain serializes the cause chain under the
reported error (depth-capped, cycle-safe) — for a NetworkError that is the
raw fetch rejection (TypeError: Failed to fetch, an SSL message, …), the only
artifact that distinguishes failure classes when the request never produced a
server-side trace. snapshot.client pins the identity the app believes it has
(xUserAgent, appVersion, isNative) plus the raw navigator.userAgent
(names the exact WebView build on Android). The http breadcrumbs of failed
requests carry a cause string for the same reason. No secrets: api-key and
tokens are never serialized.
All use the --feedback-* design tokens — never raw red. Colours flip with the
colour scheme automatically (no dark: variants).
Connectivity failure on a data screen: skeletons + a card, not a takeover
When a data screen's own load fails on the network, prefer keeping the loading
skeletons and floating a card above them to a full-screen stop — a whole-screen
red surface reads as "everything's broken", and zeroed summary cards misread as
empty data when it's really the connection. The dashboard is the reference
(app/dashboard/page.tsx): on error it keeps loading skeletons and renders a
ConnectionCard (components/notices/). Every portal data screen follows the
same pattern on a dataless failure (no cache to show): Dashboard, Planner,
Devices, Users, Exercises, Workouts, Tags, History, Insights, Profiles, and
Notifications each render the card above their still-skeletoned content with a
working Try again. (A failed refresh over usable cache stays a
StalenessHint — see the store-kit freshness doctrine.)
connectionModeFor(error, browserOnline)(@enode/core/connectivity) names the cause —offline(device has no network, viauseBrowserOnline— the device-only signal, distinct fromuseIsOnline's reachability),unreachable(backend didn't answer), ormaintenance(a503, "back in a minute").ConnectionCardrenders it prominently (tone-tinted border + accent, a haloed emblem, a soft bloom) with Try again + Contact support. It's a fatal-ish state, so — unlike a stacked notice — it can't be minimized or dismissed; it clears when the load recovers.- Early surfacing: pages decide visibility through
useShowConnectionCard( loading, failed)(exported next to the card) — it also shows the card while a load is still running onceuseIsOnlinesays we're down. That reachability signal is fed by everyapiRequestoutcome and the notifications SSE stream's heartbeat/drop detection (NotificationStreamHost→realtimeSubscribe→markOffline). A single stream drop reconnects silently — deploys, proxy idle timeouts and stale watchdog timers after laptop sleep are routine — and only a reconnect attempt that also fails counts as offline evidence, so a dead backend still surfaces a few seconds after the stream drops (instead of after the 15s request timeout) without the surfaces flashing on every routine drop. Caller-driven aborts (unmount, navigation) feed nothing. The card self-clears when the signal recovers and the pending load resolves normally.
Non-fatal, dismissible notices (launch announcements like migration/insights)
instead live in the sidebar NoticeStack (components/notices/): at most one
expanded (accordion), each dismissed on its own (persisted), up to three stacked
then a carousel. Keep those out of a data view's body — the body is for its own
data and status.
Reserve the full-screen FullStopState / RequireData errorVariant="block" for
a view that genuinely can't render anything (missing essential reference data),
not for a transient connectivity blip on a screen that can still show its frame.
Diagnostics & support
app/error-handling/ErrorEngineInit.tsx wires the engine at startup: the telemetry reporter,
the UI dispatcher (so handleError from anywhere shows a toast), the support
transport (the backend relay), and the global window.onerror /
unhandledrejection handlers. Unhandled crashes auto-submit a report silently
(consent-gated, deduped, auth-gated); handled errors and general feedback are an
explicit one-tap report. Reports are vendor-neutral — see
offline & sync and ADR 0002.
Browser page translation
Browser page translation (Safari, Chrome) replaces the text nodes React
rendered. Without protection, React's next removeChild or insertBefore on
such a node throws NotFoundError mid-commit, and the route boundary replaces
the whole screen on every update.
ErrorEngineInit therefore also calls installForeignDomGuard
(@enode/core/diagnostics/foreign-dom-guard). It makes those two calls tolerate
a node that something outside React has moved: the removal is skipped, and the
insertion appends. The first guarded call of each kind leaves an error
breadcrumb (DOM guard: …) in the next support report. A translated page can
still show stale text, but it no longer crashes. See
ADR 0013.
The guard prevents the crash, not the drift. Where exact values matter, keep the
per-site measures: render a caption as one text node (captionLine), and mark
numbers and units translate="no" (as FitValue does).
Requiring data a view can't work without
A failed request is a disposition (toast / inline / retry). A view that is
missing essential reference data is a different problem: it can't render
correctly, so it should show a loading or error state instead of a broken screen.
Model that with store-kit + the <RequireData> gate, not ad-hoc ?. chains.
-
Each prefetch/lazy store embeds a
createStoreStatus()controller and callsmarkLoading()/markReady(hasData)/markError(err)around its fetch (seeexercises/store.tsfor the reference wiring). Views read it viauseStoreState(controller)/useFreshness(controller). -
<RequireData>(@enode/ui/feedback/require-data) gates a view on one or more requirements and renders the children only when every one is satisfied — else a loading placeholder (still settling) or an error surface (a load failed, or returned empty/insufficient). The “expected count” is each requirement'ssatisfiedpredicate, e.g.categories.length >= 1:const metrics = useStoreState(metricsStatus);<RequireDatarequirements={[{state: metrics,satisfied: categories.length >= 1,label: t("metric categories"),},]}onRetry={() => loadMetrics()}errorVariant="block" // whole screen needs it; "inline" for a degradable sectionshowStaleHint // present-but-stale → a hint over the children, not an error><CategoryList categories={categories} /></RequireData>; -
The decision is the pure
combineReadiness()(unit-tested). Rule of thumb, matchinguseFreshness: a view is only an error when it is dataless — if you can work with cached/partial data, makesatisfiedtrue and show a staleness hint instead of tearing the screen down. -
The error-debug screen has a live demo (“App UI states → Required-data gate”).
Status note:
store-kitis now wired into the eager reference stores —exercises,metrics,text-content,display,tracking,equipments,exercise-variations,tags,data-package-types— each exporting anxStatuscontroller (e.g.metricsStatus). A view gates on one withuseStoreState(metricsStatus)+<RequireData>. The store-side change per store iscreateStoreStatus()+markLoading()/markReady(size > 0)/markError()around the fetch +reset()inclearX. No server change — the gate is the client comparing what it needs to what the existing response actually returned.
Debugging error states (dev screen)
app/debug/ is a hidden developer area, gated as a whole by debug/layout.tsx
and split into segments behind a shared chrome (back button + tabs): Errors
(/debug/errors) and Style guide (/debug/styleguide). Add a route + page to
add a tab (debug/DebugChrome.tsx). The bug-icon entrance (next to the gear on the
Today header) and the routes are gated by isDebugViewEnabled() (debug/flag.ts):
NEXT_PUBLIC_ENABLE_DEBUG_VIEW="true" always shows it, "false" always hides it,
unset → dev only (production 404s).
The Errors segment triggers every backend error UI state on demand — one
button per /api/debug/errors/* endpoint, grouped
into collapsible sections (Sanity, Custom 3xx, EnodeAbort, Arbitrary Vapor Abort,
Uncaught). Each fire renders a card with the final status, wall-clock duration,
pretty-printed body, the production classification, a structured drill-down for
the custom 3xx bodies (310–313), and a runnable copy-as-cURL. A top-of-screen
Accept-Language switcher (en/de/fr) re-fires the /enode/* endpoints to verify
the localized userMessage, a masked Authorization strip tells prod/staging/
local apart, and a Fire all button sweeps every endpoint into a summary table.
-
See the real UI states — every fired error can be replayed through the actual presenter: a per-card "Show in app UI" / "As inline alert" button (or the "Auto-present on fire" toggle) calls
useErrorPresenter().present(...), so the genuine toast /InlineAlert/BlockingErrorDialogappears exactly as in the app. An "App UI states" catalog panel additionally triggers every surface directly with representative data — toasts, the globalhandleErrordispatcher, inline alerts (all variants), the staleness hint, the offline banner (markOffline/markOnline), the blocking dialog, the support report sheet, and the segment error boundary — so each state can be eyeballed without a backend. -
Gating — the route lives behind
isDebugViewEnabled()(app/debug/flag.ts): always on in dev; a production static export 404s the route unlessNEXT_PUBLIC_ENABLE_DEBUG_VIEW=trueis set at build time. The entrance (a bug icon next to the gear) is gated by the same flag. Reachable at/debug/errors. -
Real pipeline, no teardown — every call goes through
rawApiRequest(packages/core/src/api/client.ts), the non-throwing sibling ofapiRequest: same headers, bearer auth, query serialization, timeout + caller-abort, connectivity signal and breadcrumbs, but it returns the status/body as data so 3xx/4xx/5xx can be inspected instead of thrown. The screen reproduces the production classification by running the realclassifyErrorover a synthesized error, without dispatching the app-shell navigation events — so firing the 401 endpoint doesn't log you out mid-debug. The "treat 310–313 as success" toggle flips whether the custom 3xx render as classified errors or as plain inspectable responses.
Tests
Pure logic is unit-tested (errors/*.test.ts, api/with-retry.test.ts,
store-kit/status.test.ts, the debug-screen helpers
app/debug/errors/{curl,result}.test.ts, and api/raw-request.test.ts): the
classifier across every disposition, the status and action tables, the body
parser, retry/backoff, the reporter/dispatcher seams, the non-throwing
rawApiRequest, and the debug cURL/normalization helpers. Run npm test.