Debrief Redesign
What Oura, Whoop and Muse actually do to make dense physiological data legible, and a three-tier visual system for EEG built on it, with every metric tethered to a personal baseline.
Bio-UX masterplan
What Oura, Whoop and Muse actually do to make dense physiological data legible, and a three-tier visual system for EEG built on it, with every metric tethered to a personal baseline.
What’s grounded and what isn’t
Part 1 comes from vendor documentation and one third-party design analysis, listed at the end. I have not inspected these apps directly, so I can’t verify pixel measurements, exact colour values, or animation behaviour. Where a figure comes from a third-party estimate rather than the vendor, it says so.
I also have no user research on Crown Debrief. Nothing here is validated preference, it’s inference from what three successful products in adjacent categories chose to do. Calling it “what users prefer” would be a claim I can’t support. The way to earn that claim is five minutes watching one person who has never seen it try to read a session.
Part 1 Industry benchmark analysis
Three products, three different answers to the same problem, and one pattern all three share.
The shared move: compression to a single learned number›
Whoop reduces “dozens of biometric signals to one Recovery score between 0 and 100.” Oura does the same with Readiness, assembled from nine contributors, resting heart rate, HRV balance, body temperature, recovery index, sleep, sleep balance, sleep regularity, previous day activity, activity balance.
The point isn’t the number. It’s that the user learns one scale, once, and then reads everything through it. Crown Debrief currently asks a person to learn six scales simultaneously, none of which they’ve seen before.
Oura: baselines as the entire product›
Every contributor is scored against your own history, not a population norm. Two windows are used: a 14-day short-term baseline for HRV Balance and Activity Balance, with recent days weighted more heavily, and a two-month long-term baseline for most others. Oura states plainly that it “can take up to two weeks for the Oura App to learn your average values.”
The wording is fixed and repeated everywhere: 85-100 Optimal, 70-84 Good, 60-69 Fair, 0-59 Pay Attention. One vocabulary across every metric, so the words do the interpreting the numbers can’t.
Their recent redesign collapsed five tabs into three, Today, Vitals, My Health. Today is described as “the ‘Top Stories’ page of a news app.” Vitals shows core metrics with individual baseline ranges and expandable contributors. My Health holds slow-moving long-term trends and periodic reports.
Whoop: semantic colour and hard tier separation›
A deliberately constrained three-colour vocabulary: “Green signals readiness and recovery. Red signals strain, risk, or low recovery. Yellow marks the in-between.” Colour carries meaning rather than decorating, so it’s learned once and applies everywhere.
The interface is dark because “coloured data points pop against black” and it’s gentler for morning and evening checks, a functional choice, not a style one.
The detail most worth stealing: their three tiers live on separate screens reached by deliberate interactions, not expandable sections on one page. Tier 1 is three numbers. Tier 2 is week-over-week trends. Tier 3 is raw biometric graphs. Someone who only wants tier 1 never sees tier 3 exists.
Third-party estimate, not vendor documentation: a design analysis puts the Recovery score at roughly 72 point, sized for arm’s-length reading.
Muse: the closest analogue, and the most cautionary›
Muse is the only benchmark actually reading EEG. Two things it does well: during a session, feedback is audio, not visual, weather sounds that shift with your state, so you’re not staring at a chart while trying to concentrate. Afterwards, the summary is a graph of time spent in calm, neutral and active states.
Three named states, not five numbers. That is the single most transferable idea in this document.
The cautionary half: Muse awards points and birds for time spent calm. That’s a real engagement mechanic and it works, but scoring someone’s brain and rewarding them for producing the “right” state is precisely the thing your safety notes rule out. Worth naming rather than quietly copying.
The four transferable patterns
1. Compress many signals into one number a person learns. 2. Tether every metric to a personal baseline with a fixed interpretive vocabulary. 3. Separate tiers onto separate screens, not accordions. 4. Name states rather than showing values, calm/neutral/active beats a decimal.
Part 2 Where the previous plan fell short
It was a layout fix for a system-level problem. Bluntly: it rearranged the furniture.
| Previous proposal | Why it’s insufficient |
|---|---|
| “Verdict at the top” | Correct but shallow. A sentence at the top is not compression, it’s a summary sitting above the same six competing numbers. Whoop and Oura don’t lead with a sentence, they lead with one figure the user has learned to read. |
| “Fold detail away on one page” | Directly contradicted by the benchmark. Whoop separates tiers onto distinct screens on purpose. An accordion still announces that fourteen more things exist and puts the burden of ignoring them on the reader. |
| “Add labels to the timeline” | Labels a line chart that shouldn’t be a line chart. Two overlapping noisy traces is an instrument readout. Muse’s answer, discrete named states over time, is a different visual object entirely. |
| “Plain language on every number” | Right instinct, no mechanism. Oura’s version is a fixed four-word vocabulary applied identically everywhere. Mine was “write a sentence next to each figure”, which produces fifteen inconsistent phrasings and no learnable scale. |
| Colour left as-is | The biggest omission. Teal for focus and amber for calm is decorative, the hues carry no meaning, so nothing can be read at a glance. Whoop’s entire legibility rests on colour being the message. |
| Baseline noted as a caveat | I put it in a closing section as a thing worth doing later. For Oura it is the product. It belongs in the data model and in every component, not in a footnote. |
Why numbers alone fail, psychologically›
A focus reading of 0.385 is not a hard number in the way that 38 kg is. It’s a model’s probability estimate on an arbitrary scale, and a person has no anchor for it, no lived experience of what 0.385 feels like, no cultural reference, not even a sense of which direction is good.
Given no anchor, people invent one, and they reliably invent the wrong one: they read it as a percentage and conclude they were 38% focused, which is both meaningless and mildly demoralising. Neurosity’s own documentation says anything above 0.3 is significant, so the intuitive reading isn’t just imprecise, it’s inverted.
A baseline replaces the missing anchor with the only one that’s actually valid: you, last week. “Twenty minutes more deep work than your usual Friday” needs no scale, no priors, and no neuroscience. It is also the only comparison that is defensible, because the between-person variation in these scores is large enough to make any population norm meaningless.
Part 3 The three-tier visual system
The semantic palette›
Four states, four hues, fixed meaning, used nowhere else in the interface. Note the deliberate departure from Whoop: their axis is good-to-bad, which is wrong here. Drifting is not a failure, it’s a state, so the palette runs cool-to-warm by engagement, not green-to-red by quality. The app should never tell someone their brain scored badly.
Semantic state tokens
Focused -st-focus · engaged
Settled -st-settle · calm, restorative
Drifting -st-drift · low engagement
No reading -st-none · hatched, never solid
“No reading” is rendered as a hatch, never a flat fill. Poor signal is the absence of a state and must never look like one, the one rule this palette exists to enforce.
Tier 1 · The macro: three seconds›
One number, one comparison, one sentence. The compressed metric is Deep Work: total time spent meaningfully above your own baseline.
Deliberately time, not a score. Minutes are a unit people already own, they can’t be misread as a percentage, and, unlike a 0-100 rating, counting minutes doesn’t grade someone’s brain, which keeps this the right side of the line your safety notes draw.
Tier 1 · session hero
2h 15m
Deep work · one session
+22 min vs your usual Friday
A strong morning, a heavy afternoon.
Your best stretch ran 09:44-11:59. Focus dropped after lunch and stayed low for 47 minutes.
The learning state: first ten sessions›
Baselines can’t be faked. Until roughly ten real sessions exist, the comparison slot shows what it’s doing instead of inventing a number, the same choice Oura makes when it says it needs two weeks.
Tier 1 · before a baseline exists
2h 15m
Deep work · one session
Learning what’s normal for you
2 of 10 sessions recorded. Comparisons switch on once there’s enough to compare against.
Tier 2 · The meso: the session, in states›
The line chart goes. In its place, a state ribbon: the session as a continuous band of named states, the same idiom consumer sleep tracking uses for sleep stages, which is the closest established solution to this exact problem.
Beneath it, two aligned lanes: a deviation strip showing how each slice compared with your baseline, and an activity lane carrying what you were doing. Vertical alignment between “what my brain did” and “what I was doing” is where insight actually comes from, and it’s impossible in the current design because the two live on different parts of the page.
The ribbon sits on a dark panel in both themes, following Whoop’s reasoning: the semantic hues need a constant ground to stay comparable, and they read strongest against near-black.
Tier 2 · state ribbon, deviation strip, activity lane
14:12 Drifting 28% below your usual afternoon
writing the parser lunch meetings review
09:00 11:00 13:00 15:00 17:30
The dashed segment is a break in recording, drawn as a gap. The hatched segment is poor signal. Neither is ever coloured as a state.
Tier 3 · The micro: exploration on demand›
Reached by a deliberate action, on its own screen, exactly as Whoop separates its deep-dive. Nothing here appears unless asked for: the raw focus and calm traces, band power, per-electrode contact, the state engine’s internals.
Desktop-appropriate interactions, not phone gestures: scrub along the ribbon with the pointer and a sticky data pill follows; drag across to select a window, which recomputes Tier 1 for just that window; click an event to open its detail with the note and activity fields in place.
Part 4 Component blueprints
4.1 · BaselineBandGauge›
The signature component, and the one that enforces the baselining rule structurally: it is unreadable without a baseline, so it cannot be shipped showing a bare number.
A horizontal track. Your usual range for this metric renders as a shaded band with hairline edges. Today’s value is a marker on that track. A person reads position, not magnitude, inside the band is normal, outside is not, and no scale needs learning.
BaselineBandGauge · three states
Deep work 2h 15m above usual
less your usual range · last 10 sessions more
Settled time 1h 02m typical
less your usual range · last 10 sessions more
Longest unbroken stretch 18 min below usual
less your usual range · last 10 sessions more
The fixed vocabulary›
Oura’s discipline, adapted. Five words, used identically for every metric, derived from where the value sits relative to your own distribution. Never “good” or “bad”.
| Word | Trigger | Colour token |
|---|---|---|
| Well above usual | z ≥ +1.5 | -st-focus |
| Above usual | +0.5 ≤ z < +1.5 | -st-settle |
| Typical | −0.5 < z < +0.5 | -st-none |
| Below usual | −1.5 < z ≤ −0.5 | -st-drift |
| Not enough data | fewer than 10 sessions | -st-none, hatched |
4.2 · StateRibbon›
Canvas, not DOM, a seven-hour session is thousands of slices. Existing classify() in core/stats.js already produces the states; this consumes them directly.
-
Binning: one bin per rendered pixel column, each taking the modal state of its rows. Modal, not mean, averaging state labels is meaningless.
-
Minimum segment: 4 px. Anything shorter merges into its neighbour, otherwise a noisy minute becomes visual confetti.
-
Poor signal: 45° hatch at 55% opacity, drawn as a repeating pattern, never a fill.
-
Recording gap: transparent with dashed 1 px edges. No interpolation across it, ever.
-
Height: 44 px. Tall enough to carry hue, short enough that the aligned lanes stay within one glance.
-
Ground: the dark panel token in both themes, so the semantic hues sit on a constant background.
4.3 · DeviationStrip›
Sits directly beneath the ribbon, sharing its x-scale exactly. One bar per bin, height proportional to that bin’s z-score against your baseline for that hour of day, not against the session, and not against a flat all-day average, because 3pm and 10am are not comparable.
Bars are coloured by the same semantic token as the ribbon segment above, so the two lanes read as one object. Before a baseline exists, this lane is hidden entirely rather than shown empty.
4.4 · ActivityLane›
Below the strip, same x-scale. Blocks of named activity, drawn as low-contrast chips so they never compete with the ribbon.
The distinction that matters: an activity spans time, a note marks a moment. Different storage, different interaction. Activities come from a small reusable tag set the user builds up, free text every session would make cross-session comparison impossible, and cross-session comparison is the entire payoff (“your deep work runs 40% lower in meeting blocks”).
4.5 · ScrubPill›
Pointer moves along the ribbon, a 1 px rule follows, and a pill tracks it carrying three lines: clock time, state name, and the baseline comparison in words. The pill clamps to the container edges rather than overflowing.
Motion: position transitions at 90 ms with an ease-out curve, fast enough to feel attached to the cursor, slow enough not to jitter. All transitions disabled under prefers-reduced-motion.
4.6 · Screen architecture
Tier 1 · lands here
Today
The hero number, the baseline comparison, two sentences, the suggestion. Three gauges. Nothing else. This is the whole app for most visits.
Tier 2 · one click
Session
Ribbon, deviation strip, activity lane, event cards with notes. Time selection recomputes the hero.
Tier 3 · deliberate
Detail›
Raw traces, band power, electrode contact, state engine internals, the guide’s retrieval. Separate screen, quiet link.
Three destinations, matching Oura’s collapse from five tabs to three and Whoop’s three tiers. This is the change that produced the shipped product: the developer panel’s five tabs became three screens, and the two that disappeared, Live and Diagnostics, were the two a consumer should never have been shown first.
4.7 · Build order›
-
Semantic token set and the fixed vocabulary Both palettes, the five words, the z-score thresholds. Everything else depends on it and it’s an afternoon’s work in CSS plus one function.
-
Cross-session baseline store Per-person, per-metric, per-hour-of-day, over the last ten sessions. Plus the learning state for when it isn’t ready. This is the load-bearing change, it’s a data model addition, not a UI one.
-
Deep Work metric and the Tier 1 hero The compression. One number, one comparison, one sentence.
-
StateRibbon on canvas Replaces the line chart. Binning, minimum segment merge, hatch and gap handling.
-
DeviationStrip and ActivityLane The aligned lanes, plus the activity tag store.
-
Scrub, select, and the three-screen split Interaction and the move of Live and Diagnostics behind the detail link.
Strategic improvements & unconsidered opportunities
Test it on one person before building any of it›
Everything above is inference from three products in adjacent categories. Sitting one person who has never seen Crown Debrief in front of the current version, and asking them to say what happened in the session, would take fifteen minutes and would be worth more than this entire document. It’s the highest-leverage thing available and it costs nothing.
Follow Muse’s real insight: don’t visualise during the session at all›
Muse gives audio feedback while you meditate because looking at a chart while trying to concentrate is self-defeating. The same logic condemns the Live tab outright, watching your focus score while working is a guaranteed way to lower it. If live feedback is ever wanted, ambient audio is the right channel, and the visual belongs entirely to the debrief afterwards.
Hour-of-day baselines beat session baselines›
Comparing 3pm to your all-day average will systematically mislabel every afternoon as a slump, because for most people afternoons genuinely are lower. Baselining per hour-of-day costs almost nothing extra to compute and removes an entire class of false conclusions. I’ve specified it in 4.3, but it’s worth flagging as a decision rather than a detail.
The activity tags are the real product›
The EEG is a sensor. The insight, “your deep work is 40% lower on meeting days”, comes from the join between brain state and what you were doing. Which means the tag set deserves more design attention than anything else here: a small controlled vocabulary, one tap to apply, consistent across sessions. Get that wrong and no amount of visualisation recovers it.
A weekly report, not just session debriefs›
Oura’s My Health tab exists because slow-moving trends are where the value accumulates. A Friday email, “your best hours this week were 9-11, and Tuesday’s meetings cost you an hour of deep work”, would probably be more useful than the session view, and it’s mostly the same computation over a wider window.
The honest risk in all of this›
Every technique here makes the data feel more authoritative, semantic colour, confident vocabulary, clean compression. The underlying signal is a consumer EEG probability score with real noise in it, and better design does not make it more valid, only more persuasive. The confidence badge and the hatched no-reading state exist to hold that line, and they should survive any future simplification pass.
Which of these would you like to pull into the build, and shall I start at the top of 4.7, or run the fifteen-minute test on someone first?
Sources. Oura: Readiness Contributors and the 2026 app redesign. Whoop: a third-party design breakdown, the 72-point figure is that author’s estimate, not vendor documentation. Muse: the app page. Neurosity thresholds from their Focus and Calm API docs. No code changed.
Explore more
This site covers what the documentation doesn't: the things I wish someone had handed me first.