Case study · 5 months · Dec 2024 — Apr 2026
An honest 5-month account of shipping a context-aware focus system. The research that reframed the problem, the seven prototypes we killed, the ethical review that changed the product, and what 140,000 daily users taught us about attention we couldn't have learned in a lab.
If you read only this
Focus tools fail because they demand focus to be used. We designed one that makes zero demands — it senses attention instead of asking for it. Seven prototypes died to reach this. The one that shipped has 8× industry-average retention. This case study is about how we got there, and what I'd do differently.
Role clarity
Case studies that claim everything are worthless. Here's what was mine, what was the team's, and where the lines blurred.
The explainer layer (Decision 02) emerged from a three-hour whiteboard session between me, Priya (eng lead), and Laila (ethicist). I wrote the spec. Priya pushed back on what was technically honest. Laila insisted the defaults be inverted. The final form was none of ours individually.
01 · The problem
The category was broken before we arrived. Six weeks shadowing 34 knowledge workers — engineers, designers, writers, researchers — watching how they actually used existing focus tools. Most were abandoned within eleven days.
The pattern was consistent. Existing tools asked the user to predict their own focus — to schedule it, commit to it, fight their impulses through friction. They worked on the assumption that a distracted person is a weak person who needs to be restrained.
But the people we watched weren't weak. They were responding sensibly to an environment constantly making claims on their attention. The tool meant to help was another thing making claims on their attention.
The reframe: anything that requires a decision to focus is already interrupting focus. We needed a system that asked nothing, showed nothing, and acted only when certain.
Pomodoro sessions abandoned mid-flow because a 25-minute timer ended just as the person got productive. Measured in 27 of 34 subjects.
Every time someone added Twitter to a blocklist, they developed a new habit of browsing it on their phone within 48 hours. The friction moved, it didn't disappear.
The most destructive moment wasn't the meeting — it was the 20 minutes after, spent re-assembling attention no tool recognized. 14k measurement sessions later, this became a core product feature.
02 · Research
Mixed-method study. Diary entries, shadow sessions, passive biometric logging, semi-structured interviews. Pre-registered hypotheses with UCL Interaction Centre. Raw data, not narrated anecdotes.
Six-week diary study. 1,820 focus entries. Four behavioral archetypes surfaced, all with distinct interruption-sensitivity profiles.
Of users given a mode picker, 91% used only the default. Manual mode-switching was a myth we were building for.
HRV coherence predicted measured deep-work sessions 4.2× more reliably than self-reported focus. Self-report was worse than chance after 3pm.
Traditional focus apps uninstalled within 11 days. The threshold was exactly: 'I notice it exists.'
A focus tool is successful Every interaction it requires is a tax on the thing it claims to protect.
We printed this on a 2m banner. It sat behind my monitor for 14 months. Any feature that violated it died in review — nine did.
03 · Design principles
Written on the studio wall. Every design review returned to them. Several features died against them. Not aspirations — filters.
The system must be more confident than the user would be, at every intervention. Ambiguity defaults to silence. Operationalized as: p(intervention helpful) > 0.82 before any action fires.
If an intervention would be noticeable as "something the app did," it is too loud. Dim, not darken. Delay, not block. Reduce, not remove.
Every intervention can be rewound and inspected. Three signals — exactly three — are always available. Fewer feels opaque. More feels defensive.
A tool that watches you must never be watched by anyone else. On-device. Append-only local ledger. Zero telemetry. Third-party audited.
When the context that justified the intervention ends, the intervention ends with it. No lingering states, no “are you still focusing?” prompts. The system exits without asking to be thanked. Measured: median exit-to-interaction was 0.3 seconds.
04 · Key decisions
A case study is only honest when it shows the decisions that could have gone the other way — and names what we traded for each.
We removed the focus slider entirely. The single scariest thing we shipped.
We lost the feeling of control to gain actual adoption. Power users complained. We let them.
The first three prototypes had a focus slider. Users could dial intensity from light to deep. Everyone on the team loved it. It tested well in the lab — 87% task completion in controlled sessions.
In the field, it died. Of 48 beta users given the slider, 91% set it to medium on day one and never touched it again. The control was a burden they never wanted. We removed it in v0.7, replaced it with inferred intensity, and engagement tripled within two weeks.
Fig. 02 · Slider became orb. Input became inference. Engagement 3.2×.
Every action can be interrogated. Built before the modes themselves.
Every intervention had to be reducible to exactly three signals. Several otherwise-good features became impossible. We shipped fewer actions, explained perfectly.
The ethics review caught us early. A system that watches you and acts on you must be explainable by default — not as a setting, not buried in a menu. We built the why view as the first feature, before the model itself, and refused to ship any intervention that couldn't populate it with exactly three signals.
Fig. 03 · The why view. Reachable from any surface via ⌘ ?
We hired a composer before shipping a button. Low-frequency textural shifts, not chimes.
Users with hearing loss or muted systems get zero audio feedback. We added haptic-only mode in v1.3 after advocacy feedback — took longer than it should have.
Mode transitions needed confirmation. Confirmation is traditionally a chime. A chime is an interruption. Four weeks on a sonic language tuned to be noticed only if you're already listening — so it never performs its presence to a focused mind.
Fig. 04 · The sonic language. Each transition a shape, not a sound.
05 · Prototype graveyard
Portfolios that only show what shipped are PR documents. The decisions that matter are in what got retired. Here's every major prototype — and the reason each one died.
A wall of charts showing historical focus patterns. Tested with 12 users. 2 of them opened it after day one. The data was interesting; looking at it was a job. Killed — replaced by a weekly single-paragraph email.
GPT-style nightly summary of "how your attention went today." Users hated it. Quote: "I don't want my laptop psychoanalyzing me." Killed — we were building surveillance we couldn't justify.
Covered in Decision 01. Killed because users didn't want to feel like pilots of their own attention.
Show friends how much deep work you did. The worst idea we had. The ethics review ended it in 20 minutes. Killed — and the question "does this shame anyone" became part of every review.
Consecutive days of deep work, with badges. Retained users 18% better in A/B. Still killed, because the retention was coming from fear of breaking the streak, not from the tool working. Fear-based retention violates principle 02.
"Hey Pulse, help me focus." Shipped to 500 users. Used an average of 2.3 times total. Killed — the voice surface demanded the kind of attention the product existed to protect.
Pulse would write focus blocks into your calendar. Users overrode it 78% of the time. The calendar is sovereign; we were squatters. Killed — replaced with suggested windows that the user accepts.
06 · Process
18 months. 7 major prototypes. 4 retired in favor of silence. The version that shipped was the smallest one.
Six-week diary study, 34 subjects. Pre-registered hypotheses. UCL working paper. Four archetypes, one thesis.
Six prototypes that gave the user control. All performed well in the lab. All failed in the field. We killed the entire branch.
Inference replaces control. The orb becomes the only surface. Engagement 3.2× in two weeks. First time the thesis felt true.
38MB binary. No marketing site for eight weeks. 1,200 private testers. Word-of-mouth to 18k daily actives.
Teams publish deep windows to each other's calendars. The social contract of focus, made visible. First team-tier revenue.
94% D30 retention. 4.9 on the App Store. Still no paid acquisition. $14M ARR.
07 · Design system + accessibility
All inference animations, orb pulses, and state transitions collapse to static in reduced-motion mode. Zero functionality is gated by motion.
Focus state colors meet 7:1 contrast. Explainer layer never drops below 8.2:1. Validated with 4 users with low vision in research.
Every sonic transition has a haptic equivalent. Hearing-impaired users get the same feedback fidelity. Took us too long — shipped in v1.3.
A hard rule: no screen introduces more than three new concepts. Applied to onboarding, settings, and the explainer. Enforced in code review.
Microcopy avoided idiom. "Deep work" translates. "In the zone" does not. Linguistic review with native speakers in all 12 languages.
The context model runs without network. The app works on a plane, in a basement, off-grid. No feature requires connectivity. Ever.
08 · Outcomes
All metrics are behavioral, not self-reported. Every number here has a how-measured note attached, and every claim is reproducible with the public datasheet.
Avg focus-depth per user, 60 days in. Measured via typing cadence + HRV coherence, not self-report. Control group: previous 60 days without Pulse.
Post-meeting recovery friction, 14k sessions. Time-to-re-engagement dropped from 19 min to 7 min. Measured passively.
8× category baseline of 11.8%. Matched-cohort analysis vs 5 comparable focus apps. Independently verified.
Third-party audited quarterly. Trail of Bits Q1 and Q3 2025. Reports public. No asterisk.
“I forgot it was running — and then I realized I'd been writing ”
09 · If you were going to challenge me
A case study I respect is one where the designer predicts the critique. Here are the five questions a senior reviewer would ask me, and what I'd say.
Possibly partially, yes. It's measured behaviorally (cadence, HRV) so it's not purely subjective — but placebo can still shift those. The more defensible metric is the 94% D30 retention, which is 8× baseline and impossible to explain with placebo alone. We've been careful not to overclaim the +41% as causal.
The Trail of Bits audits are public. The binary is reproducible-build verified. We'd prefer you didn't trust us — that's why we paid for the independent review. If you have a specific claim you want to verify, the network-capture methodology is documented in our public datasheet.
It's a fair critique and I thought about it constantly. The defense: Pulse never restricts. It dims, delays, and suggests. Any intervention is reversible in one tap, and the system exits the moment you override it. Paternalism requires restriction. We don't restrict.
The honest answer: the current training data skews toward knowledge workers at English-speaking companies. We know this. We disclose it. The v3 roadmap includes a multi-context training set and we're deliberately slowing shipping speed to make sure archetype coverage is broader before we expand markets.
The sonic language (ship with silence). Shared deep windows (v1 is solo). The consent ledger UI (keep the ledger, hide the UI behind a debug flag). I would not cut: the explainer, the diary research, or principle #2. Those are the product.
10 · Reflections
I would run the ethics review in week one, every time. The explainer layer existed because an ethicist made us answer a question we'd been avoiding: what right does a system have to act on someone it is watching? The answer shaped the entire product. It should not have taken a specialist to surface it.
I would not try to soften the kill. We spent three weeks trying to save the slider — fading it, hiding it behind a toggle, making it an “expert mode.” It didn't work. Features don't die gracefully; they die when you remove them. The version that shipped was better because we accepted that sooner.
I would trust the sound designer earlier. The sonic language is the single most-praised detail in user feedback, and it took us longest to commission. The cost of hiring a composer in month one would have been lower than four months of placeholder chimes.
I would be less proud of the 0. The “zero bytes leaving the device” metric is an engineering achievement and a marketing gift, but it slightly misrepresents the real ethical work — which is the consent ledger and the explainer. The 0 is easier to tweet. The ledger is what matters.
I would not have assumed retention would come from features. The thing that kept users wasn't any one feature; it was that Pulse never made them feel watched. That is a design outcome, not a feature decision. I wish I'd known that earlier.