Designing calm.
A progress review.
Validating an AI-adaptive music product for anxiety relief, from inside TDMusic — a profitable AI music-distribution company. This briefing walks the design-thinking journey in order — frame → empathize → define → ideate → prototype → test → decide — and every phase hands its output to the next. Completed research is clickable; the decision logic is live.
Research question
How can an AI music-distribution company use corporate entrepreneurship, data-driven demand identification and AI-enabled production to build a scalable, evidence-backed music-for-wellbeing product — and can design thinking verify both the need and our ability to serve it?
What is new since 10 Jul
A full demand-side + product-side survey (7 blocks, pre-registered), a Bayesian decision engine replacing the point-score matrix (priors → evidence → posteriors → gates), the journey re-cut so each phase's output visibly feeds the next — plus a methodology v2, a focus-group kit, the survey results and the biometric/ML research design.
Done · in progress · next — the whole phase on one screen
Design-thinking phase (≤ 3 months, $10–20k) from the 10 Jul supervision to the decision memo. Green = complete, coral = running now, dashed = scheduled, dotted = after the phase. Dates after today are planned.
- ✓Framing & asset inventory · 5W1H requirement frameJul
- ✓6 interviews · 200-row netnographyJul
- ✓6-angle literature & market sweep (fact-checked)9 Jul
- ✓10-product competitor teardownJul
- ✓Supervision briefing #1 (scope, method, tension)10 Jul
- ✓Convergence matrix v1 → v2 Monte-Carlo re-scoreJul · 17 Aug
- ✓Concepts C1–C3 → belief register with priors & LRsJul · 17 Aug
- ✓Lilt clickable prototype · landing A/B built · WoZ kitJul
- ✓Product blueprint v1 — functions ↔ data capture ↔ method; 8 screens; interface matrix; library tags3 Sep
- ✓Survey v2 instrument (7 blocks, pre-registered) → analysed, n = 134Aug
- ✓Focus groups FG-1 · FG-2 with blind stimulus testAug
- ✓Methodology v2 — mixed methods, identification, measurement error, power17 Aug
- ✓Bayesian decision engine — priors → LRs → posteriors → gates (live)17 Aug
- ✓Biometric signal & ML research design (E0–E4, architecture)18 Aug
- ●Landing-page A/B live — waitlist conversion (≥ 5% → LR 2.5 on A3 · A4)→ 20 Sep
- ●Legal check — adaptation / derivative rights on ≥ 50 tracks (masters + publishing) → A5→ 12 Sep
- ●Pilot recruitment from survey opt-ins (Apple Watch / Oura owners) · Polar H10 sub-study prep · RA scripts→ 31 Aug
- ○E0 calibration — Watch vs Polar H10, n = 8: ICC, MAPE, reliability λ25 Aug – 6 Sep
- ○E1 efficacy pilot — 3-condition within-subject crossover, 20–30 × 3 sessions; STAI-S primary, residual HR / RMSSD secondary → A2a · A2b · A1a1 – 25 Sep
- ○Bayesian update — pipeline → posteriors → pre-registered gates → verdict25 Sep – 2 Oct
- ○Decision memo — PROCEED / PIVOT / KILL, concept selection (C1-neutral · C1-real · C2 · C3), supervision #2early Oct
- ○Dissertation write-up — methods, Empathize/Define, Ideate/Prototype, Test/Results, Discussion & roadmapSep – Nov · date tbc
- ○Post-phase (if PROCEED) — MVP scoping; E2 micro-experiments → E3 bandit personalisation → E4 self-generating library2027
The whole project on one line each — why · who · what · where · when · how
The classic requirements frame, filled from the evidence. Highlighted words are the parts the design-thinking phase must still prove.
Innovating from the core, not from scratch
A corporate-entrepreneurship project: a new action-research cycle that points the company's proven engine — demand signal → AI production → measured library — at its core asset (music), in a vertical with real clinical evidence.
net income $2.2M · valuation $70M
China distributor · Tier-1 YouTube · top TME supplier
peer-reviewed recommender (PeerJ CS, 2021); ~10× marketing ROI claim
tracks, majority owned/licensed · 220+ DSPs incl. Peloton, Tesla
Ansoff
Diversification-lite: a new customer need (health) served with an adapted product (music) on existing assets.
Three Horizons
H1 distribution funds H2 (this project); H3 is "music as a measured, closed-loop intervention".
Unfair advantage
AI music production + owned rights + worldwide distribution — a pure startup has none of these.
Insider action research, run through design thinking
As founder-CEO I sit between researcher and practitioner, so the project runs as a cyclical action-research process. Within this cycle, design thinking is the method — the d.school's five modes, paced by the Double Diamond — and decision analytics (priors, likelihood ratios, pre-registered gates) is how the Evaluate step is kept honest. Full methodology — mixed-methods design, identification strategy, measurement error, power →
The journey — each phase's input, method and hand-off
The evidence — click into any deep dive
Six streams, triangulated — interviews, netnography, literature & market, competitors, the survey (instrument + analysed results) and two focus groups. Each card opens the underlying data, quotes and citations; the triangulation matrix fixes what each stream is allowed to say.
Interviews
1 expert (dementia) + 5 users/caregivers. Full reconstructed records, key findings, verbatim quotes.
Netnography
200 coded data points from App Store & forums (Kozinets). Theme frequencies, best quotes, patterns.
Literature & market
Clinical meta-analyses, HRV/cortisol, market size & prevalence, regulation, wearables — fact-checked.
Competitor teardown
10 products, verbatim health-claims audit, pricing, wearable integrations, feature-gap map.
Survey v2 — demand & product side
7 blocks · who / why / when / where / what / how / how much · GAD-2, PSS-4, TAM, Kano, ODI, Van Westendorp · every item pre-registered to a decision node.
Focus groups
2 × 6–8 (US/EU online · China offline): moment mapping, blind audio stimulus test (real song vs neutral vs voice), concept & demo think-aloud, measurement, price.
Survey results
The pre-registered analysis run end-to-end — need, ODI opportunity scores, content by state (McNemar), wearables, TAM, Kano, Van Westendorp, purchase intent — and the LRs each threshold triggers.
How the streams fit together
Literature sets the priors; interviews, netnography and focus groups update them qualitatively (shrunk); the survey quantifies (n ≥ 100); the pilot tests efficacy. The roles are fixed in the methodology; each stream is an evidence row in the decision engine.
The convergence, scored — under uncertainty, and re-scored after the evidence
Three candidate populations entered the funnel; a weighted matrix chose the exit. v2 turns each cell into a range, jitters the weights and re-scores anxiety after the interviews: still first in > 99% of 5,000 draws. Full rationale → · Interactive Monte-Carlo →
Our unfair advantage points one way; the acute user need points the other.
In Empathize, every acute-anxiety interviewee rejected the "real songs you love" idea and asked for featureless, adaptive, neutral sound — disconfirming evidence for our core assumption, surfaced before building. The decision engine now carries this as two beliefs instead of one:
Neutral, non-melodic, adaptive sound — Endel's territory. Does not lever the catalogue.
Real licensed music + artists + AI engine — our moat, but maybe wrong for panic states.
Resolution under test: segment by arousal state — neutral adaptive sound for acute/sleep-onset; familiar real music for lighter daytime wind-down. Survey block C and the pilot A/B are pre-registered to settle it.
Hybrid-work professional, 28–40
Owns a smartwatch, self-medicates with music, won't see a therapist. Secondary: sleep-anxious new parent. Out of scope: diagnosed severe GAD.
Measurable calm in minutes
"Stressed hybrid workers who already use music to cope need measurable calm in minutes, because meditation apps demand effort and generic playlists aren't tuned to their state."
Prove it · don't make it worse · zero effort
HMW use the wearable the user already owns to prove it's working? HMW guarantee nothing jarring, no lyrics that pull you in? HMW make it one gesture? (survey B9 ranks these.)
15+ ideas, converged to three concepts — each with its riskiest belief
Adaptive "calm" app
Real songs (wind-down) or neutral sound (acute) re-shaped in real time to the listener's heart rate; shows the measured result after each session.
Adaptive-audio SDK
License the catalogue + adaptation engine to wearable & hardware brands. Sidesteps consumer-payment friction and the retention cliff.
Artist "calm" line
Artist-branded calm content through existing distribution — lowest cost, tests demand for real-music calm with no app.
The assumption map, now a belief register with priors and evidence — open any chip in the decision engine:
Deliberately cheap prototypes — one per riskiest belief
The point is to learn, not to build. The interactive design-thinking map is here.
Lilt app — screen recording
Settle → pre-rest → session (no numbers) → grounding → check-in → result → Unwind via Spotify embed → history. Recorded from the live build; heart rate simulated. Open the app ↗
Lilt — product blueprint
Two-mode app (Settle / Unwind): 8 screens, journey map, data-capture matrix (HealthKit · Oura · Polar), library tags, method → interaction mapping, 4-week plan.
Lilt — clickable prototype
5 screens: connect wearable → check-in → adaptive session → measured result → paywall. Think-aloud n = 5–8.
Landing-page smoke test
Two positionings, live waitlist + poll. Conversion ≥ 5% is a pre-registered LR of 2.5.
Efficacy pilot kit
Curated adaptive playlist vs neutral sound + Apple Watch/Oura; a human plays the algorithm. No code to test the effect.
Six tests, each pre-registered as a likelihood ratio — two are in
Survey v2 (n = 134)
Real song chosen by 24% in the acute scenario vs 47% for wind-down (McNemar p < .001); owners who would connect 57%; TAM BI top-2 25%; K1/K2 attractive / one-dimensional; PSM range [$5.12, $8.55] contains $6.99; purchase intent 35% (US/EU); bundled + employer 33%.
Focus groups (2 × 6–8)
Five moments of need; blind stimulus test — neutral sound wins the acute/night moment (10/13), the calm real song wins wind-down (9/13); "the number cuts both ways" → adapt silently, show the result after; subscription fatigue → bundled preference.
Efficacy pilot (E0 calibration → E1)
Within-subject, three conditions in a Latin square — T1 adaptive neutral · T2 adaptive real song · C active control (the participant's own relaxing playlist) — 20–30 people × 3 sessions; STAI-S primary, residual HR / RMSSD secondary; ANCOVA-form mixed model + Bayesian re-analysis. Positive → LR 4; null → 0.3. Design →
Landing A/B
Two positionings, live waitlist. Visitor → waitlist ≥ 5% is a pre-registered LR of 2.5 on both engagement and willingness to pay — the only behavioural WTP signal in the phase.
Legal check on ≥ 50 tracks
Adaptation / derivative rights, masters and publishing separately. Decisive: LR 8 if feasible, 0.15 if not. A neutral-sound acute product does not need it; the wind-down real-music mode and C3 do.
Adapt silently, show the result after
Focus groups and the survey's "measured result" item agree: the after-result is the credibility hook, but a live heart-rate line during a session can itself raise anxiety. Carried into the prototype as a "hide numbers" default for the acute mode. Signal & ML design →
The phase ends in a decision, not a pitch — and the decision is a model you can argue with
Priors → evidence (with likelihood ratios and sources) → posteriors → gates. Everything below is live from the same numbers as the decision engine; toggle a survey or pilot result there and this verdict changes.
A1a is far under the pivot line (≈ 7%) → the acute product is neutral adaptive sound; A1b (≈ 80%) keeps the real-music catalogue as the wind-down mode and the C3 probe. Engagement (A3 ≈ 72%) and willingness to pay (A4 ≈ 82%) now clear their gates. What stands between here and PROCEED is the efficacy pilot — the A2 gates are set so that literature alone cannot clear them — and the legal check. Open the gates →
Methodology v2 — how the evidence is made to count
Convergent mixed methods · triangulation matrix (who may say what) · within-subject crossover with an active control, ANCOVA-form mixed model · wearable measurement error: calibration, artefact rules, residualisation · power stated honestly · insider-bias controls · stage-gate lens.
Biometric signal & ML research design
HR steers, HRV is measured, self-report is the construct · pre-registered cleaning (motion mask, Kalman, ectopic rules) · per-user baseline → residual "calmer than your usual 10 pm" · E0 calibration → E1 pilot → E2 micro-experiments → E3 bandit personalisation → E4 self-generating library.
Regulatory strategy stays wellness-only ("relieve stress", never "treat"); clinical/DTx route deferred to post-funding. Method: Fairfield & Charman 2017; GRADE; Bland & Osterwalder 2019.
Questions I'd value your guidance on
Updated for where the project stands: survey and focus groups in, pilot and legal check ahead, decision layer live.