How the evidence is made to count.
Two lenses on one design. Quantitative methods: what can be identified, with what error, at what power — a within-subject crossover with an active control, ANCOVA-form mixed models, wearable measurement error handled by calibration and residualisation. Management science: insider action research run as a stage-gate — every source is assigned a role in a Bayesian layer, every test is a purchase of information, and the phase ends at pre-registered gates.
Source: research/METHODOLOGY_DESIGN.md · companions: DECISION_MODEL.md · surveys/questionnaire-v2.md · FOCUS_GROUP_KIT.md · ML_RESEARCH_DESIGN.md
Convergent mixed methods with a sequential quantitative core
Frame: insider action research (Coghlan & Brannick). Process: design thinking. Method architecture: two strands that meet in one Bayesian integration — the qualitative strand generates and interprets, the quantitative strand quantifies and tests.
Produces the arousal-state hypothesis, the acoustic-feature list, the symptom clusters; enters the model with heavy shrinkage.
integration →
gates
Estimates shares and price points; the pilot estimates a causal within-person effect; behavioural and legal evidence close the gates.
N1 · N2 · N3
Frequent moments (B7) · under-served outcomes (ODI) · state-dependent need and content (C3, FG T1–T2).
A1 – A5
Content by state · one-session effect on STAI-S and heart-rate measures · connect & value · pay · rights.
Gates
Posterior thresholds fixed in advance; each source's role bounded by its quality — nothing double-counted.
Source × construct × role — who is allowed to say what
prior sets the starting belief · generate produces hypotheses · quantify estimates a share/mean · test estimates a causal effect · texture informs interpretation only.
| Construct | Literature | Interviews | Netnography | Focus groups | Survey | Pilot | Landing / legal |
|---|---|---|---|---|---|---|---|
| Need prevalence & intensity | prior | texture | texture | texture | quantify | — | — |
| Timing / place / triggers | — | generate | — | map | quantify | — | — |
| Under-served outcomes | — | generate | complaints | rank | quantify (ODI) | — | — |
| Content by state (A1a/A1b) | weak prior | generate | both coexist | stimulus test | quantify | A/B test | — |
| Efficacy (A2a/A2b) | prior | anecdote | anecdote | — | — | test | — |
| Engagement (A3) | precedent | absence | action gap | probe | quantify | think-aloud | conversion |
| Willingness to pay (A4) | market data | — | billing anger | reaction | PSM · PI · F9 | — | conversion · partners |
| Rights (A5) | label deals | — | — | — | — | — | legal check |
Validated where it exists, adapted and disclosed where it doesn't
GAD-2 · PSS-4 · STAI-S(6) · PSM
Kroenke 2007 · Cohen & Williamson 1988 · Marteau & Bekker 1992 · van Westendorp 1976. Wording untouched; α and convergent r reported.
TAM · Kano · ODI · DOI · expectancy
Davis 1989 · Kano 1984 · Ulwick 2002 · Rogers 2003 · Devilly & Borkovec 2000. Two items per TAM construct; CFA/PLS-SEM if n ≥ 150.
HR · RMSSD / SDNN · residual Δ
Wrist HR accurate (≈ 6%); wrist HRV MAPE ≈ 29% vs chest strap → calibration sub-study, artefact rules, residualisation. Pipeline →
Biometrics are honest about arousal, not about anxiety
HR/HRV index autonomic arousal — exertion, caffeine, excitement and fear all move them. Self-report (STAI-S) is therefore the construct measure and the primary outcome; heart-rate measures are the mechanism check and the product's control signal. The design reports the concordance between the two rather than assuming it (Panteleeva 2018 found self-report effects without physiological ones).
A within-subject crossover with an active control, analysed as ANCOVA in a mixed model
3 conditions · Latin-square order · 3 sessions per person
T1 adaptive neutral sound · T2 adaptive calm real song · C active control = a generic "relaxing" playlist the participant already uses (not silence — silence changes expectancy and confounds "any audio" with "our audio").
15 min, seated, same time-of-day, ≥ 24 h washout; pre/post STAI-S(6), expectancy items, 5-min rest baseline; a research assistant runs the script (the founder does not).
Within-person effect on change, baseline-adjusted
β1, β2 = effect vs active control at equal baseline; β1 − β2 answers A1a on the outcome itself. Report 95% CI, d_z, ICC — and a Bayesian re-analysis with the meta-analytic prior (variance ×2): P(β1 < 0 | data) drives the gate. ANCOVA form over raw change scores (Vickers & Altman 2001).
| Threat | Direction | Control |
|---|---|---|
| Expectancy / placebo | inflates T | active control; expectancy measured and entered as covariate; analyst blinded to labels; all arms described as "calming audio" |
| Regression to the mean | inflates any drop | baseline covariate; 5-min rest so baseline is not the arrival spike |
| Order / carry-over | biases later sessions | Latin square; Order and Session# covariates; washout |
| Circadian HRV | noise / bias if unbalanced | fixed slot per participant; covariate |
| Motion, caffeine, speech | artefacts | seated protocol; accelerometer masking; 2-h rule logged |
| Founder demand effects | inflates T | RA-run sessions; founder absent; blinded analysis |
| Multiple outcomes | false positives | one primary (STAI-S); secondaries with Holm correction |
| Small n | low power for physiology | 3 sessions/person; Bayesian analysis; feasibility framing |
Wearable measurement error — what it does and what to do
As an outcome
Classical (non-differential) error adds variance → loss of power, not bias. Remedy: more sessions per person; HR (accurate) as co-primary physiological outcome, RMSSD secondary; average the last 5 min.
As a regressor
Errors-in-variables → attenuation. Remedy: calibration sub-study (n ≈ 8, Polar H10 concurrent) → reliability ratio λ, regression calibration / instrument the wrist value with the chest value.
Differential error
A beat-driven track could induce motion → biased comparison. Remedy: pre-specified artefact rules, motion-flag rate compared across arms, sensitivity analysis on clean windows.
Residualisation
Two-stage: (i) person-specific baseline model on ≥ 14 days of non-session data (time-of-day, activity, recent HR); (ii) outcome = observed − predicted over the session. Interpretable as "calmer than your usual 10 pm" — the same number the product shows. Kalman smoothing before features.
Honest about what n = 20 can and cannot show
± 9.8 pp (± 8 pp at 150); cells ≥ 40 owners, ≥ 30 GAD-2+; McNemar 80% power for a 20-pp shift
saturation, stimulus reaction; stop when a group adds no new code
Cochrane STAI-S ≈ d .55 → n ≈ 28 for self-report; physiology d ≈ .4 → n ≈ 50
feasibility signal for A2a, under-powered for A2b — stated; Bayesian posterior with meta-analytic prior drives the gate, not p
Controls that are actually implementable by a founder-researcher
Confirmation
Thresholds and LRs pre-registered; the model built before the data; kill criteria written down; a negative result is a valid outcome.
Demand effects
Research assistants run sessions and groups with scripts; the founder is absent from rooms; the survey is anonymous.
Analyst freedom
The analysis code and thresholds are written and tested before the data are analysed (pre-registered plan); one primary outcome; deviations logged.
Thematic analysis
Braun & Clarke; hybrid codebook from netnography codes; second coder on 20% (κ ≥ .70); saturation criterion; member checks on reconstructed interviews.
How qualitative evidence counts
Enters the Bayesian layer only after shrinkage (k = .35–.5) and correlation discounting; can move a belief, cannot cross a gate alone — the formal answer to "how much do 3 interviews count".
Reflexivity
Researcher diary; a reflexivity section in the methods chapter; limitations stated up front (non-probability samples, hypothetical WTP, small pilot, consumer physiology, subjective LRs).
Evidence as a stage-gate: each test is a purchase of information
Order tests by information per cost
Legal check first (cheapest, most decisive: LR 8 / 0.15) → survey → pilot. The DT phase is an option on the MVP, not a commitment (Bowman & Hurry; McGrath).
C1 · C2 · C3 are hedges, not rivals
C3 probes real-music calm demand at near-zero cost; C2 hedges consumer-payment friction; C1 is the asset-leverage bet. Gates select which to fund.
A ring-fenced team can kill honestly
Small separate team, $10–20k, ≤ 3 months — the organisational condition for a real kill decision (O'Reilly & Tushman).
| Wk | Activity | Output → node |
|---|---|---|
| 1 | Legal check on ≥ 50 tracks starts; survey back-translation + n = 3 pilot; landing A/B live | A5 (LR 8/0.15) · A3/A4 conversion |
| 2–4 | Survey fielding (n ≥ 100); FG-1, FG-2 recruited from opt-ins | A1a/A1b/A3/A4 survey LRs · FG themes |
| 3–5 | Calibration sub-study (n = 8, Polar H10); WoZ pilot 20–30 × 3 sessions | A2a / A2b |
| 6 | Pre-registered pipeline → Bayesian update → gate verdict → decision memo | PROCEED / PIVOT / KILL |