Define the measurement before naming the biomarker
A wearable produces samples: acceleration, optical intensity, electrical potential, temperature, or another physical signal. Firmware transforms those samples; an algorithm may then estimate steps, rhythm, sleep, gait, or behavior. Calling the final number a biomarker makes a biological claim that the device output indicates a normal process, pathogenic process, or response. That claim requires a defined measurand, algorithm, population, setting, time window, and decision. A dashboard label is not sufficient.
Terminology remains unstable. A systematic mapping review identified 415 biomedical articles using “digital biomarker” through March 2023. Sixty-nine percent gave no definition. The 128 articles that did contained 127 distinct definitions, and only 23 definitions included all three examined components: data type, collection method, and purpose. This is not semantic housekeeping. Different definitions imply different validation evidence and make apparently similar study results difficult to compare.
Use a three-part validation chain
The V3 framework separates verification, analytical validation, and clinical validation. Verification asks whether the hardware and low-level software measure what they are specified to measure under controlled conditions. Analytical validation asks whether the algorithm transforms sensor data into an accurate physiological or behavioral measure across expected conditions. Clinical validation asks whether that measure identifies, monitors, or predicts the stated clinical construct in the defined population and context. Passing one layer does not imply the next.
A heart-rate algorithm can track a reference electrocardiogram under rest while degrading with motion, skin contact, perfusion, device placement, or firmware change. Even accurate heart rate does not establish a diagnostic claim for arrhythmia. Validation should therefore include reference methods, range and repeatability, known interference, relevant subgroups, device and operating-system versions, missing-data mechanisms, and pre-specified acceptance limits. “Clinically validated” without this claim-specific evidence is an incomplete statement.
Large pragmatic studies reveal both reach and selection
The Apple Heart Study enrolled 419,297 US participants using compatible consumer devices. Over a median 117 days, 2,161 participants, or 0.52%, received an irregular-pulse notification. Among 450 notified participants who returned an analyzable ECG patch, atrial fibrillation was present in 34%. For notifications occurring simultaneously with ECG monitoring, positive predictive value was 0.84. Each denominator answers a different question; combining them into a generic “84% accurate” claim would be misleading.
The study also demonstrates selection and workflow effects. Participants owned compatible hardware, reported no prior atrial fibrillation, and were younger on average than many populations at greatest risk. Only a subset of notified participants returned an analyzable patch, applied on average thirteen days later. The study established feasibility and specific performance estimates within its design; it did not show that population screening reduced stroke or mortality. Clinical value requires downstream comparative evidence, including consequences of alerts and missed episodes.
Continuous data is continuously conditional
Remote streams are shaped by wear behavior, battery, synchronization, connectivity, handedness, work, sleep, travel, illness, and changes in device ownership. Missingness may be informative: a person can stop wearing a device because symptoms worsen, the device irritates skin, or charging becomes burdensome. Imputing those intervals as random can bias a model. Research records should preserve raw or minimally processed data where feasible, sampling rate, quality flags, wear-time estimates, transformations, and reasons for exclusion.
A 2020 wearable study reported physiological changes associated with COVID-19 before symptoms in a small cohort of infected participants. It was valuable proof-generating research, not a universal diagnostic test. Such findings require prospective replication against an appropriate reference, assessment of other infections and confounders, and pre-specified alert performance in the target population. Retrospective alignment to a known symptom date is especially vulnerable to optimistic threshold selection.
Remote collection changes the research operation
FDA’s 2023 guidance treats digital health technologies as hardware, software, or combinations used to acquire data remotely in clinical investigations. Selection should follow fitness for purpose and participant risk. Sponsors need instructions, training, technical support, data-transfer plans, cybersecurity controls, record retention, and procedures for updates, loss, and malfunction. Participants should understand what is collected, how often, who receives it, whether it is monitored in real time, and what the research team will or will not do when a concerning signal appears.
The last point is essential. Research collection can create a false expectation of clinical surveillance. Protocols must define alert review, response windows, emergency disclaimers, incidental findings, withdrawal, and device return. Equity analysis should measure who can participate given hardware, language, disability, connectivity, and charging requirements. A technically excellent endpoint can still make a trial less representative if participation depends on an expensive phone or high digital literacy.
The minimum evidence package
A digital-measurement dossier should specify the concept of interest, context of use, sensor and algorithm versions, reference method, verification and analytical-validation results, clinical-validation population, missing-data plan, usability evidence, security model, and change controls. Metrics should include data yield, valid wear time, failure by device and subgroup, agreement with reference including bias and limits, test-retest behavior, clinical association, and the decision consequence of thresholds. Results require uncertainty intervals rather than a single accuracy value.
Evidence maturity should be labeled as feasibility, verified sensor, analytically validated measure, clinically validated measure, qualified endpoint, or outcome-tested intervention, using terms appropriate to the jurisdiction and programme. The label should travel with every chart and model input. A sensor stream is valuable precisely because it can measure life outside a visit. The laboratory obligation is to preserve the conditions under which that stream becomes interpretable, and to stop the claim at the boundary the evidence supports.
Scope and limitations
The mapping review describes terminology through March 2023. The Apple Heart Study was sponsored by Apple, used a self-selected device-owning cohort, and had incomplete follow-up; its predictive values depend on the studied population and confirmation pathway. Findings from infection-detection research were exploratory and are not diagnostic guidance. Regulatory status and endpoint qualification are context-specific. This article does not recommend a device, screening programme, diagnosis, or treatment and should not replace clinical or regulatory review.
References
Source review: 20 August 2026. Quantitative values retain their original definitions, periods, and boundaries.
- 01Digital Health Technologies for Remote Data Acquisition in Clinical Investigations
US Food and Drug Administration · 2023
www.fda.gov ↗ - 02Definitions of digital biomarkers: a systematic mapping of the biomedical literature
BMJ Health & Care Informatics · 2024
pmc.ncbi.nlm.nih.gov ↗ - 03Large-Scale Assessment of a Smartwatch to Identify Atrial Fibrillation
New England Journal of Medicine · 2019
www.nejm.org ↗ - 04Pre-symptomatic detection of COVID-19 from smartwatch data
Nature Biomedical Engineering · 2020
www.nature.com ↗ - 05Verification, analytical validation, and clinical validation (V3)
npj Digital Medicine · 2020
www.nature.com ↗

