The benchmark is not the bedside
An algorithm can discriminate cases from controls in a curated test set and still fail in practice. Prevalence changes predictive value. Acquisition devices and protocols change input distributions. Clinicians see histories, laboratory values, uncertainty, and competing diagnoses that may be withheld from a benchmark comparator. Workflow determines whether an output arrives in time, reaches the right person, and leads to a safe action. The unit of evaluation is therefore the AI-enabled care pathway, not the model in isolation.
The 2020 BMJ review is an instructive baseline. It identified 81 non-randomized studies comparing deep learning with clinicians; nine were prospective and only six used a real-world clinical environment. Fifty-eight were judged at high risk of bias. The review does not describe every contemporary clinical AI system, and the field has advanced, but it demonstrates why a performance headline cannot carry an implementation claim. Study design, comparator conditions, external data, and clinical setting determine what the number means.
Define the claim before choosing the model
A defensible programme begins with an intended-use statement: target population, setting, user, input, output, decision, time horizon, and excluded uses. The endpoint should reflect the decision. A triage tool may require sensitivity at an operationally tolerable alert rate; a risk model needs calibration and decision utility, not only area under the curve; a documentation assistant needs factuality, omission, and workload measures. The current pathway and its error distribution are the comparator. Without that baseline, an impressive model may optimize a proxy that does not improve care.
Evidence maturity should be labeled. Analytical validation establishes performance on fit-for-purpose data. External validation tests transport to a different site, time, device, or population. Silent prospective evaluation observes the live stream without influencing care. Early clinical evaluation studies human interaction and workflow at small scale. Comparative prospective evaluation tests consequences against a control. Each stage can falsify a different assumption; none should be silently substituted for the next.
External validity is a design variable
Randomly splitting one dataset does not demonstrate transportability. Development and evaluation records may share sites, devices, documentation conventions, duplicated patients, or future information leaked through preprocessing. A stronger plan separates data by the boundary the claim must cross: later time, independent institution, different acquisition system, or materially different population. It reports confidence intervals, calibration, missingness, and subgroup performance, and it examines whether thresholds chosen during development remain clinically acceptable.
Fairness cannot be reduced to equal aggregate error rates. The relevant harm depends on task and pathway: a false negative can delay escalation, a false positive can trigger invasive work, and an abstention can transfer workload. Evaluation should report sample and event counts for intersecting subgroups, but it must also state when those counts are too small for stable inference. WHO’s governance principles add autonomy, well-being, transparency, accountability, inclusion, and sustainability; these are operating requirements that shape consent, oversight, redress, and resource use.
Human factors can reverse model performance
Clinical AI is a joint cognitive system. Users may over-trust an authoritative interface, ignore a noisy alert, anchor on a recommendation, or learn workarounds that never appear in model metrics. The system may improve a novice while distracting an expert, or increase sensitivity while creating review queues that delay everyone. Early evaluation should observe who sees the output, what information accompanies it, how uncertainty is represented, when users accept or override it, and what happens after disagreement.
DECIDE-AI addresses this early live phase with 17 AI-specific reporting items and ten generic items. Its scope includes system version, clinical workflow, user experience, performance errors, and human factors. The guideline improves transparent reporting; it does not certify that a study was well designed or a product is safe. A laboratory should therefore pair reporting compliance with pre-specified usability criteria, scenario testing, escalation drills, and direct measurement of downstream clinical and operational effects.
Versioning turns approval into a lifecycle
Software changes after initial evaluation: code, weights, prompts, knowledge sources, thresholds, devices, and upstream record systems evolve. A reproducible evidence package binds results to model identifier, training-data lineage, preprocessing, dependencies, configuration, intended-use policy, and evaluation dataset. The deployed system should log input-quality checks, output, uncertainty or abstention, user action, override, latency, and safety events under appropriate privacy controls. Monitoring must distinguish a change in case mix from a change in performance.
A change protocol specifies which modifications require regression testing, local revalidation, regulatory review, or suspension. Monitors need denominators and alert thresholds; raw alert counts are not enough. Drift detection is not clinical validation, because stable feature distributions can coexist with changed outcomes and a shifted distribution can be harmless. Periodic outcome review, incident analysis, and recalibration governance remain necessary. The FDA’s public inventory of AI-enabled medical devices is useful market context, but inclusion on a list does not establish comparative benefit for a particular local workflow.
The deployment dossier
Before a clinical pilot, the dossier should contain the intended use, pathway map, hazard analysis, evidence table, data sheet, model card, threshold rationale, subgroup plan, human-factors findings, security assessment, rollback procedure, and monitoring protocol. It should identify the accountable clinical owner and the party authorized to pause operation. Results should follow the reporting framework appropriate to stage: TRIPOD+AI for prediction-model studies, DECIDE-AI for early live evaluation, and CONSORT-AI where a randomized trial of an AI intervention is conducted.
A go decision requires more than statistical significance. The effect must be clinically and operationally meaningful, uncertainty compatible with the risk, performance acceptable across pre-specified groups, and failure recoverable within the care system. A no-go or redesign result is productive evidence. The laboratory standard is a traceable argument from claim to measurement to decision, with each inference bounded by its study design. That discipline is slower than a leaderboard announcement and faster than repairing an unmeasured failure in care.
Scope and limitations
The BMJ review searched through June 2019 and should be read as a rigorous historical baseline, not a census of current clinical AI. Reporting guidelines improve completeness but do not themselves establish methodological quality, regulatory authorization, or clinical benefit. Requirements vary by jurisdiction and intended use. The proposed ladder is a synthesis for research governance and must be adapted with clinicians, patients, statisticians, security specialists, and the relevant regulator; it is not medical advice.
References
Source review: 20 August 2026. Quantitative values retain their original definitions, periods, and boundaries.
- 01Artificial intelligence versus clinicians: systematic review of design, reporting standards, and claims
The BMJ · 2020
www.bmj.com ↗ - 02DECIDE-AI: reporting guideline for early-stage clinical evaluation of AI decision support
Nature Medicine · 2022
www.nature.com ↗ - 03Ethics and governance of artificial intelligence for health
World Health Organization · 2021
www.who.int ↗ - 04Artificial Intelligence-Enabled Medical Devices
US Food and Drug Administration · 2026
www.fda.gov ↗ - 05TRIPOD+AI statement
The BMJ · 2024
www.bmj.com ↗ - 06CONSORT-AI extension
Nature Medicine · 2020
www.nature.com ↗

