The bridge from Essentials: where healthcare AI value actually appears, how to read any proposal along the chain from problem to output to decision to action to outcome to value, four questions that screen a proposal out before a business case, actionability under real capacity, and stating a proposal as a falsifiable value hypothesis.
Name which surface of value a proposal is actually pursuing, and what evidence that surface demands.
Trace a proposal along the impact chain: problem, output, decision, action, outcome, value.
Find the weakest link in that chain before any model, vendor or metric is discussed.
Screen a proposal with four questions, including whether something simpler would do.
Test actionability: the action, the owner, the moment, and the capacity to act at that moment.
State a proposal as a value hypothesis — for whom, what changes, what value, how measured, what would falsify it.
Say what an external performance result does and does not license for your own population and setting.
Lesson progress
0 / 5 steps
Not started
Learn
Building on Essentials
AI in Healthcare Essentials established the vocabulary: how AI, machine learning, deep learning and generative AI relate; prediction versus generation; what a large language model does; deterministic versus probabilistic behaviour; training versus inference; and the first test of whether a problem needs AI at all. None of that is re-taught here. This module asks what you must decide once those words are settled.
A predictive model estimates a defined target; a generative model creates content.So the first design decision is what the target is — and who decided that it counts as the event.
Training estimates parameters; inference applies the fixed model in the workflow.So the question that matters is what was visible at the moment of inference, not what existed in the training records.
Most models return a score rather than a verdict.So somebody has to choose the operating point — and that choice is a clinical and operational policy, not a setting.
Value comes from a decision changing, not from the output existing.So a proposal has to name the action, the owner, the shift and the capacity before it is fundable.
Value is real but often accrues to a different budget than the cost
Patient & staff experience
How the service feels to use
Access, communication, navigation, reduced repetition of the same story
Evidence needed: experience measures alongside access and equity checks
Legitimate on its own terms — not a substitute for clinical evidence
Healthcare AI does not create value in one place. It appears on four distinct surfaces: a clinical decision made earlier or better; time and effort removed from documentation and knowledge work; an operational flow that runs with less waste; and an experience — for patients or staff — that improves. Naming the surface is the first practitioner move, because each surface demands a different kind of evidence.
Most weak proposals are weak because the surface is unnamed. A tool described as 'improving care' is usually pursuing one surface while being justified with the evidence of another — diagnostic accuracy quoted for a drafting tool, or time saved quoted for something that changes clinical conclusions.
The chain from problem to value
01
Problem
A named pressure on care, cost, access or workload.
02
Output
What the system actually produces, for whom.
03
Decision
The decision that would be made differently.
04
Action
Who does what, when — and with what capacity.
05
Outcome
What changes for the patient or the service.
06
Value
Realised, measurable, and attributable.
Read any proposal along this chain and say which link is weakest.
Value reaches a patient or a service only through a chain: a problem worth solving, an output, a decision that changes, an action someone takes, an outcome that shifts, and value that is realised and measurable. The chain is only as strong as its weakest link, and the weakest link is almost never the model.
Reading a proposal along the chain is fast and it is diagnostic. It tells you which link is unspecified, which is the real project risk, and which module of this course you need next.
Four questions before any model
Four questions before any model
What decision will change?
If nothing changes, the output is alert volume.
Who acts, with what capacity?
Name the role and the shift before development.
Is the outcome recorded reliably?
A noisy label limits everything downstream.
Would something simpler do?
Compare with a rule or a workflow fix first.
These four questions screen proposals out before anyone commissions data work. They are deliberately blunt: if a proposal cannot answer them in a meeting, it is not ready for a business case.
Actionability, capacity and the value hypothesis
Actionability is the link that fails most often. An output is actionable when there is a defined action, a named owner, a moment at which it is taken, and the capacity to take it at that moment — including at 03:00 and in August. Capacity is not an implementation detail: where capacity is fixed, capacity effectively decides how much of the output can ever be used.
The value hypothesis is the positive counterpart to the four screening questions: how a proposal that survives states what it is for, including what would falsify it. Later modules refine each part — Module 2 the data, Module 3 the clinical decision and the operating point, Module 6 the evidence, Module 7 the controls, Module 8 the investment case, Module 10 the implementation.
The value hypothesis
Before any model is chosen, write the proposal as five short answers. It is a specification device, not a business case — the fifth answer is what keeps it honest.
1For whom?Which patients, which staff group, which service — named, not 'the organisation'.
2What changes in the workflow or decision?The specific step that is done differently, by whom, and at what moment.
3What value is expected?Which of the four value surfaces, stated as a direction of travel rather than a promised number.
4How will we measure it?One or two indicators you could actually collect, plus the baseline you compare with.
5What would falsify the hypothesis?The result that would make you stop or redesign — no change in the indicator, or benefit bought at an unacceptable cost elsewhere.
Worked example: imaging worklist triage
For whom? Patients having outpatient CT scans, and the reporting radiologists.
What changes? Studies with a suspected time-critical finding are moved to the front of the reporting queue instead of being read in arrival order.
Expected value? Clinical quality — shorter time from scan to actionable report for the subset that needs it most.
Measured how? Median and 90th-percentile time from scan to report for confirmed time-critical findings, against the pre-deployment baseline; total reporting throughput as a balancing measure.
Falsified by? No shift in reporting time for that subset, or a shift bought by delaying everyone else beyond an agreed limit.
Two patterns worth carrying forward
Transportability
The model that travelled badly
A recurring, well-documented pattern: a risk model performs strongly where it was developed, then degrades elsewhere because the population, coding practice and workflow differ.
External results are evidence about the approach, not proof about your setting.
Capability turned around
What a system can be made to do
The same generative capability enables convincing deepfakes, and deployed systems can be manipulated through jailbreaks or prompt injection hidden in the content they read.
In healthcare this is a security and governance question: what a system can be made to say or do, not only how well it performs.
Illustrative examples for discussion — not assessed content.
Optional previews
Neither preview is required, and neither is assessed in this module. Both concepts are taught properly where they belong — open them only if the question is live for you now.
Optional preview: how the target and the prediction moment are setOptionalOne level below the value hypothesis. Taught in full in Module 2 — open it only if you want the mechanics now.
A clinical prediction model learns from examples carrying a known target, so two decisions are made before any modelling: what exactly counts as the event, and at which instant the model must produce an output.
Most healthcare labels are produced by a care process rather than observed directly — ICU transfer, a coded diagnosis, an accepted referral — so a model fitted to one partly learns local decision behaviour. The same process creates the risk that variables generated by the response to an event appear to predict it. Module 2 turns this into a systematic method.
Optional preview: the operating pointOptionalWhere a score becomes an alert. Taught in full in Module 3 — open it only if the trade-off is live in your organisation now.
A score is not yet an alert. The instant it triggers an action, sorts patients or fills a fixed number of programme places, someone has chosen an operating point — and where capacity is fixed, capacity chooses it.
Move the alert threshold
An illustrative cohort — not validated clinical data.
Lower — flag more40%Higher — flag fewer
flagged, event occurred flagged, no event event missed not flagged, no event
Patients flagged
27 of 100
Events caught
7 of 12
Events missed
5
False alerts
20
Alerts that are real
26%
The model has not changed — only the operating point has. Threshold choice sets workload, missed events and alert credibility, so it is a clinical and operational policy decision, not a data science setting.
Sources & evidence · 4 sources
This module cites public or consensus guidance, scholarly literature.
Content reviewed: September 2026. Publication dates of the individual sources are shown in each citation.
Collins GS, Moons KGM, Dhiman P, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:e078378. doi:10.1136/bmj-2023-078378
Sets out what a prediction-model report should contain, including how predictors and outcomes are defined and timed. It is a reporting standard: following it makes evidence legible, but does not itself establish that a model is safe or effective.
Vasey B, Nagendran M, Campbell B, et al. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nature Medicine. 2022;28:924–933. doi:10.1038/s41591-022-01772-9
Supports the distinction between model performance and clinical evaluation of a decision-support system in live use. It covers early-stage clinical evaluation and does not replace comparative effectiveness evidence.
World Health Organization. Ethics and governance of artificial intelligence for health. WHO guidance. 2021. ISBN 978-92-4-002920-0
Supports the governance and 'does this need AI at all' framing at policy level. It is guidance rather than binding regulation, and does not determine any specific jurisdiction's legal requirements.
National Institute of Standards and Technology. Generative artificial intelligence. NIST Computer Security Resource Center Glossary; definition sourced by NIST to NIST SP 800-218A, Secure Software Development Practices for Generative AI and Dual-Use Foundation Models (July 2024). doi:10.6028/NIST.SP.800-218A. Entry status checked 25 August 2026.
Used only for a stable, citable definition of generative AI. A glossary entry standardises terminology; it supports no claim about healthcare performance, safety or risk.