Towards Operationalisable Clinical Risk Prediction Models
AKI is the current clinical demonstrator. This PhD asks how a risk model can move from a valid
prospective target, to a faithful representation of the evolving patient and clinical context, to an
operationally meaningful signal that supports structured review without pretending that prognosis is the
same as preventability or treatment response.
Current repository authority: HOLDAKI = current demonstratorContext = haemodynamic · infection · medication · procedureRisk score → structured context review, not treatment
Executive summary
One briefing, with a fast supervisor-level overview and deeper epidemiology / phenotype detail underneath.
Predict clinically meaningful deterioration while preserving prospective information boundaries.
accepted framing
Intended decision
Prioritise structured review
The treating ICU / acute-care team reviews prospectively observable context and routes onward when warranted.
DEC-0084
Primary outcome
First prospectively ascertainable Stage ≥2
Persistent / severe trajectory outcomes remain secondary; the operational endpoint is observation-process dependent.
DEC-0087
Current authority
No globally accepted target, features or model
Candidate.8 is canonical and non-executable; patient-level target/model execution remains gated.
HOLD
Working research question
In adults at repeated ICU-origin prediction times, can prospectively observable clinical context improve the clinical meaning and operational usefulness of future severe-AKI risk prediction beyond a strong time-updated physiology baseline, and which contextual representations provide stable incremental value without compromising prospective validity?
Research framework
The thesis is organised around three linked methodological problems. Operationalisability is evaluated throughout rather than added as a final deployment chapter.
Translate a clinical construct into a prospectively valid EHR prediction target with explicit risk-set, timing, component, missingness and terminal-event semantics.
Main AKI issue
Clinical AKI ≠ EHR-ascertainable AKI ≠ operational prediction endpoint.
→
2 · Represent
Represent the evolving patient and context
Move beyond a flat list of predictors to dynamic renal, physiological, treatment, exposure and observation-process states.
Main AKI issue
Context changes the meaning of the same observed value or medicine.
→
3 · Justify complexity
Add external knowledge only for a demonstrated residual problem
Test conventional representations first. Knowledge-informed or neuro-symbolic methods enter only if they address a specific limitation that remains.
Research design principle
Method follows the clinical and methodological requirement, not novelty.
Current context priority (DEC-0087): observed haemodynamic / treatment-state context first, infection / acute deterioration second, medication / pharmacology third, and procedure / surgery / contrast fourth. Observation process is a cross-cutting validity layer rather than a later competing context block. Each added context still requires its own evidence, prospective-observability and specification gate.
Epidemiological & prediction design
The detailed review portal is folded into this briefing here: who enters the risk set, when prediction is issued, what event is forecast, how follow-up works, and where uncertainty remains.
Current prospectively ascertainable Stage 0 or Stage 1 can remain eligible for first future Stage ≥2 prediction.
DEC-0004
When
Repeated ICU-origin occasions
Rolling entry opens from ICU +6 h. q6 is the accepted primary model risk-refresh cadence; q12 is the prespecified cadence sensitivity.
DEC-0011 / 0012
What
First Stage ≥2
Forecast the first prospectively ascertainable KDIGO Stage ≥2 event; persistent / severe trajectory remains secondary.
DEC-0003 / 0006 / 0014
Horizon
24 h primary
The primary horizon is (t, t+24 h]; 48 h is secondary. A prespecified sensitivity excludes events in the first 6 h after prediction.
DEC-0013 / 0025
Follow-up
Hospital-wide issued horizon
Leaving the index ICU stops new prediction issuance, but an already-issued horizon continues under hospital follow-up. Death and discharge remain distinct terminal states.
DEC-0015 / 0016 / 0023
Decision evidence trail
Why these epidemiological design decisions?
This makes the reasoning auditable: external literature defines the methodological problem, project epidemiology quantifies the trade-off in MIMIC, and a separate human decision records the selected design and what remains uncertain.
partial design · project HOLD
1 · External evidenceClinical definitions, prediction methodology and observation-process literature.
3 · Trade-offCompare event support, warning time, observability and operational consequences.
4 · Human decisionA bounded, time-stamped design choice is recorded in the decision register.
5 · Residual uncertaintyNamed limitations and reopening triggers remain visible rather than being silently resolved.
external / literature evidenceproject empirical evidenceaccepted partial decisionremaining uncertainty
Design question
Evidence examined
What the evidence showed
Current decision
Remaining uncertainty
Who enters the risk set?
EPID-001, EPID-002 + phenotype literature
Stage 1 remained a meaningful at-risk state. At q6/24 h, Stage 0/1 eligibility captured 2,669 additional unique future Stage ≥2 events versus Stage 0 alone.
Current prospectively ascertainable Stage 0 or Stage 1 is eligible for first future Stage ≥2 prediction. DEC-0004
Eligibility still depends on the final prospectively available-component phenotype.
When should prediction begin?
EPID-006 + temporal/workflow review
Governed prospective renal states can already be available at +6 h. A fixed +24 h landmark would impose a later entry point than the reviewed data require.
Rolling eligibility opens from ICU +6 h; UNKNOWN is deferred rather than forced into Stage 0/1. DEC-0011
Not every patient is classifiable at +6 h; the final component set can alter early classifiability.
How often should risk update?
EPID-002 + workflow literature
q6 h is the highest-resolution reviewed rolling substrate. Model computation cadence is conceptually separate from clinical alert/review cadence.
q6 h primary; q12 h prespecified sensitivity / efficiency analysis. DEC-0012
Whether q12 h gives similar practical performance with lower repeated-prediction burden remains empirical.
24 h or 48 h horizon?
Rolling target-design audit + warning-time review
24 h captured 5,796 unique Stage ≥2 events versus 6,223 at 48 h (93.1% of 48 h coverage). Median warning was 11.1 h versus 19.7 h; 24 h also had fewer after-transfer positives and terminally shortened windows.
24 h primary; 48 h secondary. Retain a 0–6 h minimum-warning-time sensitivity. DEC-0013
The clinical value of extra warning versus tighter near-term prediction remains an explicit trade-off.
What counts as Stage ≥2?
REV-PHENO-01, KDIGO / operationalisation literature + MIMIC phenotype evidence
SCr, UO and KRT have different observation processes. An observed qualifying component can establish AKI; an unavailable component cannot be treated as criterion-negative.
First prospectively EHR-ascertainable available-component KDIGO Stage ≥2; hospital-wide SCr Stage ≥2 retained as a major sensitivity / benchmark. DEC-0014
The complete executable target remains gated; component implementation and observation adequacy still require governed closure.
What happens at ICU transfer?
EPID-002 + ascertainment / workflow literature
6.8% of q6/24 h Stage ≥2 positive windows occurred after ICU transfer. Transfer changes the observation regime but does not biologically end AKI risk.
An already-issued horizon continues hospital-wide; no new ward predictions are created for the current MIMIC design. DEC-0015, DEC-0023
A future ward-compatible prediction model would require its own feature and prospective-availability contract.
Death and hospital discharge alive terminate the opportunity for the explicitly in-hospital endpoint and should not be collapsed into ordinary “no AKI” labels.
Preserve death-before-Stage ≥2 and hospital-discharge-before-Stage ≥2 as distinct terminal states. DEC-0016
The later statistical model still has to specify how these terminal states enter learning and evaluation.
An event already ascertainable at prediction time t is not a future outcome; a deterministic right-endpoint convention is needed for reproducibility.
Use (t, t+H], with H = 24 h primary and 48 h secondary. DEC-0025
Main residual issue is source-specific timestamp precision / ordering, not the interval convention itself.
Interpretation rule: literature evidence, local empirical evidence and a project design decision are deliberately kept separate. The evidence supports and constrains a decision; it does not make the decision automatically. These are accepted partial design decisions, while CURRENT_AUTHORITY.yaml still records the project as HOLD with no globally accepted target, feature contract or model.
Figure 1. Prospective rolling-prediction design.
Rolling prediction eligibility opens from ICU +6 h, with q6 h as the primary model risk-refresh cadence and q12 h as a prespecified sensitivity. The primary horizon is (t, t+24 h], with 48 h secondary. New prediction issuance ends when the patient leaves the index ICU, while an already-issued prediction remains under in-hospital follow-up through its original horizon. These are accepted partial design decisions; the complete executable target remains gated.
Figure 2. Development ICU-stay substrate and q6 prediction-occasion structure.
The current UO/q6 development substrate has 45,619 patients, 59,357 hospital admissions and 65,813 valid ICU stays. Of these, 64,968 ICU stays contribute at least one scheduled q6 occasion, generating 931,427 stay–time prediction occasions. These are confirmed development-substrate and scheduled-grid counts, not the final target-eligible AKI modelling cohort.
Why this is not just “predict AKI in the next 24 hours”
The phenotype must distinguish clinical KDIGO criteria from what was actually observable in the EHR at the prediction time. Baseline SCr availability, UO evaluability, KRT attribution, event-time versus availability-time, revisions, terminal events and zero-opportunity outcome ascertainment can all change what the label means.
Open epidemiological question — outcome unascertainability
Some prediction windows may contain no valid renal assessment, so Stage ≥2 status cannot be determined. These are retained as outcome unascertainable rather than classified as negative.
For discussion: if these windows are excluded from the primary binary analysis, does this introduce selection through the observation process because renal monitoring itself may be informative? Should this be addressed through reporting and sensitivity analyses, weighting, or explicit modelling of the observation process?
What remains non-authoritative / gated
The complete target is not globally accepted in CURRENT_AUTHORITY.yaml; candidate.8 remains canonical, aligned and non-executable. Global feature and model fields are also null. The page therefore distinguishes accepted partial decisions from candidate/specification state.
Current progress
A brief update on where the work is now and what needs to finish before the next stage.
project HOLD
Where I am
Urine-output data-quality check in progress
Most of the epidemiological study design is now in place. I am currently running a full check of the urine-output pipeline on the development dataset before incorporating it into the prediction framework.
Current bottleneck
The full check needs to finish and pass
The run is computationally slow and is still ongoing. I am not interpreting partial results. Once it finishes cleanly, I can assess whether the urine-output information is sufficiently available and reliable for the planned predictor set.
In short: the study design has moved forward; the immediate bottleneck is completing data processing and validation before integrating urine output into the next stage of the prediction pipeline.
A value is usable only when its clinical event and its information availability are valid for the prediction occasion. This affects both the target and predictor-side context.
A test, observation, drug administration, procedure or treatment state occurs.
2 · Information becomes available
The EHR can expose the information later than the biological or collection event.
3 · Prediction occasion t
Only information validly available by t may enter the model or current-state phenotype.
4 · Future window (t, t+H]
The outcome is sought prospectively after the prediction time.
5 · Ascertainability
Observed non-event is separated from zero valid opportunity to determine the endpoint.
Renal component
SCr, UO and KRT are non-equivalent signals
Each has different observation and timing properties. One observed qualifying component may establish AKI while another remains unavailable.
Context component
Observed proxy ≠ latent state
MAP, lactate, fluids and vasopressors may support haemodynamic interpretation but do not by themselves establish shock, hypovolaemia or a renal mechanism.
Transportability
Measurement policy can change model meaning
An EHR phenotype and its apparent missingness can shift when another hospital measures or documents differently.
How clinical context enters the prediction model
Context domains are added one at a time, with each block requiring its own clinical rationale, prospective-observability check and specification before comparison.
Start with a strong renal and physiological representation using only prospectively valid observed patient information.
↓
Question: what can the evolving clinical state already explain?
context increment
B · Add one governed context domain
Add one clinically justified context block. Current priority is observed haemodynamic/treatment state, then infection/acute deterioration, then medication/pharmacology, then procedure/surgery/contrast.
↓
Each domain is tested separately before combining domains, so any incremental value remains interpretable.
contextual representation
C · Context × patient state × time
Test whether the same context becomes more informative when represented relative to renal trajectory, physiology, treatment state and time rather than as a flat feature list.
↓
Only after fair same-information comparisons should more complex or external-knowledge mechanisms be considered.
Current first context priority: observed haemodynamic / treatment-state context.
This is the first broader-context study to develop next, based on current feasibility, local support and clinical relevance.
priority 1 Observed haemodynamic / treatment state BP/MAP trajectories, governed support state and other prospectively observed haemodynamic information; do not infer hypovolaemia or shock from one proxy.
priority 2 Infection / acute deterioration Sepsis-related and evolving acuity context, with phenotype semantics kept separate from individual treatment or laboratory signals.
priority 4 Procedure / surgery / contrast Prospectively timed procedural and exposure context where clinically justified and locally supported.
cross-cutting validity Observation process Measurement recency, availability, coverage and UNKNOWN state accompany every model comparison rather than competing as a later context block.
later after validation Latent volume / congestion / cause-specific states Potentially high impact, but only after transparent phenotype validation; low MAP ≠ hypovolaemia and heart failure ≠ congestion.
Anti-conflation rule: a clinically valid relationship does not automatically become a predictor, a predictive interaction, a model constraint, or a treatment recommendation.
Evaluation: prediction + operational usefulness
The accepted evaluation architecture avoids reducing success to AUROC alone and keeps retrospective predictive claims separate from clinical-outcome benefit.
Primary horizon; 48 h retained as secondary/sensitivity.
ΔAUPRC
Primary discrimination estimand for the central contextual representation contrast; paired ΔAUROC is key secondary evidence.
Calibration
Reliability curve, calibration-in-the-large/intercept where compatible, slope and Brier score.
Matched burden
Compare PPV, sensitivity and workload at prespecified review-burden operating points rather than inventing an “optimal” cutoff.
First alert
Patient/stay-level first-alert PPV, sensitivity, lead time, unique patients alerted and repeated-alert burden.
2,000×
Paired subject-cluster bootstrap replicates for final uncertainty, preserving within-subject repeated observations.
Claim boundary: retrospective work may establish predictive increment and, with supporting operational evidence, an operationally promising increment. It cannot establish treatment benefit or clinical utility without prospective/interventional evaluation.
Research sequence
A gate-based plan is more defensible than promising a fixed number of papers or methods in advance.
Finish the remaining executable integration and validation gates for the prospectively ascertainable Stage ≥2 phenotype.
Next
Run strong simple comparators
Establish time-updated physiology performance and the first governed haemodynamic / treatment-state context increment with prespecified evaluation.
Then
Expand context deliberately
Prioritise later phenotype/context blocks only after evidence, observability and specification review.
Conditional
Knowledge-informed methods
Introduce explicit external knowledge or neuro-symbolic mechanisms only when ordinary learning leaves a concrete residual problem.
Potential publication pathway
A three-paper framing for supervisor discussion, not a fixed commitment. Paper 1 focuses on methodological epidemiology and target ascertainment; Paper 2 uses a TraCeR-style longitudinal competing-risk survival model as the candidate backbone for context-aware prediction; Paper 3 is a neuro-symbolic extension testing whether governed directional and relational knowledge adds incremental value. The model architectures remain candidate framing, not current project authority.
Prospectively ascertainable severe AKI in routine EHR data
Candidate framing: how to define and evaluate a temporally valid severe-AKI prediction target when the clinical phenotype and its observation process evolve over time.
Research question
How should first future Stage ≥2 AKI be defined and ascertained prospectively when component availability, measurement timing, ICU transfer and terminal events vary over time?
Aims
Specify a prospective EHR endpoint with explicit information-time and UNKNOWN semantics.
Quantify how risk-set, timing, horizon, transfer and component choices change event support and ascertainment.
Separate observed non-event from outcome unascertainability and document transportability implications.
Potential novel contribution
A reproducible target-design framework that treats the observation process as part of the operational estimand rather than hidden preprocessing, with its epidemiological consequences quantified in longitudinal EHR data.
Can integrated clinical context enrich severe-AKI risk characterisation while preserving predictive validity?
Candidate framing: use a TraCeR-style longitudinal competing-risk survival Transformer as the candidate backbone for repeated prediction. The retrospective question is whether integrated context makes the risk representation more clinically meaningful while retaining strong predictive performance — not whether the model has already demonstrated clinical or operational utility.
Primary research question
Compared with a dynamic renal/physiology baseline, does adding integrated prospectively observable haemodynamic/treatment-state, infection/acute-deterioration, medication/pharmacology and procedure/exposure context preserve strong predictive validity while providing more clinically meaningful characterisation of future severe-AKI risk?
Aims
Evaluate a TraCeR-style longitudinal competing-risk candidate backbone against simpler adequate temporal comparators under the same prospective target and information set.
Compare an integrated-context model directly with a physiology-only model using the same longitudinal backbone.
Evaluate discrimination, calibration and probability quality, alongside whether context distinguishes clinically different states among patients with similar predicted risk.
Use prespecified domain attribution / ablation analyses to explain where contextual information contributes.
Report first-alert PPV, lead time and review burden only as secondary exploratory operationally relevant properties, not as proof of clinical utility.
Potential novel contribution
A context-aware longitudinal survival framework that tests whether clinically interpretable patient-state information can enrich severe-AKI risk characterisation without sacrificing predictive validity, while keeping retrospective predictive evidence distinct from claims of real-world clinical utility.
Analysis hierarchy: the primary paper-level comparison is integrated context versus dynamic physiology within the same candidate longitudinal-survival framework. Individual context domains are prespecified secondary attribution / ablation analyses. First-alert workload and lead-time analyses are exploratory operationally relevant properties, not evidence that deployment improves workflow or outcomes. “Preserving predictive validity” is conceptual here and should not be interpreted as formal non-inferiority unless a margin is prespecified.
Paper 3 · NeSy extension candidate
Directional & relational knowledge
Neuro-symbolic extension of the Paper 2 survival model
Does governed directional and relational clinical knowledge add value beyond the context-aware survival model?
Candidate framing: keep the Paper 2 target, cohort, observed patient information and TraCeR-style backbone fixed, then add explicit governed directional / relational knowledge through a NeSy / LTN-style extension so that the incremental contribution of knowledge itself can be tested.
Research question
Does injecting governed directional and relational clinical knowledge into the matched context-aware longitudinal survival model improve predictive validity, calibration, clinically meaningful risk characterisation or semantic consistency beyond the same observed patient information alone?
Aims
Represent clinically supported direction and relationship knowledge with explicit provenance, evidence strength, semantics and uncertainty.
Compare the conventional Paper 2 model with a matched NeSy / LTN extension using the same target, cohort, observed inputs and longitudinal backbone.
Evaluate predictive performance and calibration separately from rule satisfaction, semantic consistency and patient-level reasoning behaviour.
Use knowledge ablations to identify which relations, directions or weighting choices contribute to any observed change.
Potential novel contribution
A controlled test of whether explicit clinical direction and relationship knowledge provides incremental value beyond a strong context-aware longitudinal survival model, separating the effect of knowledge injection from the effect of additional patient information.
Method boundary: this is the proposed Paper 3 research framing, not acceptance of TraCeR, LTN or any NeSy architecture as current model authority. A separate model / knowledge-specification gate is still required, and added complexity must be justified against simpler adequate alternatives.
Framing boundary: Paper 1 is methodological epidemiology / phenotype-and-estimand work; Paper 2 is context-aware longitudinal prediction with applied clinical epidemiology; Paper 3 is a neuro-symbolic knowledge-injection extension. None is a causal-effect or prospective-utility study. Claims of clinical usefulness, workflow benefit or patient benefit require evidence beyond retrospective modelling. These contribution and novelty statements remain provisional until manuscript-specific external literature checks are completed.
Evidence trail for this version
The page separates live project authority, the Evidence Matrix, project literature notes and synthesis.
Live project governance checked first
Source access: Project governance files are stored in a private GitHub repository and require authorised access.
AKI context-priority decision — current context ordering: haemodynamic → infection → medication → procedure, with observation process treated as a cross-cutting validity layer.
Consulted 00_Read Me, C16_Clinical Phenotypes, C17_Phenotype Relationships and P19_Feature–Phenotype Crosswalk. The Matrix is an evidence-navigation/synthesis layer, not scientific authority.
REV-PHENO-01 — AKI operational definition, prospective EHR ascertainment and component-availability distinctions.
REV-ASCERT-01 — outcome unascertainability, informative observation and uncertain endpoints.
REV-INTUSE-03 — AKI alert / CDS intervention evidence and the limits of generic risk alerts.
REV-INTUSE-05 — temporal workflow, warning-time and cadence considerations.
REV-INTUSE-06 — MIMIC workflow observability and transportability constraints.
REV-TRANS-03 — critical methodology comparison and the requirement that added complexity address a demonstrated problem.
REV-INTUSE-07 — intended-use synthesis and clinical ownership / action architecture.
REV-MED-08 — prediction relevance and system placement of medication knowledge; action knowledge ≠ prediction knowledge.
Private reflective sources were not consulted. The phenotype relationship map is a supervisor-facing synthesis of the live Matrix and governed project evidence; it is not an executable ontology, feature contract or causal model.