Supervisor briefing · Evidence checked against live project sources · 28 Sep 2026
Evidence trail
PhD research

Towards Operationalisable Clinical Risk Prediction Models

AKI is the current clinical demonstrator. This PhD asks how a risk model can move from a valid prospective target, to a faithful representation of the evolving patient and clinical context, to an operationally meaningful signal that supports structured review without pretending that prognosis is the same as preventability or treatment response.

Current repository authority: HOLD AKI = current demonstrator Context = haemodynamic · infection · medication · procedure Risk score → structured context review, not treatment

Executive summary

One briefing, with a fast supervisor-level overview and deeper epidemiology / phenotype detail underneath.

Clinical problem

Future severe AKI

Predict clinically meaningful deterioration while preserving prospective information boundaries.

accepted framing
Intended decision

Prioritise structured review

The treating ICU / acute-care team reviews prospectively observable context and routes onward when warranted.

DEC-0084
Primary outcome

First prospectively ascertainable Stage ≥2

Persistent / severe trajectory outcomes remain secondary; the operational endpoint is observation-process dependent.

DEC-0087
Current authority

No globally accepted target, features or model

Candidate.8 is canonical and non-executable; patient-level target/model execution remains gated.

HOLD
Working research question
In adults at repeated ICU-origin prediction times, can prospectively observable clinical context improve the clinical meaning and operational usefulness of future severe-AKI risk prediction beyond a strong time-updated physiology baseline, and which contextual representations provide stable incremental value without compromising prospective validity?

Research framework

The thesis is organised around three linked methodological problems. Operationalisability is evaluated throughout rather than added as a final deployment chapter.

1 · Define

Define the prediction problem

Translate a clinical construct into a prospectively valid EHR prediction target with explicit risk-set, timing, component, missingness and terminal-event semantics.

Main AKI issue

Clinical AKI ≠ EHR-ascertainable AKI ≠ operational prediction endpoint.

→
2 · Represent

Represent the evolving patient and context

Move beyond a flat list of predictors to dynamic renal, physiological, treatment, exposure and observation-process states.

Main AKI issue

Context changes the meaning of the same observed value or medicine.

→
3 · Justify complexity

Add external knowledge only for a demonstrated residual problem

Test conventional representations first. Knowledge-informed or neuro-symbolic methods enter only if they address a specific limitation that remains.

Research design principle

Method follows the clinical and methodological requirement, not novelty.

Current context priority (DEC-0087): observed haemodynamic / treatment-state context first, infection / acute deterioration second, medication / pharmacology third, and procedure / surgery / contrast fourth. Observation process is a cross-cutting validity layer rather than a later competing context block. Each added context still requires its own evidence, prospective-observability and specification gate.

Epidemiological & prediction design

The detailed review portal is folded into this briefing here: who enters the risk set, when prediction is issued, what event is forecast, how follow-up works, and where uncertainty remains.

Who

Incident-risk state

Current prospectively ascertainable Stage 0 or Stage 1 can remain eligible for first future Stage ≥2 prediction.

DEC-0004
When

Repeated ICU-origin occasions

Rolling entry opens from ICU +6 h. q6 is the accepted primary model risk-refresh cadence; q12 is the prespecified cadence sensitivity.

DEC-0011 / 0012
What

First Stage ≥2

Forecast the first prospectively ascertainable KDIGO Stage ≥2 event; persistent / severe trajectory remains secondary.

DEC-0003 / 0006 / 0014
Horizon

24 h primary

The primary horizon is (t, t+24 h]; 48 h is secondary. A prespecified sensitivity excludes events in the first 6 h after prediction.

DEC-0013 / 0025
Follow-up

Hospital-wide issued horizon

Leaving the index ICU stops new prediction issuance, but an already-issued horizon continues under hospital follow-up. Death and discharge remain distinct terminal states.

DEC-0015 / 0016 / 0023
Decision evidence trail

Why these epidemiological design decisions?

This makes the reasoning auditable: external literature defines the methodological problem, project epidemiology quantifies the trade-off in MIMIC, and a separate human decision records the selected design and what remains uncertain.

partial design · project HOLD
1 · External evidenceClinical definitions, prediction methodology and observation-process literature.
2 · Project epidemiologyDevelopment-only EPID audits quantify cohort, timing and ascertainment consequences.
3 · Trade-offCompare event support, warning time, observability and operational consequences.
4 · Human decisionA bounded, time-stamped design choice is recorded in the decision register.
5 · Residual uncertaintyNamed limitations and reopening triggers remain visible rather than being silently resolved.
external / literature evidence project empirical evidence accepted partial decision remaining uncertainty
Design question Evidence examined What the evidence showed Current decision Remaining uncertainty
Who enters the risk set? EPID-001, EPID-002 + phenotype literature Stage 1 remained a meaningful at-risk state. At q6/24 h, Stage 0/1 eligibility captured 2,669 additional unique future Stage ≥2 events versus Stage 0 alone. Current prospectively ascertainable Stage 0 or Stage 1 is eligible for first future Stage ≥2 prediction. DEC-0004 Eligibility still depends on the final prospectively available-component phenotype.
When should prediction begin? EPID-006 + temporal/workflow review Governed prospective renal states can already be available at +6 h. A fixed +24 h landmark would impose a later entry point than the reviewed data require. Rolling eligibility opens from ICU +6 h; UNKNOWN is deferred rather than forced into Stage 0/1. DEC-0011 Not every patient is classifiable at +6 h; the final component set can alter early classifiability.
How often should risk update? EPID-002 + workflow literature q6 h is the highest-resolution reviewed rolling substrate. Model computation cadence is conceptually separate from clinical alert/review cadence. q6 h primary; q12 h prespecified sensitivity / efficiency analysis. DEC-0012 Whether q12 h gives similar practical performance with lower repeated-prediction burden remains empirical.
24 h or 48 h horizon? Rolling target-design audit + warning-time review 24 h captured 5,796 unique Stage ≥2 events versus 6,223 at 48 h (93.1% of 48 h coverage). Median warning was 11.1 h versus 19.7 h; 24 h also had fewer after-transfer positives and terminally shortened windows. 24 h primary; 48 h secondary. Retain a 0–6 h minimum-warning-time sensitivity. DEC-0013 The clinical value of extra warning versus tighter near-term prediction remains an explicit trade-off.
What counts as Stage ≥2? REV-PHENO-01, KDIGO / operationalisation literature + MIMIC phenotype evidence SCr, UO and KRT have different observation processes. An observed qualifying component can establish AKI; an unavailable component cannot be treated as criterion-negative. First prospectively EHR-ascertainable available-component KDIGO Stage ≥2; hospital-wide SCr Stage ≥2 retained as a major sensitivity / benchmark. DEC-0014 The complete executable target remains gated; component implementation and observation adequacy still require governed closure.
What happens at ICU transfer? EPID-002 + ascertainment / workflow literature 6.8% of q6/24 h Stage ≥2 positive windows occurred after ICU transfer. Transfer changes the observation regime but does not biologically end AKI risk. An already-issued horizon continues hospital-wide; no new ward predictions are created for the current MIMIC design. DEC-0015, DEC-0023 A future ward-compatible prediction model would require its own feature and prospective-availability contract.
What about death and discharge? EPID-002 + REV-ASCERT-01 / terminal-event reasoning Death and hospital discharge alive terminate the opportunity for the explicitly in-hospital endpoint and should not be collapsed into ordinary “no AKI” labels. Preserve death-before-Stage ≥2 and hospital-discharge-before-Stage ≥2 as distinct terminal states. DEC-0016 The later statistical model still has to specify how these terminal states enter learning and evaluation.
What exact future-time interval? Prospective prediction semantics + reproducible target-clock requirements An event already ascertainable at prediction time t is not a future outcome; a deterministic right-endpoint convention is needed for reproducibility. Use (t, t+H], with H = 24 h primary and 48 h secondary. DEC-0025 Main residual issue is source-specific timestamp precision / ordering, not the interval convention itself.
Interpretation rule: literature evidence, local empirical evidence and a project design decision are deliberately kept separate. The evidence supports and constrains a decision; it does not make the decision automatically. These are accepted partial design decisions, while CURRENT_AUTHORITY.yaml still records the project as HOLD with no globally accepted target, feature contract or model.
Figure 1. Prospective rolling-prediction design showing repeated ICU prediction times, prediction horizons, ICU transfer and terminal states.
Figure 1. Prospective rolling-prediction design. Rolling prediction eligibility opens from ICU +6 h, with q6 h as the primary model risk-refresh cadence and q12 h as a prespecified sensitivity. The primary horizon is (t, t+24 h], with 48 h secondary. New prediction issuance ends when the patient leaves the index ICU, while an already-issued prediction remains under in-hospital follow-up through its original horizon. These are accepted partial design decisions; the complete executable target remains gated.
Figure 2. Development ICU-stay substrate and q6 prediction-occasion structure showing 45,619 patients, 59,357 admissions, 65,813 valid ICU stays, 64,968 q6-contributing stays and 931,427 scheduled q6 prediction occasions.
Figure 2. Development ICU-stay substrate and q6 prediction-occasion structure. The current UO/q6 development substrate has 45,619 patients, 59,357 hospital admissions and 65,813 valid ICU stays. Of these, 64,968 ICU stays contribute at least one scheduled q6 occasion, generating 931,427 stay–time prediction occasions. These are confirmed development-substrate and scheduled-grid counts, not the final target-eligible AKI modelling cohort.
Why this is not just “predict AKI in the next 24 hours”

The phenotype must distinguish clinical KDIGO criteria from what was actually observable in the EHR at the prediction time. Baseline SCr availability, UO evaluability, KRT attribution, event-time versus availability-time, revisions, terminal events and zero-opportunity outcome ascertainment can all change what the label means.

Open epidemiological question — outcome unascertainability
Some prediction windows may contain no valid renal assessment, so Stage ≥2 status cannot be determined. These are retained as outcome unascertainable rather than classified as negative.

For discussion: if these windows are excluded from the primary binary analysis, does this introduce selection through the observation process because renal monitoring itself may be informative? Should this be addressed through reporting and sensitivity analyses, weighting, or explicit modelling of the observation process?
What remains non-authoritative / gated

The complete target is not globally accepted in CURRENT_AUTHORITY.yaml; candidate.8 remains canonical, aligned and non-executable. Global feature and model fields are also null. The page therefore distinguishes accepted partial decisions from candidate/specification state.

Current progress

A brief update on where the work is now and what needs to finish before the next stage.

project HOLD
Where I am

Urine-output data-quality check in progress

Most of the epidemiological study design is now in place. I am currently running a full check of the urine-output pipeline on the development dataset before incorporating it into the prediction framework.

Current bottleneck

The full check needs to finish and pass

The run is computationally slow and is still ongoing. I am not interpreting partial results. Once it finishes cleanly, I can assess whether the urine-output information is sufficiently available and reliable for the planned predictor set.

In short: the study design has moved forward; the immediate bottleneck is completing data processing and validation before integrating urine output into the next stage of the prediction pipeline.

Clinical phenotypes & relationships

Future severe AKIFirst prospectively ascertainable Stage ≥2 · operational EHR endpoint
Selected domain

Baseline vulnerability

Background susceptibility can alter future AKI risk without being the acute mechanism itself.

  • Chronic kidney disease / reduced renal reserve
  • History of prior AKI
  • Age ≥65 and selected chronic disease contexts
Relationship / guardrail

CKD → risk factor for AKI. Keep chronic baseline state separate from acute change; one elevated creatinine does not establish CKD.

Oliguriacriterion forAKI

Adequately measured low UO can establish AKI stage; missing UO cannot be treated as normal UO.

Hypotension≠Hypovolaemia

Low MAP/BP is haemodynamic information, not a deterministic label for intravascular depletion.

Medication exposure before AKI≠Drug-induced AKI

Temporal precedence is a patient fact, not causal attribution or preventability.

UNKNOWN component≠Observed negative

Absence of a valid assessment opportunity must remain explicit rather than silently becoming a negative label.

Hypovolaemiamodifies actionDiuretic

Depletion may support withholding/review, while congestion may make diuretic treatment appropriate: the same medicine can imply opposite actions.

EHR-ascertainable AKI≠Clinical / latent AKI truth

The operational phenotype depends on measurement coverage, timestamps and observation policy.

Prospective observability & temporal relationships

A value is usable only when its clinical event and its information availability are valid for the prediction occasion. This affects both the target and predictor-side context.

1 · Clinical event

A test, observation, drug administration, procedure or treatment state occurs.

2 · Information becomes available

The EHR can expose the information later than the biological or collection event.

3 · Prediction occasion t

Only information validly available by t may enter the model or current-state phenotype.

4 · Future window (t, t+H]

The outcome is sought prospectively after the prediction time.

5 · Ascertainability

Observed non-event is separated from zero valid opportunity to determine the endpoint.

Renal component

SCr, UO and KRT are non-equivalent signals

Each has different observation and timing properties. One observed qualifying component may establish AKI while another remains unavailable.

Context component

Observed proxy ≠ latent state

MAP, lactate, fluids and vasopressors may support haemodynamic interpretation but do not by themselves establish shock, hypovolaemia or a renal mechanism.

Transportability

Measurement policy can change model meaning

An EHR phenotype and its apparent missingness can shift when another hospital measures or documents differently.

How clinical context enters the prediction model

Context domains are added one at a time, with each block requiring its own clinical rationale, prospective-observability check and specification before comparison.

reference model

A · Time-updated patient state

Start with a strong renal and physiological representation using only prospectively valid observed patient information.

↓

Question: what can the evolving clinical state already explain?

context increment

B · Add one governed context domain

Add one clinically justified context block. Current priority is observed haemodynamic/treatment state, then infection/acute deterioration, then medication/pharmacology, then procedure/surgery/contrast.

↓

Each domain is tested separately before combining domains, so any incremental value remains interpretable.

contextual representation

C · Context × patient state × time

Test whether the same context becomes more informative when represented relative to renal trajectory, physiology, treatment state and time rather than as a flat feature list.

↓

Only after fair same-information comparisons should more complex or external-knowledge mechanisms be considered.

Current first context priority: observed haemodynamic / treatment-state context.
This is the first broader-context study to develop next, based on current feasibility, local support and clinical relevance.
priority 1
Observed haemodynamic / treatment state
BP/MAP trajectories, governed support state and other prospectively observed haemodynamic information; do not infer hypovolaemia or shock from one proxy.
priority 2
Infection / acute deterioration
Sepsis-related and evolving acuity context, with phenotype semantics kept separate from individual treatment or laboratory signals.
priority 3
Medication / pharmacology
Prospective medication exposure, dose/timing and medication × renal/physiological context.
priority 4
Procedure / surgery / contrast
Prospectively timed procedural and exposure context where clinically justified and locally supported.
cross-cutting validity
Observation process
Measurement recency, availability, coverage and UNKNOWN state accompany every model comparison rather than competing as a later context block.
later after validation
Latent volume / congestion / cause-specific states
Potentially high impact, but only after transparent phenotype validation; low MAP ≠ hypovolaemia and heart failure ≠ congestion.
Anti-conflation rule: a clinically valid relationship does not automatically become a predictor, a predictive interaction, a model constraint, or a treatment recommendation.

Evaluation: prediction + operational usefulness

The accepted evaluation architecture avoids reducing success to AUROC alone and keeps retrospective predictive claims separate from clinical-outcome benefit.

24 h

Primary horizon; 48 h retained as secondary/sensitivity.

ΔAUPRC

Primary discrimination estimand for the central contextual representation contrast; paired ΔAUROC is key secondary evidence.

Calibration

Reliability curve, calibration-in-the-large/intercept where compatible, slope and Brier score.

Matched burden

Compare PPV, sensitivity and workload at prespecified review-burden operating points rather than inventing an “optimal” cutoff.

First alert

Patient/stay-level first-alert PPV, sensitivity, lead time, unique patients alerted and repeated-alert burden.

2,000×

Paired subject-cluster bootstrap replicates for final uncertainty, preserving within-subject repeated observations.

Claim boundary: retrospective work may establish predictive increment and, with supporting operational evidence, an operationally promising increment. It cannot establish treatment benefit or clinical utility without prospective/interventional evaluation.

Research sequence

A gate-based plan is more defensible than promising a fixed number of papers or methods in advance.

Now

Close the target / phenotype path

Finish the remaining executable integration and validation gates for the prospectively ascertainable Stage ≥2 phenotype.

Next

Run strong simple comparators

Establish time-updated physiology performance and the first governed haemodynamic / treatment-state context increment with prespecified evaluation.

Then

Expand context deliberately

Prioritise later phenotype/context blocks only after evidence, observability and specification review.

Conditional

Knowledge-informed methods

Introduce explicit external knowledge or neuro-symbolic mechanisms only when ordinary learning leaves a concrete residual problem.

Potential publication pathway

A three-paper framing for supervisor discussion, not a fixed commitment. Paper 1 focuses on methodological epidemiology and target ascertainment; Paper 2 uses a TraCeR-style longitudinal competing-risk survival model as the candidate backbone for context-aware prediction; Paper 3 is a neuro-symbolic extension testing whether governed directional and relational knowledge adds incremental value. The model architectures remain candidate framing, not current project authority.

Paper 1 · current candidate

Target & ascertainment

Methodological epidemiology · phenotype / estimand / clinical informatics

Prospectively ascertainable severe AKI in routine EHR data

Candidate framing: how to define and evaluate a temporally valid severe-AKI prediction target when the clinical phenotype and its observation process evolve over time.

Research question

How should first future Stage ≥2 AKI be defined and ascertained prospectively when component availability, measurement timing, ICU transfer and terminal events vary over time?

Aims
  • Specify a prospective EHR endpoint with explicit information-time and UNKNOWN semantics.
  • Quantify how risk-set, timing, horizon, transfer and component choices change event support and ascertainment.
  • Separate observed non-event from outcome unascertainability and document transportability implications.
Potential novel contribution

A reproducible target-design framework that treats the observation process as part of the operational estimand rather than hidden preprocessing, with its epidemiological consequences quantified in longitudinal EHR data.

Paper 2 · next candidate

Context-aware longitudinal prediction

Methodological prediction + applied clinical epidemiology

Can integrated clinical context enrich severe-AKI risk characterisation while preserving predictive validity?

Candidate framing: use a TraCeR-style longitudinal competing-risk survival Transformer as the candidate backbone for repeated prediction. The retrospective question is whether integrated context makes the risk representation more clinically meaningful while retaining strong predictive performance — not whether the model has already demonstrated clinical or operational utility.

Primary research question

Compared with a dynamic renal/physiology baseline, does adding integrated prospectively observable haemodynamic/treatment-state, infection/acute-deterioration, medication/pharmacology and procedure/exposure context preserve strong predictive validity while providing more clinically meaningful characterisation of future severe-AKI risk?

Aims
  • Evaluate a TraCeR-style longitudinal competing-risk candidate backbone against simpler adequate temporal comparators under the same prospective target and information set.
  • Compare an integrated-context model directly with a physiology-only model using the same longitudinal backbone.
  • Evaluate discrimination, calibration and probability quality, alongside whether context distinguishes clinically different states among patients with similar predicted risk.
  • Use prespecified domain attribution / ablation analyses to explain where contextual information contributes.
  • Report first-alert PPV, lead time and review burden only as secondary exploratory operationally relevant properties, not as proof of clinical utility.
Potential novel contribution

A context-aware longitudinal survival framework that tests whether clinically interpretable patient-state information can enrich severe-AKI risk characterisation without sacrificing predictive validity, while keeping retrospective predictive evidence distinct from claims of real-world clinical utility.

Analysis hierarchy: the primary paper-level comparison is integrated context versus dynamic physiology within the same candidate longitudinal-survival framework. Individual context domains are prespecified secondary attribution / ablation analyses. First-alert workload and lead-time analyses are exploratory operationally relevant properties, not evidence that deployment improves workflow or outcomes. “Preserving predictive validity” is conceptual here and should not be interpreted as formal non-inferiority unless a margin is prespecified.
Paper 3 · NeSy extension candidate

Directional & relational knowledge

Neuro-symbolic extension of the Paper 2 survival model

Does governed directional and relational clinical knowledge add value beyond the context-aware survival model?

Candidate framing: keep the Paper 2 target, cohort, observed patient information and TraCeR-style backbone fixed, then add explicit governed directional / relational knowledge through a NeSy / LTN-style extension so that the incremental contribution of knowledge itself can be tested.

Research question

Does injecting governed directional and relational clinical knowledge into the matched context-aware longitudinal survival model improve predictive validity, calibration, clinically meaningful risk characterisation or semantic consistency beyond the same observed patient information alone?

Aims
  • Represent clinically supported direction and relationship knowledge with explicit provenance, evidence strength, semantics and uncertainty.
  • Compare the conventional Paper 2 model with a matched NeSy / LTN extension using the same target, cohort, observed inputs and longitudinal backbone.
  • Evaluate predictive performance and calibration separately from rule satisfaction, semantic consistency and patient-level reasoning behaviour.
  • Use knowledge ablations to identify which relations, directions or weighting choices contribute to any observed change.
Potential novel contribution

A controlled test of whether explicit clinical direction and relationship knowledge provides incremental value beyond a strong context-aware longitudinal survival model, separating the effect of knowledge injection from the effect of additional patient information.

Method boundary: this is the proposed Paper 3 research framing, not acceptance of TraCeR, LTN or any NeSy architecture as current model authority. A separate model / knowledge-specification gate is still required, and added complexity must be justified against simpler adequate alternatives.
Framing boundary: Paper 1 is methodological epidemiology / phenotype-and-estimand work; Paper 2 is context-aware longitudinal prediction with applied clinical epidemiology; Paper 3 is a neuro-symbolic knowledge-injection extension. None is a causal-effect or prospective-utility study. Claims of clinical usefulness, workflow benefit or patient benefit require evidence beyond retrospective modelling. These contribution and novelty statements remain provisional until manuscript-specific external literature checks are completed.

Evidence trail for this version

The page separates live project authority, the Evidence Matrix, project literature notes and synthesis.

Live project governance checked first

Source access: Project governance files are stored in a private GitHub repository and require authorised access.

AKI Evidence Matrix — live workbook

Consulted 00_Read Me, C16_Clinical Phenotypes, C17_Phenotype Relationships and P19_Feature–Phenotype Crosswalk. The Matrix is an evidence-navigation/synthesis layer, not scientific authority.

Open the live Evidence Matrix

Project literature notes consulted
  • REV-PHENO-01 — AKI operational definition, prospective EHR ascertainment and component-availability distinctions.
  • REV-ASCERT-01 — outcome unascertainability, informative observation and uncertain endpoints.
  • REV-INTUSE-03 — AKI alert / CDS intervention evidence and the limits of generic risk alerts.
  • REV-INTUSE-05 — temporal workflow, warning-time and cadence considerations.
  • REV-INTUSE-06 — MIMIC workflow observability and transportability constraints.
  • REV-TRANS-03 — critical methodology comparison and the requirement that added complexity address a demonstrated problem.
  • REV-INTUSE-07 — intended-use synthesis and clinical ownership / action architecture.
  • REV-MED-08 — prediction relevance and system placement of medication knowledge; action knowledge ≠ prediction knowledge.

Private reflective sources were not consulted. The phenotype relationship map is a supervisor-facing synthesis of the live Matrix and governed project evidence; it is not an executable ontology, feature contract or causal model.