Supervisor briefing · Working research plan · 25 Sep 2026
Evidence basis
PhD research direction

Towards Operationalisable Clinical Risk Prediction Models

A cumulative programme using acute kidney injury as the principal demonstrator: first define a prospectively valid prediction problem, then represent the evolving patient state, and only then add structured clinical knowledge where the data show a specific limitation.

Current project authority: HOLD AKI = principal demonstrator Medication = important context, not the sole core

Executive summary

The thesis is being narrowed to three linked methodological problems rather than five mandatory standalone studies.

Problem 1

Define the right target

Clinical AKI criteria are not automatically valid prediction-time labels. The first task is a prospectively ascertainable and reproducible endpoint.

Problem 2

Represent the evolving patient state

Test what temporal modelling and broader clinical context add beyond current-state predictors. Medication is one nested context family.

Problem 3

Add knowledge only when justified

Structured clinical knowledge becomes a modelling intervention only if data-driven learning shows a concrete residual problem such as sparsity or poor generalisation.

Unifying thesis question
How can clinical risk prediction models be designed so that their targets, longitudinal patient-state representations, use of external clinical knowledge, and evaluation are aligned with real clinical decision-making?

Three main empirical chapters

Operationalisability is evaluated throughout the thesis rather than being left to a separate final deployment chapter.

Chapter 1

Define the prediction problem

Prospective operationalisation of severe AKI in longitudinal EHR data.

Key question

What are we predicting, and when can it validly be said to have occurred using information available at prediction time?

→
Chapter 2

Represent the evolving patient state

Dynamic clinical-context modelling: renal trajectory, physiology, care context, medication and other supported longitudinal context.

Key question

Which evolving contexts add useful predictive information, and how should they be represented over time?

→
Chapter 3

Add knowledge where data are insufficient

Knowledge-informed learning for sparse or difficult clinical contexts, conditional on a demonstrated need.

Key question

Can structured clinical knowledge improve robustness, generalisation, calibration or interpretability?

Where medication fits

Medication is a strategically important clinical-context family because it is dynamic, context-dependent, clinically actionable and often sparse at specific drug × state combinations. It remains a medication-safety test case within Chapter 2 and potentially Chapter 3, but it does not define the entire thesis.

A: clinical / physiological comparator B_RAW: + prospective medication facts B_CONTEXT: + medication × physiology × time

Immediate focus: Paper 1

Working study: prospective operationalisation of severe AKI as a prediction endpoint in longitudinal EHR data.

Core research question

Clinical phenotype ≠ prediction-time phenotype

How should severe AKI be operationalised as a prospectively ascertainable prediction endpoint, and how materially do information timing, component observability and outcome-ascertainment choices change the resulting prediction problem?

Why it matters

Target construction affects model validity

Retrospective labels can use information that was unavailable when a prediction would have been made. Routine-care monitoring is also uneven, so “no observed AKI” is not always equivalent to a well-observed negative.

RQ1
Endpoint construction: define first prospectively ascertainable severe AKI with explicit information-time semantics.
RQ2
Baseline & availability: quantify how baseline-SCr and prospective information choices change risk-set membership and event ascertainment.
RQ3
Component contribution: compare SCr-only with prospectively observable SCr/UO/KRT.
RQ4
Outcome ascertainability: distinguish a prospectively observed non-event from a horizon with zero valid assessment opportunity.
Status: this is a planned methodology study. Current canonical authority remains on HOLD; the full target and model are not yet accepted, and the integrated patient-level SCr/UO/KRT phenotype has not yet been authorised for final execution.

Current evidence and progress

Selected decision-relevant findings only. Numerical results below are development-only SCr audits, not final Paper 1 results.

46.0%
of the broader q6/24 h future severe-event set was omitted by Stage-0-only eligibility versus Stage 0/1 in EPID-002.
94.9%
of q6/24 h unique events were retained by q12, using 50.9% as many prediction rows in EPID-002.
~47%
historical 8–365 d SCr comparator availability at +6, +12 and +24 h in EPID-006.
+905
24 h future SCr Stage ≥2 events gained at +24 h when moving from history-required to any valid SCr pathway in EPID-006.
Phenotype substrate

Substantial groundwork is already complete

  • Stage 0/1 rolling-risk design has been empirically explored in development-only SCr audits.
  • Baseline-SCr missing-history handling has been quantified and shown to materially alter eligibility and event capture.
  • Outcome unascertainability has a literature-backed conceptual framework distinct from ordinary predictor missingness.
Still pending

Integrated phenotype and later modelling

  • UO has extensive governed static/synthetic implementation, but final integrated patient-level target execution remains blocked.
  • KRT maintenance-exclusion semantics are specified, but patient-level execution remains gated.
  • Neutral medication representation is still being validated before any definitive medication-aware modelling comparison.
Interpretation of the development audits

EPID-002 suggests that waiting longer improves renal-state interpretability but loses some early prediction opportunity; q12 is an efficient comparator to q6; and 24 h versus 48 h remains a trade-off between event proximity, warning time, repeated labels and care-transition exposure. EPID-006 shows that the handling of missing historical SCr can substantially change who is classifiable and eligible for prediction. These audits inform the final design but do not themselves define it.

Indicative 2.5-year roadmap

The plan is organised around three defensible thesis chapters, not five guaranteed publications.

~0–6 months

Complete phenotype

Resolve remaining target gates, execute the integrated phenotype, complete Paper 1 analyses and move into manuscript drafting/submission.

~6–16 months

Dynamic clinical context

Build strong simple comparators first, then test longitudinal modelling and incremental context families. Medication is one nested experiment.

~16–24 months

Knowledge-informed study

Proceed only if Chapter 2 reveals a concrete residual problem. Compare conventional structured-learning solutions before neuro-symbolic methods.

~24–30 months

Integrate & write

Complete thesis integration, operational evaluation across chapters, remaining manuscripts, revisions and viva preparation.

Scope rule: promising ideas can remain future work. The programme should add methodological complexity only when a demonstrated research requirement justifies it.
Evidence basis used for this briefing
  • governance/CURRENT_AUTHORITY.yaml — live canonical project authority checked before drafting.
  • governance/RESEARCH_DESIGN_PRINCIPLES.md — programme identity, method-selection and scope principles.
  • papers/THESIS_3_CHAPTER_FRAMING_DRAFT_2026-09-24.md — current corrected thesis framing.
  • papers/paper1/PROTOCOL_DRAFT_2026-09-24.md — current Paper 1 research plan.
  • reports/qc/EPID-002_rolling_aki_target_design_epidemiology.md — development-only rolling SCr target-design audit.
  • reports/qc/EPID-006_missing_historical_baseline_handling.md — development-only baseline-SCr handling audit.
  • REV-PHENO-01 and REV-ASCERT-01 — literature/methodology evidence for prospective operationalisation and outcome ascertainment.

Private reflective sources were not consulted. This page is a supervisor-facing planning summary, not a scientific authority document.