Skip to content
FlowScript templates

Diagnostics — ROC, calibration and decision curve

External validation of a clinical risk model in three figures: discrimination (ROC, AUC and a paired DeLong test), calibration (slope, intercept, Brier score) and clinical usefulness (decision-curve net benefit). Typed FlowScript for clinical trials. Keywords: CONSORT, randomised, randomized, trial, RCT, flow diagram, PRISMA, cohort.

Template previewFlowScript
ROC curve — Discrimination — TRIAGE-S versus qSOFA, external validation cohortFigure: Discrimination — TRIAGE-S versus qSOFA, external validation cohort. Line chart. Horizontal axis: 1 − specificity (false-positive rate) (probability), linear scale, from 0 to 1. Vertical axis: Sensitivity (true-positive rate) (probability), linear scale, from 0 to 1. 2 series. TRIAGE-S predicted risk (0.808646 area under the curve, interval 0.7335 to 0.883792, n = 160); qSOFA score (0 to 3) (0.73125 area under the curve, interval 0.646513 to 0.815987, n = 160). Observations: n = 160 subjects. Uncertainty: 95% confidence intervals. Comparison: TRIAGE-S predicted risk is higher than qSOFA score (0 to 3) (absolute difference 0.0773958; 95% CI 0.00887557 to 0.145916; DeLong paired test for two correlated areas; p = 0.027). Annotations: Chance diagonal at 0.5 (reference line). Source: Prevalence 25.0%; statistics computed from the per-subject scores..Discrimination — TRIAGE-S versus qSOFA, external validation cohort160 subjects, 40 with the outcome (25.0% prevalence) · 133 distinct scores · 95% intervalsChance (AUC 0.50)Chance (AUC 0.50)TRIAGE-S predicted risk: AUC 0.809 over 160 subjectsqSOFA score (0 to 3): AUC 0.731 over 160 subjectsYouden-optimal cut-off 0.271: sensitivity 0.750, specificity 0.775TRIAGE-S predicted risk — AUC 0.81 (95% CI 0.73 to 0.88)qSOFA score (0 to 3) — AUC 0.73 (95% CI 0.65 to 0.82)ΔAUC 0.077 (95% CI 0.009 to 0.146), DeLong P = 0.027Diamond: cut-off 0.27 (Youden J 0.53)0.000.250.500.751.001 − specificity (false-positive rate)0.000.250.500.751.00Sensitivity (true-positive rate)— Some scores are shared by a positive and a negative subject. Those ties are counted as ½ of a concordant pair, which is what makes thetrapezoidal area equal to the Mann–Whitney U statistic; the curve is diagonal across them.— Youden's J assumes a false positive and a false negative cost the same. If they do not, use decisionCurve: its threshold probability ISthe cost ratio.— Area and interval computed from the scores by DeLong et al. (1988); the drawn curve is joined with straight segments, so the trapezoidunder it IS the area printed.

Make it your own.

// External validation of a clinical risk model in three figures: discrimination (ROC, AUC and a paired DeLong test), calibration (slope, intercept, Brier score) and clinical usefulness (decision-curve net benefit).
//
// Every string below is EXAMPLE text from a fictional study: replace it.
//
// TRIPOD asks for discrimination and calibration, and for a measure of
// clinical usefulness when a model is meant to guide decisions, because
// the three answer different questions. A model can rank patients well
// (high AUC) and still be systematically too confident, which only the
// calibration plot shows; and a model can be accurate and calibrated and
// still not change a decision, which only the decision curve shows.
//
// The data are one row per patient: x is the model's predicted
// probability of sepsis within 48 hours of triage, y is the outcome
// adjudicated by two physicians blind to the score (1 = sepsis). Every
// statistic on the three figures is computed from these vectors, so each
// figure carries its own copy of the same 160 patients: a series belongs
// to the block it is written inside. The second ROC series is qSOFA on the
// same patients; it omits y, inherits the first series' outcomes, and that
// pairing is what licenses the DeLong test between the two curves.

roc triage_s_roc {
  title: "Discrimination — TRIAGE-S versus qSOFA, external validation cohort"
  model: triage_s
  validation: external
  n: 160
  events: 40
  level: 0.95
  series triage_s_scores {
    label: "TRIAGE-S predicted risk"
    x: [
      0.108, 0.061, 0.12, 0.615, 0.054, 0.122, 0.041, 0.544, 0.162, 0.301, 0.282, 0.338, 0.044, 0.149, 0.187, 0.513,
      0.12, 0.08, 0.05, 0.423, 0.685, 0.028, 0.028, 0.018, 0.246, 0.457, 0.301, 0.299, 0.13, 0.053, 0.013, 0.294,
      0.063, 0.028, 0.089, 0.025, 0.027, 0.153, 0.021, 0.047, 0.041, 0.085, 0.104, 0.038, 0.028, 0.023, 0.095, 0.082,
      0.485, 0.155, 0.568, 0.481, 0.574, 0.037, 0.076, 0.524, 0.154, 0.855, 0.253, 0.018, 0.352, 0.337, 0.238, 0.197,
      0.241, 0.006, 0.03, 0.101, 0.04, 0.139, 0.01, 0.689, 0.097, 0.454, 0.005, 0.503, 0.153, 0.772, 0.04, 0.166,
      0.201, 0.725, 0.203, 0.128, 0.082, 0.047, 0.064, 0.153, 0.45, 0.391, 0.519, 0.012, 0.016, 0.663, 0.164, 0.812,
      0.05, 0.271, 0.05, 0.14, 0.049, 0.289, 0.105, 0.354, 0.11, 0.242, 0.817, 0.529, 0.201, 0.38, 0.054, 0.27,
      0.332, 0.064, 0.285, 0.423, 0.441, 0.127, 0.038, 0.056, 0.017, 0.327, 0.011, 0.421, 0.138, 0.508, 0.444, 0.294,
      0.054, 0.182, 0.636, 0.352, 0.018, 0.474, 0.189, 0.191, 0.206, 0.372, 0.041, 0.37, 0.355, 0.068, 0.046, 0.029,
      0.087, 0.037, 0.008, 0.297, 0.022, 0.792, 0.179, 0.068, 0.024, 0.273, 0.226, 0.688, 0.254, 0.751, 0.301, 0.015
    ]
    y: [
      0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 1, 0, 0, 1,
      1, 0, 0, 0, 0, 1, 1, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,
      0, 0, 0, 1, 0, 0, 1, 0, 1, 0, 0, 1, 0, 0, 0, 1, 0, 0, 0, 0,
      1, 1, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0,
      0, 1, 0, 1, 0, 0, 0, 0, 0, 0, 1, 0, 0, 1, 1, 1, 0, 1, 0, 0,
      1, 1, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0, 0,
      0, 1, 0, 0, 0, 1, 1, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 1,
      0, 0, 0, 0, 0, 0, 0, 1, 0, 1, 1, 0, 0, 0, 0, 1, 0, 1, 0, 0
    ]
  }
  series qsofa_scores {
    label: "qSOFA score (0 to 3)"
    x: [
      1, 1, 2, 1, 0, 1, 1, 3, 1, 2, 1, 0, 0, 0, 1, 2, 1, 0, 1, 2,
      3, 0, 0, 0, 2, 1, 0, 1, 0, 0, 0, 2, 0, 0, 0, 0, 0, 0, 0, 2,
      0, 1, 0, 0, 0, 0, 1, 0, 1, 2, 2, 2, 2, 0, 1, 1, 1, 2, 2, 0,
      0, 2, 1, 2, 2, 0, 0, 1, 1, 2, 0, 2, 0, 3, 0, 3, 1, 3, 0, 0,
      0, 2, 0, 0, 0, 1, 0, 1, 2, 1, 1, 0, 0, 3, 2, 3, 0, 1, 0, 0,
      1, 1, 1, 2, 1, 0, 3, 2, 1, 3, 0, 0, 2, 0, 1, 3, 3, 1, 0, 0,
      0, 3, 0, 2, 1, 3, 3, 1, 0, 2, 2, 2, 2, 2, 1, 2, 0, 2, 0, 3,
      2, 1, 0, 0, 0, 0, 0, 2, 0, 2, 1, 0, 0, 2, 1, 2, 3, 3, 1, 0
    ]
  }
}

calibration triage_s_cal {
  title: "Calibration — TRIAGE-S, external validation cohort"
  model: triage_s
  validation: external
  method: decile
  bins: 8
  n: 160
  events: 40
  series triage_s_cal_scores {
    label: "TRIAGE-S predicted risk"
    x: [
      0.108, 0.061, 0.12, 0.615, 0.054, 0.122, 0.041, 0.544, 0.162, 0.301, 0.282, 0.338, 0.044, 0.149, 0.187, 0.513,
      0.12, 0.08, 0.05, 0.423, 0.685, 0.028, 0.028, 0.018, 0.246, 0.457, 0.301, 0.299, 0.13, 0.053, 0.013, 0.294,
      0.063, 0.028, 0.089, 0.025, 0.027, 0.153, 0.021, 0.047, 0.041, 0.085, 0.104, 0.038, 0.028, 0.023, 0.095, 0.082,
      0.485, 0.155, 0.568, 0.481, 0.574, 0.037, 0.076, 0.524, 0.154, 0.855, 0.253, 0.018, 0.352, 0.337, 0.238, 0.197,
      0.241, 0.006, 0.03, 0.101, 0.04, 0.139, 0.01, 0.689, 0.097, 0.454, 0.005, 0.503, 0.153, 0.772, 0.04, 0.166,
      0.201, 0.725, 0.203, 0.128, 0.082, 0.047, 0.064, 0.153, 0.45, 0.391, 0.519, 0.012, 0.016, 0.663, 0.164, 0.812,
      0.05, 0.271, 0.05, 0.14, 0.049, 0.289, 0.105, 0.354, 0.11, 0.242, 0.817, 0.529, 0.201, 0.38, 0.054, 0.27,
      0.332, 0.064, 0.285, 0.423, 0.441, 0.127, 0.038, 0.056, 0.017, 0.327, 0.011, 0.421, 0.138, 0.508, 0.444, 0.294,
      0.054, 0.182, 0.636, 0.352, 0.018, 0.474, 0.189, 0.191, 0.206, 0.372, 0.041, 0.37, 0.355, 0.068, 0.046, 0.029,
      0.087, 0.037, 0.008, 0.297, 0.022, 0.792, 0.179, 0.068, 0.024, 0.273, 0.226, 0.688, 0.254, 0.751, 0.301, 0.015
    ]
    y: [
      0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 1, 0, 0, 1,
      1, 0, 0, 0, 0, 1, 1, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,
      0, 0, 0, 1, 0, 0, 1, 0, 1, 0, 0, 1, 0, 0, 0, 1, 0, 0, 0, 0,
      1, 1, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0,
      0, 1, 0, 1, 0, 0, 0, 0, 0, 0, 1, 0, 0, 1, 1, 1, 0, 1, 0, 0,
      1, 1, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0, 0,
      0, 1, 0, 0, 0, 1, 1, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 1,
      0, 0, 0, 0, 0, 0, 0, 1, 0, 1, 1, 0, 0, 0, 0, 1, 0, 1, 0, 0
    ]
  }
}

decision_curve triage_s_dca {
  title: "Net benefit — act on TRIAGE-S, treat all, or treat none"
  model: triage_s
  validation: external
  thresholds: [0.05, 0.075, 0.1, 0.125, 0.15, 0.175, 0.2, 0.225, 0.25, 0.275, 0.3, 0.325, 0.35, 0.375, 0.4]
  n: 160
  events: 40
  series triage_s_dca_scores {
    label: "TRIAGE-S predicted risk"
    x: [
      0.108, 0.061, 0.12, 0.615, 0.054, 0.122, 0.041, 0.544, 0.162, 0.301, 0.282, 0.338, 0.044, 0.149, 0.187, 0.513,
      0.12, 0.08, 0.05, 0.423, 0.685, 0.028, 0.028, 0.018, 0.246, 0.457, 0.301, 0.299, 0.13, 0.053, 0.013, 0.294,
      0.063, 0.028, 0.089, 0.025, 0.027, 0.153, 0.021, 0.047, 0.041, 0.085, 0.104, 0.038, 0.028, 0.023, 0.095, 0.082,
      0.485, 0.155, 0.568, 0.481, 0.574, 0.037, 0.076, 0.524, 0.154, 0.855, 0.253, 0.018, 0.352, 0.337, 0.238, 0.197,
      0.241, 0.006, 0.03, 0.101, 0.04, 0.139, 0.01, 0.689, 0.097, 0.454, 0.005, 0.503, 0.153, 0.772, 0.04, 0.166,
      0.201, 0.725, 0.203, 0.128, 0.082, 0.047, 0.064, 0.153, 0.45, 0.391, 0.519, 0.012, 0.016, 0.663, 0.164, 0.812,
      0.05, 0.271, 0.05, 0.14, 0.049, 0.289, 0.105, 0.354, 0.11, 0.242, 0.817, 0.529, 0.201, 0.38, 0.054, 0.27,
      0.332, 0.064, 0.285, 0.423, 0.441, 0.127, 0.038, 0.056, 0.017, 0.327, 0.011, 0.421, 0.138, 0.508, 0.444, 0.294,
      0.054, 0.182, 0.636, 0.352, 0.018, 0.474, 0.189, 0.191, 0.206, 0.372, 0.041, 0.37, 0.355, 0.068, 0.046, 0.029,
      0.087, 0.037, 0.008, 0.297, 0.022, 0.792, 0.179, 0.068, 0.024, 0.273, 0.226, 0.688, 0.254, 0.751, 0.301, 0.015
    ]
    y: [
      0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 1, 0, 0, 1,
      1, 0, 0, 0, 0, 1, 1, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,
      0, 0, 0, 1, 0, 0, 1, 0, 1, 0, 0, 1, 0, 0, 0, 1, 0, 0, 0, 0,
      1, 1, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1, 0, 0,
      0, 1, 0, 1, 0, 0, 0, 0, 0, 0, 1, 0, 0, 1, 1, 1, 0, 1, 0, 0,
      1, 1, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 1, 1, 1, 0, 0, 0,
      0, 1, 0, 0, 0, 1, 1, 0, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 1,
      0, 0, 0, 0, 0, 0, 0, 1, 0, 1, 1, 0, 0, 0, 0, 1, 0, 1, 0, 0
    ]
  }
}

view discrimination: roc(triage_s_roc)
view calibration_plot: calibration(triage_s_cal)
view net_benefit: decision_curve(triage_s_dca)