Skip to content
Flow metrics (CFD) templates

CFD — ML Experiment Queue

Twelve days of a shared GPU cluster's training queue, contrasting a 4.5-day committed cycle time with a submission backlog that grows tenfold, and breaching the concurrent-training limit.

Template previewFlow metrics (CFD)
Model training queue — shared GPU cluster0100200300cumulative runsQueued growingTraining limit 162026-02-022026-02-032026-02-042026-02-052026-02-062026-02-072026-02-082026-02-092026-02-102026-02-112026-02-122026-02-13Queued (queue)Data prepTrainingEval (queue)ShippedWIP40 runscommitted at Data prepTHROUGHPUT62.4 / week8.91 runs/dayCYCLE TIME ≈4.5 daysLittle's Law: WIP ÷ throughputLEAD TIME ≈15.4 daysall 137 runs in systemFLOW EFFICIENCY75%30 active · 10 waitingDELIVERY TREND↑ 32%accelerating · 98 doneStageTypeWIP (runs)LimitTime (days)Queuedwaiting97—10.9Data prepactive11—1.2Trainingactive19162.1Evalwaiting10—1.1Shippeddone98——!Queue growth: Queued widened for 11 straight periods (2026-02-02 → 2026-02-13, +88 runs)!Training holds 19 runs against a WIP limit of 16!Arrivals outran deliveries by 106 runs across the window

Make it your own.

title "Model training queue — shared GPU cluster"
unit runs
stages: Queued*, Data prep, Training, Eval*, Shipped
commit Data prep
wip-limit "Training" 16

# Cumulative runs that have entered each stage. The Queued band is the
# cluster's admission backlog; cycle time is measured from Data prep.
2026-02-02 |  31,  22,  16,   4,   0
2026-02-03 |  44,  32,  25,  11,   6
2026-02-04 |  56,  40,  32,  19,  13
2026-02-05 |  70,  49,  42,  27,  21
2026-02-06 |  87,  60,  51,  37,  30
2026-02-07 | 103,  69,  61,  45,  38
2026-02-08 | 122,  80,  70,  55,  47
2026-02-09 | 142,  91,  82,  65,  57
2026-02-10 | 163, 102,  91,  75,  66
2026-02-11 | 185, 113, 103,  85,  76
2026-02-12 | 210, 126, 114,  97,  87
2026-02-13 | 235, 138, 127, 108,  98