CFD — ML Experiment Queue
Twelve days of a shared GPU cluster's training queue, contrasting a 4.5-day committed cycle time with a submission backlog that grows tenfold, and breaching the concurrent-training limit.
Make it your own.
title "Model training queue — shared GPU cluster"
unit runs
stages: Queued*, Data prep, Training, Eval*, Shipped
commit Data prep
wip-limit "Training" 16
# Cumulative runs that have entered each stage. The Queued band is the
# cluster's admission backlog; cycle time is measured from Data prep.
2026-02-02 | 31, 22, 16, 4, 0
2026-02-03 | 44, 32, 25, 11, 6
2026-02-04 | 56, 40, 32, 19, 13
2026-02-05 | 70, 49, 42, 27, 21
2026-02-06 | 87, 60, 51, 37, 30
2026-02-07 | 103, 69, 61, 45, 38
2026-02-08 | 122, 80, 70, 55, 47
2026-02-09 | 142, 91, 82, 65, 57
2026-02-10 | 163, 102, 91, 75, 66
2026-02-11 | 185, 113, 103, 85, 76
2026-02-12 | 210, 126, 114, 97, 87
2026-02-13 | 235, 138, 127, 108, 98