Skip to content
Design structure matrix templates

DSM — Pre-Training Parameter Coupling

Parameter DSM of the fourteen design choices behind a foundation-model pre-training run, read in inputs-in-columns convention and taken in config-file order, computing that 12 of the 31 couplings run backwards there and resolving the parameters into three blocks: data mixture with vocabulary, context and evals; hidden width with layer depth; and the sharding/optimiser/batch/checkpointing memory loop.

Template previewDesign structure matrix
Foundation model pre-training — parameter couplingParameter DSM · IC/FBD — read down a column for that element's inputs; marks below the diagonal are feedback.As declared223232323222323322232322332Hidden width1Layer depth2Attention heads3Context length4Vocabulary size5Global batch size6Learning rate7Step count8Optimiser state9Activation checkpoint10Sharding plan11Token budget12Data mixture13Eval suite14Hidden widthLayer depthAttention headsContext lengthVocabulary sizeGlobal batch sizeLearning rateStep countOptimiser stateActivation checkpointSharding planToken budgetData mixtureEval suitePartitioned sequence223232323222323322232322332123Token budget1Hidden width2Layer depth3Data mixture4Context length5Vocabulary size6Eval suite7Attention heads8Global batch size9Activation checkpoint10Optimiser state11Sharding plan12Learning rate13Step count14Token budgetHidden widthLayer depthData mixtureContext lengthVocabulary sizeEval suiteAttention headsGlobal batch sizeActivation checkpointOptimiser stateSharding planLearning rateStep countinout062724222311212033313142302012 feedback marks reduced to 4 across 3 iteration blocks.14 parameters · 31 dependency marks · density 17% · 4 sequenced stages · largest block 4Sequence runs top-left to bottom-right; 4 remaining feedback marks sit inside the boxed blocks.Highest fan-out: Hidden width (feeds 7) · highest fan-in: Sharding plan (needs 4)Block 1 · positions 2–3 · Hidden width, Layer depthBlock 2 · positions 4–7 · Data mixture, Context length, Vocabulary size, Eval suiteBlock 3 · positions 9–12 · Global batch size, Activation checkpoint, Optimiser state, Sharding plandependencystrength 3feedback (rework)iteration blockdiagonal (self)

Make it your own.

title "Foundation model pre-training — parameter coupling"
mode parameter
convention ic-fbd

# Listed in the order the training config file declares them: the model
# block, then the optimiser block, then the data block at the bottom.
parameters: Hidden width, Layer depth, Attention heads, Context length, Vocabulary size
parameters: Global batch size, Learning rate, Step count, Optimiser state, Activation checkpoint
parameters: Sharding plan, Token budget, Data mixture, Eval suite

Data mixture          <- Token budget(2), Eval suite(2)
Vocabulary size       <- Data mixture(3)
Context length        <- Data mixture(2), Token budget
Hidden width          <- Token budget(3), Layer depth(2)
Layer depth           <- Token budget(3), Hidden width(2)
Attention heads       <- Hidden width(3), Context length(2)
Global batch size     <- Token budget(2), Sharding plan(2), Hidden width
Learning rate         <- Global batch size(3), Hidden width(2), Layer depth
Optimiser state       <- Hidden width(3), Layer depth(3), Sharding plan(2)
Sharding plan         <- Hidden width(2), Layer depth(2), Optimiser state(3), Activation checkpoint(2)
Activation checkpoint <- Context length(3), Hidden width(2), Global batch size(2)
Step count            <- Token budget(3), Global batch size(3)
Eval suite            <- Vocabulary size, Context length(2)