DSM — Pre-Training Parameter Coupling
Parameter DSM of the fourteen design choices behind a foundation-model pre-training run, read in inputs-in-columns convention and taken in config-file order, computing that 12 of the 31 couplings run backwards there and resolving the parameters into three blocks: data mixture with vocabulary, context and evals; hidden width with layer depth; and the sharding/optimiser/batch/checkpointing memory loop.
Make it your own.
title "Foundation model pre-training — parameter coupling"
mode parameter
convention ic-fbd
# Listed in the order the training config file declares them: the model
# block, then the optimiser block, then the data block at the bottom.
parameters: Hidden width, Layer depth, Attention heads, Context length, Vocabulary size
parameters: Global batch size, Learning rate, Step count, Optimiser state, Activation checkpoint
parameters: Sharding plan, Token budget, Data mixture, Eval suite
Data mixture <- Token budget(2), Eval suite(2)
Vocabulary size <- Data mixture(3)
Context length <- Data mixture(2), Token budget
Hidden width <- Token budget(3), Layer depth(2)
Layer depth <- Token budget(3), Hidden width(2)
Attention heads <- Hidden width(3), Context length(2)
Global batch size <- Token budget(2), Sharding plan(2), Hidden width
Learning rate <- Global batch size(3), Hidden width(2), Layer depth
Optimiser state <- Hidden width(3), Layer depth(3), Sharding plan(2)
Sharding plan <- Hidden width(2), Layer depth(2), Optimiser state(3), Activation checkpoint(2)
Activation checkpoint <- Context length(3), Hidden width(2), Global batch size(2)
Step count <- Token budget(3), Global batch size(3)
Eval suite <- Vocabulary size, Context length(2)