Flame Graph — PyTorch Training Loop, Main Process (py-spy)
Where the Python main process of a ResNet-50 training run spends its time (4,760 samples): 45% blocked waiting for the DataLoader queue, so the GPU is starved by JPEG decoding in too few workers, and a further 15.1% stalled in a per-step loss.item() that forces a CUDA synchronise. Forward, backward and the optimiser together are only 33.5%. Illustrative values.
Make it your own.
title "ResNet-50 training — main process, py-spy 4,760 samples (illustrative)"
# py-spy record --native --rate 50 -d 60 --format raw -- python train.py --workers 4
python;train.py:main;train.py:train_epoch;torch.utils.data:DataLoader.__next__;torch.utils.data:_MultiProcessingDataLoaderIter._get_data;queue:Queue.get 2140
python;train.py:main;train.py:train_epoch;torch.nn:Module._call_impl;torchvision:ResNet.forward;cudaLaunchKernel 610
python;train.py:main;train.py:train_epoch;torch.nn:Module._call_impl;torchvision:ResNet.forward;torch:conv2d 285
python;train.py:main;train.py:train_epoch;torch:Tensor.backward;torch.autograd:run_backward;cudaLaunchKernel 540
python;train.py:main;train.py:train_epoch;torch.optim:SGD.step;torch.optim:_multi_tensor_sgd 160
python;train.py:main;train.py:train_epoch;torch:Tensor.item;cudaStreamSynchronize 720
python;train.py:main;train.py:train_epoch;torch:Tensor.to;cudaMemcpyAsync 230
python;train.py:main;train.py:train_epoch;wandb:log 75