Skip to content
Flame Graph templates

Flame Graph — PyTorch Training Loop, Main Process (py-spy)

Where the Python main process of a ResNet-50 training run spends its time (4,760 samples): 45% blocked waiting for the DataLoader queue, so the GPU is starved by JPEG decoding in too few workers, and a further 15.1% stalled in a per-step loss.item() that forces a CUDA synchronise. Forward, backward and the optimiser together are only 33.5%. Illustrative values.

Template previewFlame Graph
ResNet-50 training — main process, py-spy 4,760 samples (illustrative)all (4760)pythontrain.py:maintrain.py:train_epochtorch.utils.data:DataLoader.__next__torch.utils.data:_MultiProcessingDataLoaderIter._get_dataqueue:Queue.gettorch.nn:Module._call_i…torchvision:ResNet.forw…cudaLaunchKerneltorch…torch:Tensor.itemcudaStreamSynchron…torch:Tensor.…torch.autogra…cudaLaunchKer…torc…cuda…

Make it your own.

title "ResNet-50 training — main process, py-spy 4,760 samples (illustrative)"
# py-spy record --native --rate 50 -d 60 --format raw -- python train.py --workers 4
python;train.py:main;train.py:train_epoch;torch.utils.data:DataLoader.__next__;torch.utils.data:_MultiProcessingDataLoaderIter._get_data;queue:Queue.get 2140
python;train.py:main;train.py:train_epoch;torch.nn:Module._call_impl;torchvision:ResNet.forward;cudaLaunchKernel 610
python;train.py:main;train.py:train_epoch;torch.nn:Module._call_impl;torchvision:ResNet.forward;torch:conv2d 285
python;train.py:main;train.py:train_epoch;torch:Tensor.backward;torch.autograd:run_backward;cudaLaunchKernel 540
python;train.py:main;train.py:train_epoch;torch.optim:SGD.step;torch.optim:_multi_tensor_sgd 160
python;train.py:main;train.py:train_epoch;torch:Tensor.item;cudaStreamSynchronize 720
python;train.py:main;train.py:train_epoch;torch:Tensor.to;cudaMemcpyAsync 230
python;train.py:main;train.py:train_epoch;wandb:log 75