Did bf16 train cleanly per arm? The 12B full-SFT arm was chosen deliberately because Gemma-class models are reported to hit bf16/NaN on gfx1151.
Each measured run carries a NaN/Inf counter and first/last window loss. The probe is a short steady-state throughput window, not a full convergence run — a longer run could still surface instability — but within the measured window, the reported gfx1151 Gemma bf16 issue did not reproduce on this stack (axolotl image, torch 2.8 / rocm 7.12).