36Your training run isn't crashing, but step time doubled overnight. MFU dropped from 45% to 22%. Triage it.▼hardAnthropicMetaxAI3 replies◆ premiumNo error, no hang, the loss still moves, but the run is suddenly half as fast and burning the same dollars. The triage that separates a straggler from a fabric problem from broken comm/compute overlap, using the signals nvidia-smi can't give you.Open full answer →
45Plan a 7B pretrain from scratch on 64 H100s in 30 days, end to end.▼expertOpenAIMistralCohere1 replies◆ premium64 H100s for 30 days is a real budget with a real token count. The answer sizes the run from Chinchilla, picks parallelism for this exact cluster, and has a concrete plan for loss spikes and silent data corruption.Open full answer →