14InfiniBand vs RoCE for a GPU training cluster, how do you choose?▼mediumxAIMetaCoreWeave1 replies○ sign inBoth carry RDMA at 400Gb/s; the decision is about loss behavior, tuning burden, and who operates it. A defensible recommendation with the conditions that flip it, the format this question is scored on.Open full answer →
34Design the failure domains for a 16k-GPU training cluster. What's the blast radius of a single fault?▼hardMetaxAIAnthropic1 replies◆ premiumOne synchronous job, 16k GPUs, and any single component can stall all of them. The blast-radius analysis interviewers want: which faults take down a rack vs the run, where the shared single points of failure hide, and how parallelism layout maps onto failure domains.Open full answer →