1 article
A 2026 planning model for per-rank LLM training memory, ZeRO and FSDP sharding, GPU count, CPU or NVMe offload, topology rounding, and measured validation.