GPUSmith
Articles (English) - Page 1 / 5

CPU vs GPU LLM Inference: Low-Concurrency Cost Calculator
A 2026 guide to private CPU vs GPU LLM inference costs: matched-model tests, p95 latency objectives, whole-server energy, and a trace-driven calculator for reuse, purchases and fleet growth.

vLLM KV Cache Capacity Planning: Context and Concurrency
A 2026 engineering guide to vLLM KV cache capacity planning: calculate architecture-derived bytes per token, reconcile startup allocation, and canary context, concurrency, FP8 and preemption controls.

GPU Power Cap Optimization for Inference: Cost and SLOs
A 2026 protocol for choosing inference GPU power caps using same-run energy, latency-qualified throughput, replica rounding and monthly cost scenarios, with NVIDIA, NERSC and Dell evidence.

VMware Private AI Cloud in VCF 9.1.1: Now vs Next
A 2026 implementation guide to VCF 9.1.1 Private AI Cloud: what is generally available, which AI Gateway and sandbox features are future releases, and the hardware, license and pilot evidence buyers should request.

MIG vs MPS vs Time-Slicing for Kubernetes GPU Sharing
A 2026 decision guide for Kubernetes GPU sharing: compare isolation, memory behavior and scheduler resources, then use a workload worksheet, canary scorecard and rollback plan.

BlueField-4 DPU Qualification for Private AI Systems
A 2026 engineering guide to BlueField-4 DPU qualification for private AI, covering scale-in roles, STX differences, DOCA support, measured acceptance tests, and procurement evidence.

LLM Inference Metrics: TTFT, ITL, TPS and Goodput
A 2026 worksheet for comparing LLM inference benchmarks: define TTFT, ITL, per-user and aggregate TPS, normalize workload and timing boundaries, and reject unsupported goodput claims.

DCGM Exporter GPU Monitoring: Metrics, Alerts, Triage
A 2026 operator guide to DCGM exporter GPU monitoring: field types, alert routing, MIG labels, PromQL templates, cardinality controls, and an incident triage checklist.

MIG vs MPS vs Time-Slicing: GPU Cost and Isolation
A 2026 engineering comparison of exclusive GPUs, MIG, MPS and time-slicing for inference, with isolation guarantees, model-fit profiles, SLO tests and cost per qualified tenant.

NVIDIA AI Enterprise 8.2 vs 7.8 LTSB: Upgrade Decision
A 2026 production decision guide comparing R595 and R580 drivers, version-pinned support boundaries, planned April 2027 and July 2028 end-of-life dates, and an upgrade evidence gate.

Confidential GPU Attestation: H100 to B300 Matrix
2026 comparison of H100, H200, B200, B300 and RTX PRO confidential GPU support, CPU TEE dependencies, attestation evidence, verifier ownership and procurement acceptance tests.

Triton 26.08 Upgrade Runbook: CUDA 13.4, Canary and Rollback Testing
A 2026 upgrade runbook for Triton Inference Server 26.08 covering CUDA 13.4.1 compatibility, container image selection, backend validation, canary testing, and rollback criteria.

Cold-Plate GPU Rack Commissioning: Field Checklist
A 2026 commissioning checklist for cold-plate GPU racks: reconcile CDU duty points, flush and sample the TCS loop, test alarms, quantify heat balance, and assemble turnover evidence.

How to Qualify GPUDirect Storage with GDSIO: Runbook
A 2026 qualification runbook for GPUDirect Storage: check filesystem and topology support, verify gdscheck and gdsio results, detect CPU fallback, and record reproducible local baselines.

Disaggregated Prefill/Decode Private Fleet Break-Even
A 2026 break-even model for private disaggregated prefill/decode fleets, covering SLO goodput, KV transfer, idle pool-hours, HBM opportunity cost, and conditional routing.

HP ZGX Fury RHEL Certification and Procurement
A 2026 procurement analysis of HP ZGX Fury RHEL certification, its 748 GB coherent memory design, planned Red Hat AI Factory stack, support boundaries, and acceptance tests.

MIG vs MPS vs Time-Slicing vs Passthrough Compared
A 2026 constraint-first comparison of NVIDIA MIG, CUDA MPS, GPU time-slicing, passthrough, and dedicated GPUs across isolation, latency, scheduling, and observability.

CUDA Driver Toolkit Container Compatibility Gate
A 2026 production guide to CUDA driver, toolkit, and container compatibility, with dated driver ranges, framework checks, canary commands, and rollback criteria.

Intel E835 Private AI Clusters: 200GbE RoCE Guide
A 2026 qualification guide to Intel E835 private AI clusters, covering 200GbE port modes, RoCEv2 and iWARP, GPU-direct evidence, SKUs, and acceptance testing.

nvidia-smi topo Rank Binding: GPU, CPU and NIC Affinity
A 2026 operator runbook for reading nvidia-smi topology, mapping ranks to GPU UUIDs, effective CPU and NUMA sets, selecting NICs, and validating placement with NCCL tests.

AMD MI355X Production Clusters: Buyer Evidence
A 2026 evidence audit of AMD MI355X production clusters, covering HUMAIN's live deployment, 1,400 W power, 800G networking, acceptance gates, availability, and B200 context.