GPUSmith

Articles (English) - Page 1 / 5

CPU vs GPU LLM Inference: Low-Concurrency Cost Calculator

CPU vs GPU LLM Inference: Low-Concurrency Cost Calculator

A 2026 guide to private CPU vs GPU LLM inference costs: matched-model tests, p95 latency objectives, whole-server energy, and a trace-driven calculator for reuse, purchases and fleet growth.

10/8/2026• 26 min read
cpu vs gpu llm inferencellm inference costcpu inference
vLLM KV Cache Capacity Planning: Context and Concurrency

vLLM KV Cache Capacity Planning: Context and Concurrency

A 2026 engineering guide to vLLM KV cache capacity planning: calculate architecture-derived bytes per token, reconcile startup allocation, and canary context, concurrency, FP8 and preemption controls.

10/5/2026• 25 min read
vllm kv cache capacity planningvllm kv cache memory calculationvllm concurrency
GPU Power Cap Optimization for Inference: Cost and SLOs

GPU Power Cap Optimization for Inference: Cost and SLOs

A 2026 protocol for choosing inference GPU power caps using same-run energy, latency-qualified throughput, replica rounding and monthly cost scenarios, with NVIDIA, NERSC and Dell evidence.

10/4/2026• 27 min read
gpu power cap optimization for inferencegpu power cappinginference latency
VMware Private AI Cloud in VCF 9.1.1: Now vs Next

VMware Private AI Cloud in VCF 9.1.1: Now vs Next

A 2026 implementation guide to VCF 9.1.1 Private AI Cloud: what is generally available, which AI Gateway and sandbox features are future releases, and the hardware, license and pilot evidence buyers should request.

10/3/2026• 27 min read
vmware private ai cloudvcf 9.1.1vmware ai factory
MIG vs MPS vs Time-Slicing for Kubernetes GPU Sharing

MIG vs MPS vs Time-Slicing for Kubernetes GPU Sharing

A 2026 decision guide for Kubernetes GPU sharing: compare isolation, memory behavior and scheduler resources, then use a workload worksheet, canary scorecard and rollback plan.

10/1/2026• 26 min read
kubernetes gpu sharingmig vs mpsgpu time-slicing
BlueField-4 DPU Qualification for Private AI Systems

BlueField-4 DPU Qualification for Private AI Systems

A 2026 engineering guide to BlueField-4 DPU qualification for private AI, covering scale-in roles, STX differences, DOCA support, measured acceptance tests, and procurement evidence.

9/29/2026• 22 min read
bluefield-4 dpuprivate aiscale-in networking
LLM Inference Metrics: TTFT, ITL, TPS and Goodput

LLM Inference Metrics: TTFT, ITL, TPS and Goodput

A 2026 worksheet for comparing LLM inference benchmarks: define TTFT, ITL, per-user and aggregate TPS, normalize workload and timing boundaries, and reject unsupported goodput claims.

9/28/2026• 24 min read
llm inference metricsttftinter-token latency
DCGM Exporter GPU Monitoring: Metrics, Alerts, Triage

DCGM Exporter GPU Monitoring: Metrics, Alerts, Triage

A 2026 operator guide to DCGM exporter GPU monitoring: field types, alert routing, MIG labels, PromQL templates, cardinality controls, and an incident triage checklist.

9/27/2026• 25 min read
dcgm exportergpu monitoringprometheus metrics
MIG vs MPS vs Time-Slicing: GPU Cost and Isolation

MIG vs MPS vs Time-Slicing: GPU Cost and Isolation

A 2026 engineering comparison of exclusive GPUs, MIG, MPS and time-slicing for inference, with isolation guarantees, model-fit profiles, SLO tests and cost per qualified tenant.

9/26/2026• 25 min read
mig vs mps vs time slicinggpu sharinggpu inference
NVIDIA AI Enterprise 8.2 vs 7.8 LTSB: Upgrade Decision

NVIDIA AI Enterprise 8.2 vs 7.8 LTSB: Upgrade Decision

A 2026 production decision guide comparing R595 and R580 drivers, version-pinned support boundaries, planned April 2027 and July 2028 end-of-life dates, and an upgrade evidence gate.

9/25/2026• 22 min read
nvidia ai enterprise 8.2 vs 7.8 ltsbnvidia ai enterpriseinfrastructure 8.2
Confidential GPU Attestation: H100 to B300 Matrix

Confidential GPU Attestation: H100 to B300 Matrix

2026 comparison of H100, H200, B200, B300 and RTX PRO confidential GPU support, CPU TEE dependencies, attestation evidence, verifier ownership and procurement acceptance tests.

9/24/2026• 28 min read
confidential gpu attestationh100 attestationh200 attestation
Triton 26.08 Upgrade Runbook: CUDA 13.4, Canary and Rollback Testing

Triton 26.08 Upgrade Runbook: CUDA 13.4, Canary and Rollback Testing

A 2026 upgrade runbook for Triton Inference Server 26.08 covering CUDA 13.4.1 compatibility, container image selection, backend validation, canary testing, and rollback criteria.

9/24/2026• 24 min read
triton 26.08 upgrade guidetriton upgrade runbookinference server upgrades
Cold-Plate GPU Rack Commissioning: Field Checklist

Cold-Plate GPU Rack Commissioning: Field Checklist

A 2026 commissioning checklist for cold-plate GPU racks: reconcile CDU duty points, flush and sample the TCS loop, test alarms, quantify heat balance, and assemble turnover evidence.

9/24/2026• 22 min read
cold plate gpu rack commissioning checklistcdu commissioninggpu liquid cooling
How to Qualify GPUDirect Storage with GDSIO: Runbook

How to Qualify GPUDirect Storage with GDSIO: Runbook

A 2026 qualification runbook for GPUDirect Storage: check filesystem and topology support, verify gdscheck and gdsio results, detect CPU fallback, and record reproducible local baselines.

9/23/2026• 26 min read
gpudirect storagegdsiogdscheck
Disaggregated Prefill/Decode Private Fleet Break-Even

Disaggregated Prefill/Decode Private Fleet Break-Even

A 2026 break-even model for private disaggregated prefill/decode fleets, covering SLO goodput, KV transfer, idle pool-hours, HBM opportunity cost, and conditional routing.

9/22/2026• 22 min read
disaggregated prefill decodeprefill decode disaggregationllm inference
HP ZGX Fury RHEL Certification and Procurement

HP ZGX Fury RHEL Certification and Procurement

A 2026 procurement analysis of HP ZGX Fury RHEL certification, its 748 GB coherent memory design, planned Red Hat AI Factory stack, support boundaries, and acceptance tests.

9/21/2026• 22 min read
hp zgx fury rhel certificationhp zgx furynvidia gb300
MIG vs MPS vs Time-Slicing vs Passthrough Compared

MIG vs MPS vs Time-Slicing vs Passthrough Compared

A 2026 constraint-first comparison of NVIDIA MIG, CUDA MPS, GPU time-slicing, passthrough, and dedicated GPUs across isolation, latency, scheduling, and observability.

9/20/2026• 23 min read
mig vs mpsgpu time-slicinggpu passthrough
CUDA Driver Toolkit Container Compatibility Gate

CUDA Driver Toolkit Container Compatibility Gate

A 2026 production guide to CUDA driver, toolkit, and container compatibility, with dated driver ranges, framework checks, canary commands, and rollback criteria.

9/20/2026• 22 min read
cuda compatibilitycuda drivercuda toolkit
Intel E835 Private AI Clusters: 200GbE RoCE Guide

Intel E835 Private AI Clusters: 200GbE RoCE Guide

A 2026 qualification guide to Intel E835 private AI clusters, covering 200GbE port modes, RoCEv2 and iWARP, GPU-direct evidence, SKUs, and acceptance testing.

9/20/2026• 20 min read
intel e835 private ai clustersintel e835200gbe roce
nvidia-smi topo Rank Binding: GPU, CPU and NIC Affinity

nvidia-smi topo Rank Binding: GPU, CPU and NIC Affinity

A 2026 operator runbook for reading nvidia-smi topology, mapping ranks to GPU UUIDs, effective CPU and NUMA sets, selecting NICs, and validating placement with NCCL tests.

9/20/2026• 20 min read
nvidia-smi topo rank bindinggpu cpu affinitygpu nic affinity
AMD MI355X Production Clusters: Buyer Evidence

AMD MI355X Production Clusters: Buyer Evidence

A 2026 evidence audit of AMD MI355X production clusters, covering HUMAIN's live deployment, 1,400 W power, 800G networking, acceptance gates, availability, and B200 context.

9/19/2026• 21 min read
amd mi355x production clustersmi355x deploymentmi355x cluster