llm inference

2 articles

AWS Inferentia2 vs NVIDIA L4: True Migration Cost

AWS Inferentia2 vs NVIDIA L4: True Migration Cost

A 2026 buyer's guide to Inf2 versus G6/L4, covering pinned-stack feasibility, successful-token cost, migration ledgers, and measured payback.

9/19/202620 min read
aws inferentia2 vs nvidia l4inferentia2 costnvidia l4
vLLM 0.29 Upgrade Guide: Runner V2 and API Checks

vLLM 0.29 Upgrade Guide: Runner V2 and API Checks

A 2026 production upgrade guide for vLLM 0.29.0 covering Model Runner V2 validation, immutable CUDA and ROCm artifacts, endpoint exposure, canary tests, and rollback evidence.

9/19/202621 min read
vllm 0.29 upgrade guidevllm migrationmodel runner v2