Below is the range we source. Each generation has a different job; picking the wrong one is not a matter of paying slightly more - it is the work not fitting, or the budget not buying what you needed.
NVIDIA H100
Workhorse for large-model training and high-throughput inference
Hopper generation with HBM3 memory, available with NVLink and in 8-GPU HGX systems. The most mature ecosystem of the current generations, with the most reproducible public material - usually where a team's first large training run starts.
NVIDIA H200
When memory, not maths, is the bottleneck
Same Hopper architecture, larger HBM3e memory. When the constraint is that the work does not fit - long-context inference, larger batches - it addresses the actual problem better than an H100 does.
NVIDIA B200
Next-generation training clusters
Blackwell generation, aimed at larger training and inference clusters. Typically delivered as HGX systems with multi-node interconnect, with materially higher rack power and cooling requirements than Hopper.
NVIDIA B300
The higher tier of the Blackwell family
Positioned for the highest compute density in the Blackwell line. Supply and lead times for this class fluctuate considerably - align your timing window early rather than planning as if it were shelf stock.
NVIDIA A100
Cost-sensitive training and inference
Ampere generation in 40GB and 80GB memory configurations. Strong price-performance with comparatively deep second-hand and rental supply, which suits budget-led projects at moderate model scale.
Domestic accelerators
Projects with sovereignty or data-residency constraints
For projects with explicit requirements on chip origin, data-centre jurisdiction or data not leaving a territory. Which brands and parts apply is decided by your compliance terms, so those need to be on the table first.
Exact availability, form factor (SXM / NVL / PCIe), volume and lead time depend on supply at the time of enquiry. This page makes no in-stock promise.