GPU COMPUTE

GPU compute: H100 / H200 / B200 / B300, rented or delivered

Training, fine-tuning and inference do not want the same cards or the same commercial shape. We offer both on-demand rental and outright purchase with server delivery, across NVIDIA Hopper and Blackwell generations as well as domestic accelerators. Availability and lead times move with the market, so we quote against your actual workload and timing.

Quotes within 24h · from a single GPU · onshore and offshore sites

Accelerators we cover

Below is the range we source. Each generation has a different job; picking the wrong one is not a matter of paying slightly more - it is the work not fitting, or the budget not buying what you needed.

NVIDIA H100

Workhorse for large-model training and high-throughput inference

Hopper generation with HBM3 memory, available with NVLink and in 8-GPU HGX systems. The most mature ecosystem of the current generations, with the most reproducible public material - usually where a team's first large training run starts.

NVIDIA H200

When memory, not maths, is the bottleneck

Same Hopper architecture, larger HBM3e memory. When the constraint is that the work does not fit - long-context inference, larger batches - it addresses the actual problem better than an H100 does.

NVIDIA B200

Next-generation training clusters

Blackwell generation, aimed at larger training and inference clusters. Typically delivered as HGX systems with multi-node interconnect, with materially higher rack power and cooling requirements than Hopper.

NVIDIA B300

The higher tier of the Blackwell family

Positioned for the highest compute density in the Blackwell line. Supply and lead times for this class fluctuate considerably - align your timing window early rather than planning as if it were shelf stock.

NVIDIA A100

Cost-sensitive training and inference

Ampere generation in 40GB and 80GB memory configurations. Strong price-performance with comparatively deep second-hand and rental supply, which suits budget-led projects at moderate model scale.

Domestic accelerators

Projects with sovereignty or data-residency constraints

For projects with explicit requirements on chip origin, data-centre jurisdiction or data not leaving a territory. Which brands and parts apply is decided by your compliance terms, so those need to be on the table first.

Exact availability, form factor (SXM / NVL / PCIe), volume and lead time depend on supply at the time of enquiry. This page makes no in-stock promise.

Lead times and minimums

Most of the trade does not publish these. We do, because they decide whether you can plan around us. Windows below are the normal case; the binding figure is supply at the time of enquiry.

H100 systems
2-4 weeksNormal delivery window for Hopper-generation systems.
Blackwell (B200 / B300)
6-12 weeksTight supply and long allocation cycles - lock your window early.
Minimum rental
From one GPUNo need to buy a system first: validate on a single card, then talk scale.
Quote turnaround
Within 24 hoursSend the part, quantity, workload and timing; you get a plan and a quote inside one business day.

Where it runs, and the rules that apply

Compute can sit offshore or onshore. The two are governed differently and the available parts differ - which is why this has to be settled before choosing hardware, not after.

Offshore sites

Delivery of mainstream international accelerators (H100 / H200 / B200 / B300) is primarily via offshore sites. Cross-border use is confirmed case by case against export-control and data-transfer rules; the rules that apply to you govern the design.

Onshore sites

Onshore delivery focuses on parts that are lawfully available there, including domestic accelerators. Projects with hard requirements on chip origin, facility jurisdiction or data residency are better served this way.

We make no blanket representation about the availability of any particular part in any particular jurisdiction. Tell us your compliance constraints first - export control, cross-border data, sector regulation - and the design is worked back from them.

Typical shapes of demand

These are not customer case studies. They are the request patterns we see repeatedly, grouped by the timing and scale of the work, so you can locate yourself. Real configurations are settled at enquiry.

Starting from a single card

Budget just landed and the model size is not fixed. Rent one GPU, get the pipeline running, find out whether the bottleneck is compute or memory, then decide where to scale. The failure mode here is buying a system on day one and having the direction change.

A training run with a defined end

Three to eight weeks of training, then it is over. Renting fits: no capital outlay and scale follows the work. The thing to get right is locking the delivery window early, especially on Blackwell.

Inference that stays online

The model is settled and has to run as a service. Stability and unit cost dominate, so long-term rental or outright ownership both make sense. If the model needs to be exposed as an API, it can sit under the same account and wallet as this platform.

Memory-bound, not compute-bound

Long context or large batches, and the error is that it does not fit rather than that it is slow. Changing memory class beats adding cards - the larger HBM3e capacity on H200 is often the direct answer.

Residency or export constraints

Finance, public sector or sensitive data: settle the compliance boundary before choosing hardware. This usually means onshore sites and domestic accelerators, designed against the terms that bind you.

Two ways to take it: rent, or own

Rent compute on demand

Suits work with a defined start and end, or teams still validating: no capital outlay, and scale follows the workload. Billing granularity, available sites and minimum commitment are set per plan.

Buy servers outright

Suits teams that need to hold the hardware long term or have data-residency requirements: configuration, racking and tuning can be handled together. Lead time follows the part and current supply; you get a firm window at quotation.

Frequently asked

What is the lead time?

H100 systems normally 2-4 weeks; Blackwell (B200 / B300) 6-12 weeks - supply is tight there and allocation cycles are long, so lock your window early. These are the normal ranges; the binding figure is supply and delivery location at the time of enquiry.

What is the minimum order or rental?

Rental starts from a single GPU - you do not have to buy a system first: get the pipeline running, confirm the bottleneck, then talk scale. Purchase minimums vary by part and configuration and are given at quotation.

Should we rent or buy?

The deciding factor is the shape of the work over time, not the headline total. A defined start and end, or a scale still in flux, favours renting. Long-term ownership or hard data-residency requirements favour buying. We do both, and renting first then buying is a normal path.

What if we have data-residency or export constraints?

State the constraint first, then choose the hardware. Projects with jurisdiction or non-export requirements can go to domestic accelerators or a specified facility. The compliance terms decide the configuration, not the other way round.

Do you cover anything beyond the cards?

For purchases we can handle configuration, racking and tuning. For rentals we provide the environment and visible usage. We also run a model access platform, so if a trained model needs to be exposed as an API it can sit under the same account and wallet.

Talk to us about your workload

Send the part, the quantity, the type of job and your timing. We come back with a plan and a quote based on real availability. Phone and email reach us directly - no ticket queue.