Serve models on GPUs we own. Single-tenant if you need it, on-demand overflow when traffic spikes, load balancing included on reserved and bare metal. Latency is a number you can put in a contract, not a blog post.
Illustrative USD before tax. Spot is typical, not a guarantee. Reserved discounts are the published 18 / 24 / 28 / 31% schedule.
Most teams do not buy a ‘solution’. They have a job: serve a model, finish a sequence, train a checkpoint, move off a hyperscaler bill. The GPUs are the same. What changes is isolation, term, queue and whether engineering is in the room.
Inference wants a reserved or bare-metal baseline behind load balancing, with on-demand overflow when QPS jumps. Training wants a host that will not disappear mid-epoch - bare metal or a reserved block, InfiniBand if you scale past one node. Render wants the farm software on L40S / 4090, or a reserved block if the delivery date is real.
Batch and ETL should be on spot if they checkpoint, on-demand if they do not. Reserved is wasted money on a job that runs twice a week. Scientific codes that already scale on CUDA use the same H100 / A100 SKUs as training; the research programme is a separate conversation about hybrid scheduling, not a SKU.
Migration is inventory first: map instance types, storage and network to our card, run one job on on-demand, then reserve or go bare metal if the numbers work. We will not promise global edge latency from two sites. If your users are far from our sites, we measure before we quote, and sometimes the honest answer is keep inference where they are.
Engineering retainers exist for the cases where the hard part is the system, not the card: pipelines, pricing engines, inference control planes, farm-to-deadline wiring. Same commercial model as advertising - flat monthly, compute à la carte. Discovery is two weeks against real systems before anyone signs a retainer.
Switch the tab. The recommendation changes. If none of these is your job, say so on the contact form. We will tell you whether rentals, a retainer or a no is the honest answer.
Keep a reserved or bare-metal baseline behind load balancing. Overflow onto on-demand in the same site when traffic spikes. Same image, same network, higher unit price only on the extra seconds.
A small reserved or bare metal set behind load balancing. This is the capacity you quote to your own customers.
When QPS jumps, overflow onto on-demand in the same site. Same image, same VPC-style network, higher unit price only for the extra seconds.
Containerised endpoints can switch models without a new machine. We do not lock you to a hosted inference API.
Full table including reserved discount math lives on Pricing. Spot is typical. Fabric is InfiniBand at InfiniBand fabric, or 100 GbE (25 GbE on smaller cards) depending on the SKU.
| Instance | Site | VRAM | CPU / RAM | Fabric | On-demand | Spot typ. |
|---|---|---|---|---|---|---|
| H100 PCIe 80GB | InfiniBand | 80GB HBM3 | 24 / 200 GB | InfiniBand | $2.19 | $0.66 |
| A100 PCIe 80GB | InfiniBand | 80GB HBM2e | 16 / 180 GB | InfiniBand | $1.24 | $0.37 |
| L40S 48GB | 100 GbE | 48GB GDDR6 | 12 / 128 GB | 100 GbE | $0.86 | $0.26 |
| RTX 4090 24GB | 25 GbE | 24GB GDDR6X | 8 / 64 GB | 25 GbE | $0.41 | $0.12 |
| RTX A6000 48GB | Multi-site | 48GB GDDR6 | 12 / 96 GB | 25–100 GbE | $0.95 | $0.29 |
| CPU render node | Multi-site | 64 / 256 GB | 25 GbE | $0.68 | $0.20 |