VFX, inference, rendering, scientific jobs, streaming and migration - all on the same owned GPUs. If the fit is a retainer instead of rent, that is Services.
Most teams do not buy a ‘solution’. They have a job: serve a model, finish a sequence, train a checkpoint, move off a hyperscaler bill. The GPUs are the same. What changes is isolation, term, queue and whether engineering is in the room.
Inference wants a reserved or bare-metal baseline behind load balancing, with on-demand overflow when QPS jumps. Training wants a host that will not disappear mid-epoch - bare metal or a reserved block, InfiniBand if you scale past one node. Render wants the farm software on L40S / 4090, or a reserved block if the delivery date is real.
Batch and ETL should be on spot if they checkpoint, on-demand if they do not. Reserved is wasted money on a job that runs twice a week. Scientific codes that already scale on CUDA use the same H100 / A100 SKUs as training; the research programme is a separate conversation about hybrid scheduling, not a SKU.
Migration is inventory first: map instance types, storage and network to our card, run one job on on-demand, then reserve or go bare metal if the numbers work. We will not promise global edge latency from two sites. If your users are far from our sites, we measure before we quote, and sometimes the honest answer is keep inference where they are.
Engineering retainers exist for the cases where the hard part is the system, not the card: pipelines, pricing engines, inference control planes, farm-to-deadline wiring. Same commercial model as advertising - flat monthly, compute à la carte. Discovery is two weeks against real systems before anyone signs a retainer.