Sundancæ Research Inc.
Pricing
Solutions

Real-time streaming

Encode, transcode and light inference on live video without paying hyperscaler egress for every viewer. L40S and 4090s are the usual cards; load balancing sits in front if you run more than one worker.

Talk streamingLoad balancing
Fleet availability
Updated continuously from the scheduler
Live
Instance
Utilisation
Free now
H100 PCIe 80GB
InfiniBand
92%
6
A100 80GB
InfiniBand
78%
19
L40S 48GB
100 GbE
61%
24
RTX 4090 24GB
25 GbE
44%
41
CPU render node
Multi-site
55%
33
312 GPUs · both sites
Full price list
Cost estimator

Pick a SKU, a mode and a run length

Illustrative USD before tax. Spot is typical, not a guarantee. Reserved discounts are the published 18 / 24 / 28 / 31% schedule.

Instance
How you pay
Estimate · $2.19/GPU-hour
$420
8× H100 PCIe 80GB for 24 hours on on-demand. Illustrative, USD, before tax. Spot is typical, not a guarantee.
Per hour
$17.52
Per day, 24h
$420
30-day continuous
$12614
This run
$420
Reserved discounts are off the public on-demand rate (18 / 24 / 28 / 31%). In-region transfer is included. NVMe and object storage are extra.
In practice

How this actually runs

Most teams do not buy a ‘solution’. They have a job: serve a model, finish a sequence, train a checkpoint, move off a hyperscaler bill. The GPUs are the same. What changes is isolation, term, queue and whether engineering is in the room.

Inference wants a reserved or bare-metal baseline behind load balancing, with on-demand overflow when QPS jumps. Training wants a host that will not disappear mid-epoch - bare metal or a reserved block, InfiniBand if you scale past one node. Render wants the farm software on L40S / 4090, or a reserved block if the delivery date is real.

Batch and ETL should be on spot if they checkpoint, on-demand if they do not. Reserved is wasted money on a job that runs twice a week. Scientific codes that already scale on CUDA use the same H100 / A100 SKUs as training; the research programme is a separate conversation about hybrid scheduling, not a SKU.

Migration is inventory first: map instance types, storage and network to our card, run one job on on-demand, then reserve or go bare metal if the numbers work. We will not promise global edge latency from two sites. If your users are far from our sites, we measure before we quote, and sometimes the honest answer is keep inference where they are.

Engineering retainers exist for the cases where the hard part is the system, not the card: pipelines, pricing engines, inference control planes, farm-to-deadline wiring. Same commercial model as advertising - flat monthly, compute à la carte. Discovery is two weeks against real systems before anyone signs a retainer.

Workload

Same GPUs, four different jobs

Switch the tab. The recommendation changes. If none of these is your job, say so on the contact form. We will tell you whether rentals, a retainer or a no is the honest answer.

Always-on replicas, burst when QPS jumps

Keep a reserved or bare-metal baseline behind load balancing. Overflow onto on-demand in the same site when traffic spikes. Same image, same network, higher unit price only on the extra seconds.

  • H100 / A100 / L40S are the usual SKUs
  • vLLM, TensorRT, Triton or your container
  • Cross-site failover is opt-in
  • Data and weights stay in-country
Live

GPU time on a clock, not a contract

Encode / transcode

Hardware encode on L40S or 4090. Scale workers with the number of concurrent streams, not with a reserved instance year.

Live inference

Captions, moderation or vision models on the same box or a sidecar GPU. Overflow to on-demand for events.

Egress

In-region transfer is included. We will tell you honestly if your viewer map does not match either site.

Signal

Illustrative week · mix of work on one account

Inference
Train / render
62%47%31%16%0%
Mon
Tue
Wed
Thu
Fri
Sat
Sun
How to read it
Same GPUs, different jobs. Inference holds the baseline; render and training burst. Hover a point for the exact value. These series are illustrative of how the product behaves, not a live feed from your account.
Rate card excerpt

Published SKUs, both sites

Full table including reserved discount math lives on Pricing. Spot is typical. Fabric is InfiniBand at InfiniBand fabric, or 100 GbE (25 GbE on smaller cards) depending on the SKU.

InstanceSiteVRAMCPU / RAMFabricOn-demandSpot typ.
H100 PCIe 80GBInfiniBand80GB HBM324 / 200 GBInfiniBand$2.19$0.66
A100 PCIe 80GBInfiniBand80GB HBM2e16 / 180 GBInfiniBand$1.24$0.37
L40S 48GB100 GbE48GB GDDR612 / 128 GB100 GbE$0.86$0.26
RTX 4090 24GB25 GbE24GB GDDR6X8 / 64 GB25 GbE$0.41$0.12
RTX A6000 48GBMulti-site48GB GDDR612 / 96 GB25–100 GbE$0.95$0.29
CPU render nodeMulti-site64 / 256 GB25 GbE$0.68$0.20
Reserved baseline for the always-on encoder set
On-demand overflow for premieres and peaks
Bare metal if you need a locked-down appliance image
Engineering retainer for the control plane if you do not want to run it
Questions

Straight answers

Both. Rentals are ordinary nodes. The farm is optional software on the same hardware. Inference usually wants load balancing, not the farm.
Related
Talk to us about capacity
Tell us the workload. We’ll route it to rentals, engineering or the render farm.
Talk streamingLoad balancing