Most datacenters run near 40% utilization. This estimates the incremental revenue from selling that idle capacity as inference through HZ AI. Adjust the inputs to match your site.
Wholesale to the inference buyer. Market medians, July 2026: H100 ≈ $2.30–3.00 · H200 ≈ $3.95 · B300 ≈ $5.60–9.20 on-demand retail.
Sustained, batched serving. 70B-class on H100 ≈ 400–800 · H200 ≈ 600–1,100 · B300 ≈ 1,200–2,000. Small 8B models reach 2,500–5,000. Ballpark from published vLLM benchmarks — not yet measured on HZ AI.
Implied revenue at full load: per GPU-hour — compare against the GPU-hour sheet.
HZ AI platform fee is the remainder.
| Line | Per GPU / mo | Site / mo | Site / yr |
|---|---|---|---|
| Incremental GPU-hours sold | |||
| Tokens billed (input + output) | |||
| Gross inference revenue | |||
| HZ AI platform fee | |||
| Datacenter earnings |
Assumes 730 hours per month. Earnings require matched buyer demand and committed (reserved) volume makes these numbers contractual rather than best-case.
Token mode assumes the tokens/sec figure above is sustained across all sold hours. That number is a published-benchmark ballpark until HZ AI's own throughput benchmark lands and quote GPU-hour pricing where a guarantee is needed.
Prices are July 2026 market anchors and move with supply. Blackwell-class rates are still settling. Estimate only and not an offer.