◢◤ GENAI_CALC_
SYS_TIME UTC+0

Self-host vs API

When does owning GPUs beat paying per token? Configure the exact cluster (same controls as Modelling), set your API price, find the crossover.

How to use this tab

This tab answers "should I run this model myself, or just call an API?". Self-hosting means renting GPUs and running the model yourself — you pay for the whole cluster around the clock, busy or not. An API charges per token — only for what you use, but each token costs a bit more.

  1. Configure the cluster with the full Modelling controls — model, GPU, GPU count, parallelism, quantization, and the request shape (input/output tokens). The forward model computes the cluster's throughput (the same engine as the Modelling tab).
  2. Set the duty cycle: what fraction of the month that cluster is actually busy. Flat-out all day is very different from spiking at lunch and idle overnight.
  3. Type the API price for the same model (input and output $/1M). Seeded from the model's real market price — replace with your quote.

The app shows which is cheaper today, the break-even duty cycle, and a chart of both. The API is a straight line up from zero; self-hosting is a staircase — one cluster serves only so much, so cost jumps a step as you add servers.

These are first-order estimates. Self-host cost here is GPU rental only. It leaves out engineering/on-call, storage, data transfer, failover headroom, and cold-start. A managed API bundles that in. Treat the crossover as a starting point, then check it against real quotes from your providers.
At 40% duty cycle
The API wins
Saves ~$18.0k/mo vs the other option.
Self-host
$20.3k
/mo · $28/hr · 8× H100 SXM 80GB
API
$2.3k
/mo at this volume
Break-even
never
API cheaper at any utilisation
Break-even duty cycle
—
api-always
Self-host $/1M
$3.56
blended, at 40%
API $/1M
$0.40
blended in+out
Per-user speed
Self-host streams faster
Your min-throughput SLO is 30 tok/s/user — both meet it.
Self-host / user
85 tok/s
at batch 32
API / user
80 tok/s
your entered limit
Cost vs volume
$0$27.4k$54.7k$82.1k$109.4k02.2M4.3M6.5M8.7M
Self-host (steps = +1 cluster) API (per token) your demand
x: requests / month · y: $ / month

This month, at 40% duty cycle

Requests695k
Input tokens2.8B
Output tokens2.8B
Peak rate0.7 req/s
Cluster throughput2.7k tok/s
Cluster8× H100 SXM 80GB · TP8·PP1

Self-host is billed 24/7 for the cluster you configured. API volume is derived from the cluster's peak throughput scaled by the duty cycle, so both sides serve the same traffic.

runs in your browser prices as of 2026-10-07 first-order estimates // not a benchmark