Workload Sizing
Describe the demand and per-user SLOs. The smallest cluster that meets them is sized live.
How to use this tab
This tab answers "how much hardware do I need to buy?". It is the reverse of the Modelling tab. There you pick the hardware and see how it runs. Here you say how busy your service will be and the promises you want to keep, and the app finds the smallest cluster of GPUs that can do it. Every box has an i button — tap it to learn what that box means.
- Pick your model (the AI brain) and a GPU to try (the chip that runs it). You can upload a model's config.json to add your own.
- Describe the demand: how many people use it at once (concurrency), and whether that number is the busiest moment or a calmer average. The app always plans for the busiest moment.
- Set the average length of questions (input) and answers (output), in tokens (word-pieces).
- Set your promises (SLOs): how long a user waits for the first word (TTFT), and how fast words then stream (throughput).
The app then shows the smallest cluster that both fits the model in memory and keeps every promise, plus how other GPUs would compare. The numbers are quick physics-based estimates, not a real benchmark.
Supported model types
- Language models — chatbots and text models (Llama, Qwen, DeepSeek, GLM, Nemotron). Sized by concurrent users and per-user speed promises (TTFT and throughput).
- Diffusion models — image generators (SDXL, FLUX.1, SD 3.5, PixArt-Σ, DiT-XL) and video (CogVideoX). Sized by target images or clips per second and a max time per image/clip.
- Vision-language (VLM) — LLMs that also read images (Qwen2.5-VL, Pixtral, LLaVA-OneVision). Sized like an LLM plus an "images per request" knob that adds visual-patch tokens.
- Vision-language-action (VLA) — robotics policies (pi-zero, SmolVLA). Sized by target robots × control frequency and a max per-chunk latency.
- Speech (ASR) — Whisper family (large-v3, large-v3-turbo, medium). Sized by concurrent real-time streams and minimum RTF.
- Embeddings & rerankers — text encoders (BGE-M3, Jina v3, E5-Mistral) and rerankers (BGE reranker). Sized by target docs (or query-doc pairs) per second and a max time per batch.
- World models (JEPA) — video models (V-JEPA, V-JEPA 2, and the action-conditioned V-JEPA 2-AC) that turn a clip into an embedding. Sized by target clips per second and a max time per clip.
- Autoregressive image (VAR) — image generators built like a language model (VAR-d30). Sized like a language model; set output tokens to the image's token count.
Pick the type from the grouped model menu. The demand and SLO controls change to match: users and tokens for language, images or clips per second for the others.
- Token — a word-piece. Models read and write tokens, not whole words. Roughly ¾ of a word each.
- Concurrency — how many requests are being answered at the same moment.
- P90 — the busy level that only the busiest 10% of moments go above. Higher than the average, lower than the all-time peak.
- TTFT — "time to first token": how long a user waits before the first word appears.
- Throughput — how many tokens per second the answer streams at.
- SLO — a service promise you commit to keeping (like "under 1.5 seconds to first word").
- TP / PP — two ways to split one model across several GPUs: TP splits each layer's math, PP puts different layers on different GPUs.
Default rates are the median on-demand price across GPU cloud providers (GetDeploying GPU price index, 2026-10-07). Providers vary 2-3x around it, so enter your own quote when you have one. Committed and spot apply a typical discount.
First-order: GPU draw = idle 30% of TDP + linear to TDP with utilization. Host overhead 30% of GPU. PUE 1.20. Grid intensities are directional estimates anchored to public 2024 grid-mix data.
Wall-clock ≈ weights ÷ min(source, node PCIe). Sharded assumes a well-parallelised loader (e.g. FSDP with parallel range reads). This is where Gen5 PCIe hosts (H100 and newer) pay off vs Gen4 (A100, L4, L40S): ~2× faster once the source can keep up.
| GPU | GPUs | Nodes | TP·PP·DP | $/hour | $/1M tok | Result |
|---|---|---|---|---|---|---|
| ★B300 (Blackwell Ultra) 288GB | 4 | 1 | 2·1·2 | $35.28 | $0.553 | meets SLOs |
| B200 (Blackwell) 192GB | 4 | 1 | 2·1·2 | $34.20 | $0.536 | meets SLOs |
| H200 SXM 141GB | 8 | 1 | 4·1·2 | $43.20 | $0.519 | meets SLOs |
| H100 SXM 80GB | 8 | 1 | 8·1·1 | $27.76 | $0.468 | meets SLOs |
| A100 SXM 80GB | 16 | 2 | 8·1·2 | $29.28 | $0.490 | meets SLOs |
| A100 SXM 40GB | 32 | 4 | 8·1·4 | $58.56 | $0.867 | meets SLOs |
| RTX PRO 6000 Blackwell 96GB | 32 | 4 | 8·1·4 | $68.80 | $1.05 | meets SLOs |
| RTX 4090 24GB | 88 | 11 | 8·1·11 | $39.60 | $0.693 | meets SLOs |
| RTX PRO 4500 Blackwell 32GB | 128 | 16 | 8·1·16 | — | — | meets SLOs |
| L40S 48GB | 176 | 22 | 8·1·22 | $276 | $4.69 | meets SLOs |
| L4 24GB | — | min throughput 30 tok/s/user exceeds the hardware ceiling (best 13 tok/s even at batch 1, full tensor-parallel) — a per-user decode limit, not a cluster-size cap | ||||
| A10G 24GB | — | min throughput 30 tok/s/user exceeds the hardware ceiling (best 26 tok/s even at batch 1, full tensor-parallel) — a per-user decode limit, not a cluster-size cap | ||||
| T4 16GB | — | min throughput 30 tok/s/user exceeds the hardware ceiling (best 14 tok/s even at batch 1, full tensor-parallel) — a per-user decode limit, not a cluster-size cap | ||||
First-order sizing: TTFT is a single request's prefill; per-user throughput is the decode rate at the sized batch. Cluster = TP·PP model-parallel groups replicated (DP) until the provisioned concurrency is served. Estimates, not a benchmark.