Show the math
The calculation chain behind the current config — every formula symbolically, then with your numbers, then the result. Tied to the shared config, so it matches Modelling / Workload / Training exactly.
How to use this tab
Pick the view. Inference is the Modelling / Workload roofline (memory fit, step time, throughput, TTFT). Training is the fine-tune memory + step-time story. Change the config below (or on any other tab) and the math updates live. Results are sourced from the same functions the tool computes with — this tab can't drift from them.
Setup
Model parameters
70.6B
GPU peak (compute · bandwidth)
99 TFLOP/s (fp16), 3.35 TB/s
Parallelism
TP 8 · PP 1 · EP 1 · DP 1 (of 8 GPUs)
Sequence
8192 tokens
Memory (per GPU)
What one GPU must hold for this replica.
Weights
b_w = 2 bytes/param (BF16 / FP16)
17.6 GB
KV per token · layer
GQA: 8 KV heads
4096 B
KV cache
paged allocation, KV on 80/80 layers, sharded 8×
11 GB
Activations
1.07 GB
Overhead
CUDA context + framework + fragmentation
1.61 GB
Total vs capacity
31.3 GB / 85.9 GB — fits
Throughput (roofline)
A step is bounded by the slower of memory traffic and compute, plus in-stage collectives.
Bytes read per step (decode: weights + KV)
28.6 GB
Memory time
MBU = 0.8
10.7 ms
FLOPs per step
0.565 TFLOP (32 tokens/step)
Compute time
MFU = 0.74, per-GPU factor φ = 1
0.771 ms
Step time (roofline + comm)
bottleneck: memory
11.8 ms
Aggregate throughput
2706 tok/s
Per-user throughput
84.6 tok/s per user
Latency (TTFT)
Prefill FLOPs (uncached input)
2313 TFLOP
Time to first token
3159 ms