gpuburn

tune the knobs on your imaginary inference business, watch the furnace tell you the truth

GPU
crunching the numbers...
running total since you opened this: $0.00
revenue / day
$0
compute cost / day
$0
profit / day
$0
margin
0%
gpus needed
0

real GPU rental + actual OpenRouter API pricing + published batched-decode throughput, Aug 2026 — H100 pricing, B200 pricing, DeepSeek-V3 wide-EP decode throughput on H100 (LMSYS/SGLang), DeepSeek's own $2/hr H800, 545% theoretical margin disclosure, DeepSeek R1 on OpenRouter, Qwen3-Coder-30B-A3B throughput on H100, Qwen3.8-27B on OpenRouter, Kimi K3 on OpenRouter, Kimi K3 cost-per-token math, 8×B300 at $7.39/hr, Kimi K3 self-hosting vs. API breakeven analysis, The Inference Ledger

the knobs

what you bill customers per million input (prompt) tokens — real APIs meter this separately, and it's always cheaper than output
what you bill customers per million output (generated) tokens
prompt/context tokens you bill for, in millions — usually the bulk of real traffic
generated tokens your API ships in a day, in millions — this is what drives GPU load below
output tokens/sec one GPU produces flat-out for this model (decode, not prefill)
how much of that throughput you actually capture — batching, bursty traffic, cold starts all eat this
what you pay per GPU-hour, on-demand or amortized reserved

brag or confess