No tiers, no minimums, no egress fees. Pay only for the tokens you consume — input and output priced identically.
| Provider | $ / Mtok in | $ / Mtok out | p50 ttft | Peak tok/s |
|---|---|---|---|---|
| Tachyon · HC1 | $0.05 | $0.05 | 4.2 ms | 17,000 |
| Groq · LPU | $0.05 | $0.08 | 18 ms | 4,800 |
| Cerebras | $0.10 | $0.10 | 42 ms | 1,800 |
| Together · H100 | $0.18 | $0.18 | 120 ms | 1,200 |
| OpenAI · gpt-4o-mini | $0.15 | $0.60 | 280 ms | 800 |
Note · figures measured 2026-05-12 from us-east. Public benchmarks.
Suppose your application sends 50M input tokens and receives 50M output tokens in a billing cycle. At the flat rate of $0.05 per million tokens, the total cost is $2.50. This includes all inference, no separate egress or API-call charges.
Input tokens : 50,000,000 Output tokens : 50,000,000 Total tokens : 100,000,000 Rate : $0.05 / 1,000,000 ────────────────────────────── Cost : 100 × $0.05 = $2.50