Models & pricing
15 models, one balance
Prices are each lab's list price in USD per 1M tokens, verified on 2026-10-04. One credit buys $1 of compute at these prices; your tier's gateway fee is added on top.
Catalog
Routes are tried in order; if a provider errors or its circuit breaker is open, the request fails over to the next route automatically. Aliases (each lab's native model id) resolve to the same model and price.
Long-context pricing
Announced price changes
Gateway fee by tier
The gateway fee is a percentage of metered compute. It falls as your matured stake grows.
How a request is billed
- Before forwarding, the gateway reserves the most the request could cost: the prompt plus
max_tokens(or the model's default output allowance), at list price plus your fee. - If your balance can't cover that but can cover the prompt, the request still runs with a cap and stops cleanly with
insufficient_creditsif it reaches it. - When the response completes, you are charged for the exact tokens used — uncached input, cached input, cache writes and output each at their own rate — rounded up to the micro-credit. The unused reservation is released immediately.
- Some open-weight models are priced per serving provider, and some providers charge regional or peak-hour rates. When the upstream reports that a request cost more than list price, compute is charged at that reported cost instead, so every credit stays backed by real compute.
- Spend is drawn from developer and referral credits first, then staker credits, then purchased credits.
Example
1,000 input + 500 output tokens on a model priced $3 / $15 per 1M costs 0.003 + 0.0075 = 0.0105 credits of compute. With the 5.0% Standard fee the total is 0.011025 credits.