Pricing
Pay for the tokens you use.
Prepaid credits, billed by the token, with no subscription. Prices are in US dollars per 1M tokens and exclude VAT.
Model prices
A request is billed for the tokens it reads and the tokens it writes. Cached input costs less.
SLK 1slk-1
US dollars per 1M tokens
- Input
- $3.45
- Output
- $17.25
- Cache read
- $0.345
- Cache write
- $4.3125
What you send: messages, images, tool definitions.
What comes back, reasoning included.
Input that was served from the prompt cache.
Input that was written to the prompt cache.
For example
A request to slk-1 with 10,000 input tokens and 2,000 output tokens costs $0.0345 for the input and $0.0345 for the output: $0.069 in all.
Service tiers
Set service_tier on a request to trade price for priority. The multiplier applies to every price of the model.
| Tier | Price | Input / 1M | Output / 1M | What it is for |
|---|---|---|---|---|
default | 1× | $3.45 | $17.25 | Standard processing. The tier of a request that names none. |
flex | 0.5× | $1.725 | $8.625 | Best effort, at a lower price. For work that is not urgent. |
priority | 1.75× | $6.0375 | $30.1875 | Priority processing, at a higher price. |
The response says which tier a request was billed at. Service tiers in the docs.
Rate limits
Limits are counted per minute. Every account starts at the first tier and moves up on its own as its lifetime top-ups grow.
| Tier | Lifetime top-ups | Requests per minute | Tokens per minute |
|---|---|---|---|
| Tier 1 | None | 60 | 500,000 |
| Tier 2 | $50 | 300 | 2,000,000 |
| Tier 3 | $500 | 1,000 | 5,000,000 |
An API key can be given a lower request limit of its own.
How billing works
Credits first, usage after. There is no invoice at the end of the month, because there is nothing left to pay.
Prepaid credits
You top up your balance before you use it. The smallest top-up is $20 and the largest is $2,000. Every request is paid from the balance.
Prices exclude VAT
Every price on this page is in US dollars and before tax. Where VAT applies, it is added at the checkout.
A hold while a request runs
When a request starts, we place a temporary hold on your balance for the most the request could cost: its input, plus the longest output it is allowed to write. When the request finishes, the hold is released and you are charged for the tokens that were used.
When the balance is short
If you did not set
max_tokens, the output limit is lowered to what your balance, or the monthly limit of the key, can cover. If you setmax_tokensyourself and it does not fit, the request is refused with a 402 and costs nothing.
Get a key. Make the first call.
Sign up, top up, and send your first request to slk-1.