Docs
The Milkxi API
An API in the OpenAI Chat Completions format for SLK 1. This page covers all of it.
01
Quickstart
If you have code that talks to OpenAI, it can talk to Milkxi: change the base URL, the API key and the model. The OpenAI SDKs work as they are.
- Create an account and add credits.
- Create an API key in the dashboard. The whole key is shown once, so copy it then, and keep it in an environment variable.
- Send a request to the base URL below.
- Base URL
https://api.milkxi.com/v1
curl https://api.milkxi.com/v1/chat/completions \
-H "Authorization: Bearer $MILKXI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "slk-1",
"messages": [{"role": "user", "content": "Why is the sky blue?"}]
}'import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.milkxi.com/v1",
api_key=os.environ["MILKXI_API_KEY"],
)
response = client.chat.completions.create(
model="slk-1",
messages=[{"role": "user", "content": "Why is the sky blue?"}],
)
print(response.choices[0].message.content)import OpenAI from 'openai';
const client = new OpenAI({
baseURL: 'https://api.milkxi.com/v1',
apiKey: process.env.MILKXI_API_KEY,
});
const response = await client.chat.completions.create({
model: 'slk-1',
messages: [{ role: 'user', content: 'Why is the sky blue?' }],
});
console.log(response.choices[0].message.content);The response has the shape of an OpenAI response, with the model’s reasoning next to its answer:
{
"id": "chatcmpl-...",
"object": "chat.completion",
"model": "slk-1",
"choices": [
{
"index": 0,
"finish_reason": "stop",
"message": {
"role": "assistant",
"content": "Sunlight is scattered by the air, and blue light is scattered the most...",
"reasoning_content": "The question is about Rayleigh scattering. I should explain..."
}
}
],
"usage": {
"prompt_tokens": 68,
"completion_tokens": 412,
"total_tokens": 480,
"prompt_tokens_details": {
"cached_tokens": 0
},
"completion_tokens_details": {
"reasoning_tokens": 290
}
},
"service_tier": "default"
}An example. The text and the token counts are there to show the shape.
The Python examples further down use the client made in the Python tab above, and the TypeScript one the client of the TypeScript tab. The curl examples stand on their own.
02
Authentication
Every request carries an API key in the Authorization header, as a bearer token. Keys start with sk-slk-.
Authorization: Bearer sk-slk-...You create and revoke keys in the dashboard. A key belongs to a workspace, and what it spends comes out of the credits of that workspace. Keep it on your server: whoever has the key can spend those credits.
What a key may do
A key can be allowed everything, or only the endpoints you choose:
| Permission | Endpoints | What it allows |
|---|---|---|
models.read | GET /v1/modelsGET /v1/models/{id} | List the models that are available and read the details of one. |
chat.completions | POST /v1/chat/completions | Send messages to a model and receive its replies, streamed or not. |
A key can also be limited to certain models, and given a request limit, a monthly spend limit, an expiry date and an interaction tone of its own.
A request with no key, or with a key that is unknown, revoked or expired, gets a 401. A key that lacks the permission for an endpoint gets a 403.
03
Models
List the models your key can use, or look one up by its id.
GET/v1/models
GET/v1/models/{id}
curl https://api.milkxi.com/v1/models \
-H "Authorization: Bearer $MILKXI_API_KEY"{
"object": "list",
"data": [
{
"id": "slk-1",
"object": "model",
"created": 1790640000,
"owned_by": "milkxi"
}
]
}| Model | Context window | Output, at most | Output, by default | Input |
|---|---|---|---|---|
SLK 1slk-1slk-latest | 1,048,576 | 131,072 | 32,768 | Text and images |
Limits are in tokens. Output is text.
slk-latest is an alias: today it points to slk-1. A response names the model by its id, whichever name the request used.
04
Chat completions
Send a conversation and get the next message.
POST/v1/chat/completions
| Parameter | Support | Notes |
|---|---|---|
model | Required | slk-1, or its alias slk-latest. |
messages | Required | Text, and images as image_url content parts. |
max_tokens, max_completion_tokens | Supported | 32,768 when not set. At most 131,072. |
stream, stream_options | Supported | Server-sent events. include_usage adds a last chunk with the usage. |
tools, tool_choice | Supported | Function tools, with parallel tool calls. |
response_format | Supported | json_schema (strict) and json_object. |
reasoning_effort | Supported | low, high or max. max when not set. |
service_tier | Supported | default, flex or priority. auto is the same as no tier. |
temperature | Supported | At most 2. |
interaction_tone | Supported | Ours, not OpenAI’s. See Interaction tone. |
n | Limited | Must be 1. |
presence_penalty, frequency_penalty | Limited | Must be 0. |
logprobs | Ignored | Accepted and ignored: no log probabilities are returned. |
| Anything else | Ignored | A parameter the API does not know is ignored, not refused. |
Output length
max_tokens and max_completion_tokens mean the same thing: the most tokens one request may write. Without either, the limit is 32,768. The most you can ask for is 131,072, and asking for more is a 400.
Images
Send an image as an image_url content part, next to the text that goes with it. The model reads text and images, and answers in text.
import base64
with open("chart.png", "rb") as f:
image = base64.b64encode(f.read()).decode()
response = client.chat.completions.create(
model="slk-1",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "What does this chart show?"},
{"type": "image_url", "image_url": {"url": f"data:image/png;base64,{image}"}},
],
}],
)05
Streaming
Set stream to true and the reply arrives as server-sent events, in the chunk format of OpenAI. The reasoning comes in delta.reasoning_content and the answer in delta.content.
Set stream_options.include_usage to true and one more chunk follows the end of the answer, with the usage of the whole request.
stream = client.chat.completions.create(
model="slk-1",
messages=[{"role": "user", "content": "Write a haiku about maps."}],
stream=True,
stream_options={"include_usage": True},
)
for chunk in stream:
if chunk.usage: # the last chunk: no choices, only the usage
print("\n", chunk.usage)
elif chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="")curl -N https://api.milkxi.com/v1/chat/completions \
-H "Authorization: Bearer $MILKXI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "slk-1",
"stream": true,
"stream_options": {"include_usage": true},
"messages": [{"role": "user", "content": "Write a haiku about maps."}]
}'06
Tool calling
tools and tool_choice work as they do with OpenAI, and the model can ask for several calls in one turn. The id of a tool call looks like get_weather_0.
When you send the results back, send the assistant message exactly as you received it, with its reasoning_content and its tool_calls. The model needs both to go on.
import json
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.milkxi.com/v1",
api_key=os.environ["MILKXI_API_KEY"],
)
def get_weather(city: str) -> dict:
# A stand-in for your own code. Here every city has the same weather.
return {"city": city, "temperature_c": 18, "conditions": "cloudy"}
tools = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a city.",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}]
messages = [{"role": "user", "content": "What is the weather in Oslo and in Lima?"}]
first = client.chat.completions.create(model="slk-1", messages=messages, tools=tools)
assistant = first.choices[0].message
# Send the assistant message back as it came: reasoning_content and tool_calls included.
messages.append(assistant.model_dump(exclude_none=True))
for call in assistant.tool_calls or []: # the model can ask for several calls at once
args = json.loads(call.function.arguments)
messages.append({
"role": "tool",
"tool_call_id": call.id, # for example "get_weather_0"
"content": json.dumps(get_weather(**args)),
})
final = client.chat.completions.create(model="slk-1", messages=messages, tools=tools)
print(final.choices[0].message.content)This one is complete in itself: it makes its own client, and get_weather stands in for your own code.
07
Structured outputs
response_format takes two forms. With json_schema and strict: true, the answer follows your JSON Schema. With json_object, the answer is valid JSON, with no schema.
import json
response = client.chat.completions.create(
model="slk-1",
messages=[{"role": "user", "content": "I flew from Oslo to Lima on 3 May. Extract the trip."}],
response_format={
"type": "json_schema",
"json_schema": {
"name": "trip",
"strict": True,
"schema": {
"type": "object",
"properties": {
"origin": {"type": "string"},
"destination": {"type": "string"},
"date": {"type": "string"},
},
"required": ["origin", "destination", "date"],
"additionalProperties": False,
},
},
},
)
trip = json.loads(response.choices[0].message.content)The first request with a new schema can take longer than the ones after it.
Interaction tone is skipped for both forms, so nothing is added to your JSON.
08
Reasoning
SLK 1 reasons before it answers, on every request: reasoning is always on. reasoning_effort says how much, from low to max. Without it, the effort is max.
| Field | What it holds |
|---|---|
message.reasoning_content | The reasoning, as text. |
delta.reasoning_content | The same, piece by piece, when streaming. |
usage.completion_tokens_details.reasoning_tokens | How many tokens the reasoning took. They are part of completion_tokens and are billed as output. |
response = client.chat.completions.create(
model="slk-1",
messages=[{"role": "user", "content": "Is 2,147,483,647 a prime number?"}],
reasoning_effort="low",
)
message = response.choices[0].message
print(message.reasoning_content) # how the model got there
print(message.content) # the answer
print(response.usage.completion_tokens_details.reasoning_tokens)The OpenAI SDKs have no type for reasoning_content. In Python it is an attribute of the message all the same; in TypeScript, cast the message to read it.
In a conversation that uses tools, send each assistant message back with its reasoning_content: see Tool calling.
09
Interaction tone
Interaction tone changes how the model answers without changing your prompt. One tone applies at a time.
| Tone | What the answer looks like |
|---|---|
default | Unchanged: the model's normal answer. Nothing is added to the request. |
detailed | Thorough. Explains the reasoning, covers edge cases and includes examples where they help. |
concise | The shortest complete answer. No preamble, no repetition. |
summary | The normal answer, then a short Summary section at the end. |
bullets | The normal answer, then a Key points section at the end with 1–3 bullets: what was done, plus any action points. |
Set it on a key
Every API key has a tone, which you choose in the dashboard. A new key starts on default. A change reaches the API within about 60 seconds.
Override it for one request
Send interaction_tone in the body of the request. With the OpenAI SDK for Python, a field the SDK does not know goes in extra_body:
response = client.chat.completions.create(
model="slk-1",
messages=[{"role": "user", "content": "Plan my database migration"}],
extra_body={"interaction_tone": "bullets"}, # optional: without it, the key's tone applies
)curl https://api.milkxi.com/v1/chat/completions \
-H "Authorization: Bearer $MILKXI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "slk-1",
"interaction_tone": "concise",
"messages": [{"role": "user", "content": "Plan my database migration"}]
}'For a tool that lets you add a header but not a field in the body, send the X-Interaction-Tone header instead:
const response = await client.chat.completions.create(
{
model: 'slk-1',
messages: [{ role: 'user', content: 'Plan my database migration' }],
},
{ headers: { 'X-Interaction-Tone': 'summary' } },
);Which one applies
interaction_tonein the body, if the request has it.- Otherwise the
X-Interaction-Toneheader, if the request has it. - Otherwise the tone of the API key.
A value that is not one of the five tones is a 400. So is a blank value in the body. A blank header is ignored.
{
"error": {
"message": "Invalid 'interaction_tone': must be one of default, detailed, concise, summary, bullets",
"type": "invalid_request_error",
"param": "interaction_tone",
"code": "invalid_value"
}
}When it is skipped
A tone of default adds nothing. Any other tone is left out, and the request goes through as you sent it, when the answer can only be JSON or a tool call:
response_formatasks for JSON:{"type": "json_object"}, JSON{"type": "json_schema", …}, JSON that follows a schema
tool_choiceforces a tool call:"required", a call to some tool{"type": "function", …}, a call to a named function{"type": "custom", …}, a call to a custom tool{"type": "allowed_tools", "allowed_tools": {"mode": "required", …}}, a call to one of the allowed tools
With tools and no forced call, the tone still applies. A summary or key points are added to the final answer only, never inside a tool call.
What it costs
A tone other than default adds a short instruction to your request. Its tokens are billed as normal input, and what the tone adds to the answer is billed as normal output.
If your prompt asks for something else
When your own prompt asks for a different format, the tone usually wins. That is intended: setting a tone is a choice. Use default to keep the answer entirely under your own prompt.
10
Service tiers
service_tier sets how a request is processed and what it costs.
| Tier | Price | What it is |
|---|---|---|
default | 1× | Standard processing. The tier of a request that names none. |
flex | 0.5× | Best effort, at a lower price. For work that is not urgent. |
priority | 1.75× | Priority processing, at a higher price. |
auto is accepted too, and means the same as leaving the parameter out: the request names no tier. Any other value is a 400, with service_tier as its param.
When you name no tier
A request that names no tier is normally served on the default tier. When the default tier is slow, it may be served on the flex tier instead: it is then billed at the flex rate, and the response says "service_tier": "flex". A request is never moved to a tier that costs more.
"service_tier": "default" keeps a request on the default tier, however slow that is: a tier you name is the tier your request is served on, always.
The service_tier of the response is the tier the request was billed at.
response = client.chat.completions.create(
model="slk-1",
messages=[{"role": "user", "content": "Summarise this report: ..."}],
service_tier="flex",
)
print(response.service_tier) # the tier the request was billed atThe pricing page has the price of each tier.
11
Errors
An error has the body of an OpenAI error: an error object with a message, a type, the param at fault and a code. The status says what kind of problem it is.
{
"error": {
"message": "The model `slk-1-turbo` does not exist or you do not have access to it.",
"type": "invalid_request_error",
"param": null,
"code": "model_not_found"
}
}| Status | Type | Code | What it means |
|---|---|---|---|
| 400 | invalid_request_error | – | A parameter is missing, has the wrong type or breaks a limit. param names it, and code says what is wrong with it. |
| 400 | invalid_request_error | – | The body is not valid JSON, or is not a JSON object. |
| 401 | authentication_error | – | There is no API key in the Authorization header. |
| 401 | authentication_error | invalid_api_key | The key is not a valid one, or it has been revoked. |
| 401 | authentication_error | api_key_expired | The key is past its expiry date. Create a new one. |
| 402 | insufficient_quota | insufficient_quota | The credit balance does not cover the most the request could cost. Add credits, or lower max_tokens. |
| 402 | insufficient_quota | spend_limit_reached | The key has reached its monthly spend limit, or the max_tokens you set does not fit in what is left of it. |
| 403 | permission_error | account_suspended | The account is suspended. Contact support. |
| 403 | permission_error | insufficient_permissions | The key does not have permission to use this endpoint. |
| 404 | invalid_request_error | model_not_found | The model does not exist, or the key is not allowed to use it. |
| 404 | invalid_request_error | – | There is no such endpoint. |
| 413 | invalid_request_error | request_too_large | The request body is larger than 16 MB. |
| 429 | rate_limit_error | rate_limit_exceeded | A rate limit is used up. Wait for the number of seconds in the retry-after header, then try again. |
| 429 | rate_limit_error | model_overloaded | The model is at capacity. Try again after a short wait. |
| 500 | api_error | – | Something went wrong on our side. Try again. |
| 503 | api_error | model_unavailable | The model is temporarily unavailable. Try again later. |
| 504 | api_error | timeout | The request took too long to answer. Try again. |
The OpenAI SDKs retry a 429 and a 5xx on their own before they raise an error.
Errors after the response has started
A request that is still waiting for the model after 10 seconds has its response started with status 200, so that nothing between you and the API times it out. An error that comes after that is inside that 200: for a request that is not streamed, the body holds an error object in place of the completion; for a stream, the error is the last event, and no [DONE] follows it. So check a 200 response for error before you read choices.
12
Rate limits
Limits are counted per minute, for requests and for tokens. They depend on the tier of your account, which rises with its lifetime top-ups.
| Tier | Lifetime top-ups | Requests per minute | Tokens per minute |
|---|---|---|---|
| Tier 1 | None | 60 | 500,000 |
| Tier 2 | $50 | 300 | 2,000,000 |
| Tier 3 | $500 | 1,000 | 5,000,000 |
Requests
A key may make as many requests a minute as the tier allows, or fewer when the key has a limit of its own. All the keys of a workspace together may make no more than the tier allows.
A response that goes through carries these headers for its key, and a 429 from a request limit carries them for the limit that refused it, the key's or the workspace's:
| Header | What it holds |
|---|---|
x-ratelimit-limit-requests | The requests allowed in a minute. |
x-ratelimit-remaining-requests | The requests left in this minute. 0 on a 429. |
x-ratelimit-reset-requests | The time until the count starts again, in seconds, as 12s. |
Tokens
The input and output tokens of the requests a workspace finishes are counted by the minute. Once the count reaches the limit of the tier, new requests are refused until the minute ends. The 429 that refuses them carries the token headers, which no other response has:
| Header | What it holds |
|---|---|
x-ratelimit-limit-tokens | The tokens the workspace may use in a minute. |
x-ratelimit-remaining-tokens | 0: the limit is used up. |
x-ratelimit-reset-tokens | The time until the minute ends, in seconds, as 12s. |
Over a limit
A request over a limit gets a 429 with retry-after, the seconds to wait before trying again, and the three headers of the limit that refused it.
HTTP/1.1 429 Too Many Requests
retry-after: 12
x-ratelimit-limit-requests: 60
x-ratelimit-remaining-requests: 0
x-ratelimit-reset-requests: 12sHTTP/1.1 429 Too Many Requests
retry-after: 12
x-ratelimit-limit-tokens: 500000
x-ratelimit-remaining-tokens: 0
x-ratelimit-reset-tokens: 12sA key with a monthly spend limit has its requests let in one at a time, so that the limit holds however many arrive together; the others wait for their turn. A turn takes about 15 milliseconds, so the last of 1,000 requests sent at once waits about 15 seconds. The line holds 2,048 requests of one key on the server that takes them, more than the limits above let through at one moment, so it is those limits that refuse a burst.
Should 2,048 requests of one key be waiting already, the next gets a 429 with retry-after: 1 and no x-ratelimit headers. That request has been counted against your requests per minute, as the one you send in its place will be.
13
Billing
The API is paid from prepaid credits, at the prices here.
Holds
When a request starts, we place a temporary hold on your balance for the most the request could cost: its input, plus as many output tokens as max_tokens allows. The input is held at the most it can come to, one token for every byte of the request: English text uses about a quarter of that. When the request finishes, the hold is released and the tokens that were used are charged.
When the balance is short
If you did not set max_tokens and the balance, or the monthly limit of the key, cannot cover the default, the limit is lowered automatically to what it can cover.
If you set max_tokens yourself and it does not fit, the request is refused with a 402: insufficient_quota when it is the balance, spend_limit_reached when it is the limit of the key. A limit you set is never lowered for you.
If the balance cannot cover even the input of the request, it is refused with a 402 insufficient_quota, whatever max_tokens is.
Requests that are stopped
A stream that you stop, or whose connection closes, is ended at once, before its usage is counted, and is billed on an estimate: the input by the text of the request, the output by what had been sent. The estimate is on the high side of the real count, and your usage marks the tokens of such a request with a ~.
A request that is not streamed runs to its end even when its connection closes, and is billed for the tokens it used.
No request runs longer than 20 minutes. One that reaches the limit ends with an error, and a stream that had begun to answer by then is billed on the same estimate. For a request that may take long, set stream to true: the answer arrives as it is written.
Prompt caching
Caching is automatic: there is nothing to turn on. Input that is read from the cache is billed at the cache-read price, and counted in usage.prompt_tokens_details.cached_tokens.