9Router API
OpenAI-compatible gateway to many AI providers, with one key and automatic routing.
Modified by SantaiNetwork
Introduction
9Router exposes an OpenAI-compatible API. Point any OpenAI SDK or tool at the base URL below and use your 9Router API key. Requests are routed to the right upstream provider automatically, with rotation and failover handled for you.
Authentication
Send your key as a Bearer token in the Authorization header:
Authorization: Bearer sk-your-9router-key
Each key only ever sees and spends its own quota. Check yours on the usage page.
Base URL
https://YOUR-DOMAIN/v1
Use /v1 as the OpenAI baseURL. All endpoints below live under it.
POST/v1/chat/completions
Create a chat completion. Supports streaming via "stream": true (Server-Sent Events).
Request body
| Field | Type | Notes |
|---|---|---|
model | string | Model id, e.g. your-model. See list models. |
messages | array | OpenAI chat messages (role + content). |
stream | boolean | Optional. Stream tokens as SSE. Add stream_options.include_usage for token counts. |
max_tokens | number | Optional cap on output tokens. |
temperature | number | Optional sampling temperature. |
Example — cURL
curl https://YOUR-DOMAIN/v1/chat/completions \
-H "Authorization: Bearer sk-your-9router-key" \
-H "Content-Type: application/json" \
-d '{
"model": "your-model",
"messages": [{"role": "user", "content": "Hello!"}],
"max_tokens": 100
}'
Example — OpenAI Python SDK
from openai import OpenAI
client = OpenAI(
base_url="https://YOUR-DOMAIN/v1",
api_key="sk-your-9router-key",
)
resp = client.chat.completions.create(
model="your-model",
messages=[{"role": "user", "content": "Hello!"}],
)
print(resp.choices[0].message.content)
Example — streaming (Node / fetch)
const res = await fetch("https://YOUR-DOMAIN/v1/chat/completions", {
method: "POST",
headers: {
"Authorization": "Bearer sk-your-9router-key",
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "your-model",
messages: [{ role: "user", content: "Hello!" }],
stream: true,
stream_options: { include_usage: true },
}),
});
// res.body is a text/event-stream of `data: {...}` chunks
GET/v1/models
List every model id you can call, in OpenAI { "data": [...] } shape.
curl https://YOUR-DOMAIN/v1/models \
-H "Authorization: Bearer sk-your-9router-key"
{
"object": "list",
"data": [
{ "id": "your-model", "object": "model" },
{ "id": "another-model", "object": "model" }
]
}
Call this endpoint to get the exact ids available to your key, then pass an id verbatim as model.
Browse available models
Paste your key to list the models you can call. Read-only — nothing is sent anywhere except this gateway.
GET/v1/usage
Return limits, access, and usage for the key in the Authorization header.
Query
| Param | Notes |
|---|---|
period | Window, e.g. 1d, 7d, 14d, 30d. Default 7d. |
curl "https://YOUR-DOMAIN/v1/usage?period=7d" \
-H "Authorization: Bearer sk-your-9router-key"
Response
{
"key": { "name": "my-key", "createdAt": "..." },
"limits": { "requestsPerMinute": 0, "queueTimeoutMs": 0, "tokenQuota": 0 },
"access": { "restricted": false, "allowedModels": [] },
"usage": {
"tokensUsedAllTime": 50631235,
"tokensRemaining": null,
"period": { "days": 7, "since": "..." },
"requests": 776,
"promptTokens": 50479321,
"completionTokens": 138857,
"totalTokens": 50618178,
"byModel": { "your-model": { "requests": 100, "promptTokens": 1, "completionTokens": 2 } },
"byDay": { "2026-08-15": { "requests": 10, "totalTokens": 123 } }
}
}
Prefer a UI? The usage checker renders this for your key.
Errors
Errors follow the OpenAI shape with an error object.
| Status | Meaning |
|---|---|
401 | Missing or invalid API key. |
403 | Model not permitted for this key (allowlist) or quota exhausted. |
429 | Rate limit / queue timeout — retry after a moment. |
5xx | Upstream provider error; 9Router retries/fails over where possible. |
{ "error": { "message": "Model not permitted for this API key: your-model",
"type": "permission_error", "code": "insufficient_quota" } }
Rate limits & quotas
Keys may have a per-minute request limit (RPM), a request queue with timeout, a total token
quota, and a model allowlist. When RPM is hit, requests queue up to the configured timeout
before returning 429. See your current limits under limits in the
usage endpoint or on the usage page.