Errors
Errors use the OpenAI shape. OpenAI SDKs raise their usual exceptions.
{
"error": {
"message": "Rate limit reached: 60 requests per minute. Retry in 12s.",
"type": "rate_limit_error",
"param": null,
"code": "rate_limit_requests"
}
}Codes
| Status | Code | Meaning | Retry |
|---|---|---|---|
| 400 | invalid_json | The body is not valid JSON. | Don't retry |
| 400 | invalid_request | A parameter is missing or has the wrong type. | Don't retry |
| 400 | context_length_exceeded | Prompt plus reply exceed the 16,384-token context window. | Don't retry |
| 400 | unsupported_parameter | We don't support this parameter, e.g. n > 1 or logit_bias. | Don't retry |
| 400 | unsupported_content | A message has non-text content, e.g. an image. | Don't retry |
| 401 | missing_api_key | No Authorization: Bearer header. | Don't retry |
| 401 | invalid_api_key | The key does not exist. | Don't retry |
| 403 | key_revoked | The key was revoked. Create a new one. | Don't retry |
| 403 | account_disabled | The account is disabled. Contact us. | Don't retry |
| 404 | model_not_found | Unknown model. See GET /v1/models. | Don't retry |
| 404 | unknown_endpoint | We serve chat completions and models only. | Don't retry |
| 429 | rate_limit_requests | Over 60 requests in a minute. | Retry after Retry-After |
| 429 | rate_limit_tokens | Over 500,000 tokens today. | Retry after Retry-After |
| 429 | concurrency_limit | Over 2 requests at once on this key. | Retry after Retry-After |
| 500 | internal_error | Something failed on our side. | Retry with backoff |
| 503 | model_unavailable | The model is offline or loading. | Retry with backoff |
| 503 | engine_busy | The model is at capacity. | Retry with backoff |
| 503 | model_timeout | The model did not answer in time. Try a shorter prompt. | Retry with backoff |
| 503 | stream_interrupted | The model stopped mid-reply. | Retry with backoff |
When to retry
429: wait for theRetry-Afterheader, in seconds, then retry.500and503: retry with exponential backoff, for example 1 s, 2 s, 4 s.- Other
4xx: don't retry. Fix the request.
Rate-limit headers
Every accepted request returns these headers. A 429 returns Retry-After instead.
| Header | Meaning |
|---|---|
x-ratelimit-limit-requests | Requests allowed per minute. |
x-ratelimit-remaining-requests | Requests left in the current minute. |
x-ratelimit-reset-requests | Time until the minute resets, e.g. 12s. |
x-ratelimit-limit-tokens | Tokens allowed per day. |
x-ratelimit-remaining-tokens | Tokens left today. |
x-request-id | This request's id. Include it when you contact us. |
Errors in a stream
If the model stops unexpectedly during a stream, the last event is an error, not [DONE]:
data: {"error":{"message":"The model stopped unexpectedly. Retry the request.","type":"server_error","param":null,"code":"stream_interrupted"}}Check each event for an error field before you read choices.