Reference
Errors
The API uses conventional HTTP status codes and returns a JSON error object in the OpenAI format, so existing error handling in the OpenAI SDKs keeps working.
Every error response has a body of this shape:
{
"error": {
"message": "Human-readable description of the problem.",
"type": "invalid_request_error",
"code": "invalid_parameter",
"param": "temperature"
}
}| Field | Description |
|---|---|
message | Human-readable explanation. Do not parse it; wording can change. |
type | Broad category: invalid_request_error, authentication_error, permission_error, rate_limit_error or server_error. |
code | Stable machine-readable code. Branch on this in your code. |
param | The request parameter the error relates to, or null. |
Every response — successful or not — includes an x-request-id header. Log it, and include it when you contact [email protected].
| Status | Type | Code | Meaning | Action |
|---|---|---|---|---|
| 400 | invalid_request_error | invalid_parameter | The request is malformed or a parameter is invalid. Also returned as context_length_exceeded when prompt plus max_tokens exceed the served context. | Fix request |
| 401 | authentication_error | invalid_api_key | The key is missing, malformed or revoked. | Fix request |
| 403 | permission_error | model_not_in_plan | The key is valid but not allowed to use this model or endpoint. | Do not retry |
| 403 | permission_error | ip_not_allowed | IP allowlisting is enabled and the source address is not on the list. | Do not retry |
| 404 | invalid_request_error | model_not_found | Unknown model ID or path. Check the ID against GET /v1/models. | Fix request |
| 413 | invalid_request_error | request_too_large | The body or uploaded file exceeds the size limit of your plan. | Fix request |
| 429 | rate_limit_error | rate_limit_exceeded | The request rate for your key or plan is exceeded. | Retry with backoff |
| 500 | server_error | internal_error | Unexpected error on our side. Safe to retry with backoff; report persistent errors with the x-request-id. | Retry with backoff |
| 503 | server_error | model_loading | The model is being loaded into memory, for example after an update or restart. | Retry with backoff |
| 503 | server_error | overloaded | Capacity for this model is temporarily exhausted. | Retry with backoff |
400Bad request
{
"error": {
"message": "temperature must be between 0 and 2.",
"type": "invalid_request_error",
"code": "invalid_parameter",
"param": "temperature"
}
}401Unauthorized
{
"error": {
"message": "Invalid API key provided.",
"type": "authentication_error",
"code": "invalid_api_key",
"param": null
}
}403Forbidden — model not in plan
{
"error": {
"message": "The model 'llama-3.3-70b-instruct' is not enabled for this API key.",
"type": "permission_error",
"code": "model_not_in_plan",
"param": "model"
}
}403Forbidden — IP not allowed
{
"error": {
"message": "Requests from this IP address are not allowed for this endpoint.",
"type": "permission_error",
"code": "ip_not_allowed",
"param": null
}
}404Not found
{
"error": {
"message": "The model 'qwen3-31b' does not exist.",
"type": "invalid_request_error",
"code": "model_not_found",
"param": "model"
}
}413Payload too large
{
"error": {
"message": "Request body exceeds the maximum size for this endpoint.",
"type": "invalid_request_error",
"code": "request_too_large",
"param": null
}
}429Too many requests
Retry-After: 2
{
"error": {
"message": "Rate limit reached for requests. Retry after the time given in Retry-After.",
"type": "rate_limit_error",
"code": "rate_limit_exceeded",
"param": null
}
}500Internal server error
{
"error": {
"message": "The server had an error while processing your request.",
"type": "server_error",
"code": "internal_error",
"param": null
}
}503Service unavailable — model loading
Retry-After: 10
{
"error": {
"message": "The model 'qwen3-30b-a3b' is loading. Retry after the time given in Retry-After.",
"type": "server_error",
"code": "model_loading",
"param": null
}
}503Service unavailable — overloaded
Retry-After: 5
{
"error": {
"message": "The model is temporarily overloaded. Retry after the time given in Retry-After.",
"type": "server_error",
"code": "overloaded",
"param": null
}
}Retry only errors that can succeed on a second attempt:
- Retry
429,500,503and network errors, with exponential backoff and jitter. When aRetry-Afterheader is present, wait at least that many seconds. - Do not retry
400,401,403,404or413unchanged — the same request will fail again. Fix the request, key or plan first. - Cap the number of attempts and the total wait, and make retried operations idempotent on your side (for example, do not double-write results).
The official OpenAI SDKs already retry some of these errors automatically. Often, tuning the built-in behaviour is enough:
client = OpenAI(
base_url=os.environ["LIRUX_BASE_URL"],
api_key=os.environ["LIRUX_API_KEY"],
max_retries=4, # SDK retries 429, 5xx and connection errors with backoff
timeout=60.0,
)For full control — for example in queue workers — implement the loop yourself:
import os
import random
import time
import openai
from openai import OpenAI
client = OpenAI(
base_url=os.environ["LIRUX_BASE_URL"],
api_key=os.environ["LIRUX_API_KEY"],
max_retries=0, # we handle retries ourselves below
)
RETRYABLE = (openai.RateLimitError, openai.InternalServerError, openai.APIConnectionError)
def create_with_backoff(max_attempts=5, base_delay=1.0, max_delay=30.0, **kwargs):
for attempt in range(max_attempts):
try:
return client.chat.completions.create(**kwargs)
except RETRYABLE as err:
if attempt == max_attempts - 1:
raise
# Prefer the server's Retry-After header when present.
retry_after = None
response = getattr(err, "response", None)
if response is not None:
retry_after = response.headers.get("retry-after")
if retry_after is not None:
delay = float(retry_after)
else:
# Exponential backoff with full jitter avoids synchronised retries.
delay = random.uniform(0, min(max_delay, base_delay * 2 ** attempt))
time.sleep(delay)
response = create_with_backoff(
model="qwen3-30b-a3b",
messages=[{"role": "user", "content": "Classify this ticket."}],
)Sustained 429s or 503s