Skip to content

Reference

Errors

The API uses conventional HTTP status codes and returns a JSON error object in the OpenAI format, so existing error handling in the OpenAI SDKs keeps working.

Every error response has a body of this shape:

error.json
{
  "error": {
    "message": "Human-readable description of the problem.",
    "type": "invalid_request_error",
    "code": "invalid_parameter",
    "param": "temperature"
  }
}
Error object fields
FieldDescription
messageHuman-readable explanation. Do not parse it; wording can change.
typeBroad category: invalid_request_error, authentication_error, permission_error, rate_limit_error or server_error.
codeStable machine-readable code. Branch on this in your code.
paramThe request parameter the error relates to, or null.

Every response — successful or not — includes an x-request-id header. Log it, and include it when you contact [email protected].

HTTP status codes and error codes
StatusTypeCodeMeaningAction
400invalid_request_errorinvalid_parameterThe request is malformed or a parameter is invalid. Also returned as context_length_exceeded when prompt plus max_tokens exceed the served context.Fix request
401authentication_errorinvalid_api_keyThe key is missing, malformed or revoked.Fix request
403permission_errormodel_not_in_planThe key is valid but not allowed to use this model or endpoint.Do not retry
403permission_errorip_not_allowedIP allowlisting is enabled and the source address is not on the list.Do not retry
404invalid_request_errormodel_not_foundUnknown model ID or path. Check the ID against GET /v1/models.Fix request
413invalid_request_errorrequest_too_largeThe body or uploaded file exceeds the size limit of your plan.Fix request
429rate_limit_errorrate_limit_exceededThe request rate for your key or plan is exceeded.Retry with backoff
500server_errorinternal_errorUnexpected error on our side. Safe to retry with backoff; report persistent errors with the x-request-id.Retry with backoff
503server_errormodel_loadingThe model is being loaded into memory, for example after an update or restart.Retry with backoff
503server_erroroverloadedCapacity for this model is temporarily exhausted.Retry with backoff

400Bad request

400 · invalid_parameter
{
  "error": {
    "message": "temperature must be between 0 and 2.",
    "type": "invalid_request_error",
    "code": "invalid_parameter",
    "param": "temperature"
  }
}

401Unauthorized

401 · invalid_api_key
{
  "error": {
    "message": "Invalid API key provided.",
    "type": "authentication_error",
    "code": "invalid_api_key",
    "param": null
  }
}

403Forbidden — model not in plan

403 · model_not_in_plan
{
  "error": {
    "message": "The model 'llama-3.3-70b-instruct' is not enabled for this API key.",
    "type": "permission_error",
    "code": "model_not_in_plan",
    "param": "model"
  }
}

403Forbidden — IP not allowed

403 · ip_not_allowed
{
  "error": {
    "message": "Requests from this IP address are not allowed for this endpoint.",
    "type": "permission_error",
    "code": "ip_not_allowed",
    "param": null
  }
}

404Not found

404 · model_not_found
{
  "error": {
    "message": "The model 'qwen3-31b' does not exist.",
    "type": "invalid_request_error",
    "code": "model_not_found",
    "param": "model"
  }
}

413Payload too large

413 · request_too_large
{
  "error": {
    "message": "Request body exceeds the maximum size for this endpoint.",
    "type": "invalid_request_error",
    "code": "request_too_large",
    "param": null
  }
}

429Too many requests

429 · rate_limit_exceeded
Retry-After: 2

{
  "error": {
    "message": "Rate limit reached for requests. Retry after the time given in Retry-After.",
    "type": "rate_limit_error",
    "code": "rate_limit_exceeded",
    "param": null
  }
}

500Internal server error

500 · internal_error
{
  "error": {
    "message": "The server had an error while processing your request.",
    "type": "server_error",
    "code": "internal_error",
    "param": null
  }
}

503Service unavailable — model loading

503 · model_loading
Retry-After: 10

{
  "error": {
    "message": "The model 'qwen3-30b-a3b' is loading. Retry after the time given in Retry-After.",
    "type": "server_error",
    "code": "model_loading",
    "param": null
  }
}

503Service unavailable — overloaded

503 · overloaded
Retry-After: 5

{
  "error": {
    "message": "The model is temporarily overloaded. Retry after the time given in Retry-After.",
    "type": "server_error",
    "code": "overloaded",
    "param": null
  }
}

Retry only errors that can succeed on a second attempt:

  • Retry 429, 500, 503 and network errors, with exponential backoff and jitter. When a Retry-After header is present, wait at least that many seconds.
  • Do not retry 400, 401, 403, 404 or 413 unchanged — the same request will fail again. Fix the request, key or plan first.
  • Cap the number of attempts and the total wait, and make retried operations idempotent on your side (for example, do not double-write results).

The official OpenAI SDKs already retry some of these errors automatically. Often, tuning the built-in behaviour is enough:

Built-in SDK retries
client = OpenAI(
    base_url=os.environ["LIRUX_BASE_URL"],
    api_key=os.environ["LIRUX_API_KEY"],
    max_retries=4,   # SDK retries 429, 5xx and connection errors with backoff
    timeout=60.0,
)

For full control — for example in queue workers — implement the loop yourself:

backoff.py
import os
import random
import time

import openai
from openai import OpenAI

client = OpenAI(
    base_url=os.environ["LIRUX_BASE_URL"],
    api_key=os.environ["LIRUX_API_KEY"],
    max_retries=0,  # we handle retries ourselves below
)

RETRYABLE = (openai.RateLimitError, openai.InternalServerError, openai.APIConnectionError)

def create_with_backoff(max_attempts=5, base_delay=1.0, max_delay=30.0, **kwargs):
    for attempt in range(max_attempts):
        try:
            return client.chat.completions.create(**kwargs)
        except RETRYABLE as err:
            if attempt == max_attempts - 1:
                raise
            # Prefer the server's Retry-After header when present.
            retry_after = None
            response = getattr(err, "response", None)
            if response is not None:
                retry_after = response.headers.get("retry-after")
            if retry_after is not None:
                delay = float(retry_after)
            else:
                # Exponential backoff with full jitter avoids synchronised retries.
                delay = random.uniform(0, min(max_delay, base_delay * 2 ** attempt))
            time.sleep(delay)

response = create_with_backoff(
    model="qwen3-30b-a3b",
    messages=[{"role": "user", "content": "Classify this ticket."}],
)

Sustained 429s or 503s

If you hit limits regularly, your workload has outgrown its allocation. Check service status, then talk to us about a larger plan or dedicated capacity.