Skip to content

Reference

API reference

All endpoints are served from https://api.lirux.ai/v1 and authenticated with a Bearer token. Request and response shapes follow the OpenAI API for supported patterns; differences are noted per endpoint.

GET/v1/models

Returns the models your API key can call, in the OpenAI list format. The list reflects your plan and, on dedicated nodes, the models deployed on your node. This endpoint takes no parameters.

Request

curl https://api.lirux.ai/v1/models \
  -H "Authorization: Bearer $LIRUX_API_KEY"

Response

200 OK
{
  "object": "list",
  "data": [
    {
      "id": "qwen3-30b-a3b",
      "object": "model",
      "created": 1759276800,
      "owned_by": "lirux"
    },
    {
      "id": "qwen3-8b",
      "object": "model",
      "created": 1759276800,
      "owned_by": "lirux"
    },
    {
      "id": "bge-m3",
      "object": "model",
      "created": 1759276800,
      "owned_by": "lirux"
    }
  ]
}
POST/v1/chat/completions

Generates a model response for a conversation. Works with chat, reasoning and — with image content parts — vision models. Use stream: true for token-by-token output.

Request body

Chat completions request parameters
ParameterTypeRequiredDescription
model
stringRequired
ID of the model to use, e.g. qwen3-30b-a3b. See Models.
messages
arrayRequired
The conversation so far. Each message has a role (system, user, assistant or tool) and content — a string, or an array of content parts (text, image_url) on vision-capable models.
temperature
numberOptional
Sampling temperature between 0 and 2. Lower is more deterministic. If omitted, the model's recommended default is used.
top_p
numberOptional
Nucleus sampling: only tokens within the top top_p probability mass are considered. Adjust this or temperature, not both.
max_tokens
integerOptional
Maximum number of tokens to generate. Prompt plus completion must fit within the served context window of the deployment.
stream
booleanOptional
If true, partial deltas are sent as server-sent events, ending with data: [DONE]. Default false.
stop
string | string[]Optional
Up to 4 sequences at which generation stops. The stop sequence is not included in the output.
response_format
objectOptional
{"type": "json_object"} constrains output to valid JSON, on supported models. Also instruct the model to answer in JSON in your prompt.
tools
arrayOptional
Function definitions the model may call, in the OpenAI {"type": "function", ...} format, on supported models.
tool_choice
string | objectOptional
auto, none or a specific function, on models that support tool calling.
seed
integerOptional
Best effort: requests with the same seed and parameters tend to return the same result, but determinism is not guaranteed.
logprobs
booleanOptional
Return log probabilities of output tokens, where the model runtime supports it. See compatibility.

Request

from openai import OpenAI
import os

client = OpenAI(
    base_url="https://api.lirux.ai/v1",
    api_key=os.environ["LIRUX_API_KEY"],
)

response = client.chat.completions.create(
    model="qwen3-30b-a3b",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Summarise our Q3 support tickets."},
    ],
)

print(response.choices[0].message.content)

Response

200 OK
{
  "id": "chatcmpl-9b2e4c1a",
  "object": "chat.completion",
  "created": 1759312800,
  "model": "qwen3-30b-a3b",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "The customer reports failed invoice exports since the last update. A workaround was shared; a fix is scheduled."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": { "prompt_tokens": 41, "completion_tokens": 27, "total_tokens": 68 }
}

Response fields

  • choices[].message — the assistant message. For tool calls, content may be null and tool_calls is set.
  • choices[].finish_reason — stop (natural end or stop sequence), length (hit max_tokens or the context limit) or tool_calls.
  • usage — prompt_tokens, completion_tokens and total_tokens, counted with the model's own tokenizer.

JSON mode

On supported models, response_format: {"type": "json_object"} constrains output to syntactically valid JSON. It does not enforce a schema — validate the result in your code.

json_mode.py
response = client.chat.completions.create(
    model="qwen3-30b-a3b",
    response_format={"type": "json_object"},
    messages=[
        {"role": "system", "content": "Extract fields as JSON with keys: customer, product, issue."},
        {"role": "user", "content": ticket_text},
    ],
)

data = json.loads(response.choices[0].message.content)

Tool calling

On models that support function calling, pass tools; the model may answer with tool_calls instead of text. Execute the function in your code, append the result as a tool message with the matching tool_call_id, and call the endpoint again. Reliability varies by model — evaluate with your own tool definitions.

tools.py
response = client.chat.completions.create(
    model="qwen3-30b-a3b",
    messages=[{"role": "user", "content": "Where is order 10442?"}],
    tools=[
        {
            "type": "function",
            "function": {
                "name": "get_order_status",
                "description": "Look up the shipping status of an order.",
                "parameters": {
                    "type": "object",
                    "properties": {"order_id": {"type": "string"}},
                    "required": ["order_id"],
                },
            },
        }
    ],
)

call = response.choices[0].message.tool_calls[0]
# call.function.name == "get_order_status"
# call.function.arguments == '{"order_id": "10442"}'

Notes

Parameters not listed here are not supported; depending on the parameter they are ignored or rejected with a 400 invalid_request_error. Reasoning models may spend part of max_tokens on internal reasoning before the final answer — leave enough headroom.

With stream: true the response is text/event-stream. Each event carries a chat.completion.chunk whose choices[].delta holds the new content. The final chunk sets finish_reason, followed by data: [DONE]. Errors that occur before the first token are returned as a normal JSON error response with the appropriate status code.

Request

cURL
curl -N https://api.lirux.ai/v1/chat/completions \
  -H "Authorization: Bearer $LIRUX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3-30b-a3b",
    "messages": [{"role": "user", "content": "Write a release note."}],
    "stream": true
  }'

Event stream

text/event-stream
data: {"id":"chatcmpl-7f1c","object":"chat.completion.chunk","created":1759312800,"model":"qwen3-30b-a3b","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]}

data: {"id":"chatcmpl-7f1c","object":"chat.completion.chunk","created":1759312800,"model":"qwen3-30b-a3b","choices":[{"index":0,"delta":{"content":"Release"},"finish_reason":null}]}

data: {"id":"chatcmpl-7f1c","object":"chat.completion.chunk","created":1759312800,"model":"qwen3-30b-a3b","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: [DONE]

Vision-capable models accept images as content parts in a user message. Pass each image as an image_url part containing a base64 data URL (data:image/png;base64,… or data:image/jpeg;base64,…). Images count towards the context window.

Vision-capable models in the catalog: gemma-3-27b-it, mistral-small-3.2-24b-instruct, qwen2.5-vl-7b-instruct. Availability depends on your plan; text-only models reject image parts with 400 invalid_request_error.

vision.py · gemma-3-27b-it
import base64

with open("invoice.png", "rb") as f:
    image_b64 = base64.b64encode(f.read()).decode()

response = client.chat.completions.create(
    model="gemma-3-27b-it",
    messages=[
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "Extract the invoice number and total."},
                {
                    "type": "image_url",
                    "image_url": {"url": f"data:image/png;base64,{image_b64}"},
                },
            ],
        }
    ],
)
POST/v1/embeddings

Returns a vector representation of each input. Use the same embedding model for indexing and querying — vectors from different models are not comparable.

Request body

Embeddings request parameters
ParameterTypeRequiredDescription
model
stringRequired
ID of an embedding model, e.g. bge-m3.
input
string | string[]Required
Text to embed. Pass an array to embed several inputs in one request; each input must fit the model's maximum input length.
encoding_format
stringOptional
float (default) or base64.

Request

result = client.embeddings.create(
    model="bge-m3",
    input=["Wartungsintervall Hydraulikpumpe", "Pump maintenance interval"],
)

vector = result.data[0].embedding

Response

200 OK
{
  "object": "list",
  "data": [
    { "object": "embedding", "index": 0, "embedding": [0.0187, -0.0342, 0.0051, ...] },
    { "object": "embedding", "index": 1, "embedding": [0.0193, -0.0298, 0.0064, ...] }
  ],
  "model": "bge-m3",
  "usage": { "prompt_tokens": 14, "total_tokens": 14 }
}

Notes

Vector dimensionality is fixed per model (for example, bge-m3 returns 1024-dimensional dense vectors). The dimensions parameter is not supported. Inputs longer than the model's maximum input length return 400 invalid_request_error.
POST/v1/audio/transcriptions

Transcribes audio into text in the spoken language. The request is multipart/form-data. Long recordings are processed in windows and returned as one transcript.

Request body (multipart/form-data)

Audio transcription request parameters
ParameterTypeRequiredDescription
file
fileRequired
The audio file, in a common format such as mp3, m4a, wav or webm. Maximum upload size depends on your plan.
model
stringRequired
ID of a speech model, e.g. whisper-large-v3-turbo.
language
stringOptional
Language of the audio as an ISO-639-1 code (e.g. de, en). Improves accuracy and latency; detected automatically if omitted.
prompt
stringOptional
Optional text to guide spelling and style — for example product names or a previous segment.
response_format
stringOptional
json (default), text or verbose_json. verbose_json adds language, duration and segment timestamps; see notes.
temperature
numberOptional
Sampling temperature between 0 and 1. Default 0.

Request

with open("call-2026-10-01.mp3", "rb") as audio:
    transcript = client.audio.transcriptions.create(
        model="whisper-large-v3-turbo",
        file=audio,
    )

print(transcript.text)

Response

200 OK · json
{
  "text": "Guten Tag, Sie sprechen mit dem Kundenservice. Wie kann ich helfen?"
}
200 OK · verbose_json
{
  "task": "transcribe",
  "language": "german",
  "duration": 6.4,
  "text": "Guten Tag, Sie sprechen mit dem Kundenservice. Wie kann ich helfen?",
  "segments": [
    { "id": 0, "start": 0.0, "end": 3.1, "text": "Guten Tag, Sie sprechen mit dem Kundenservice." },
    { "id": 1, "start": 3.1, "end": 6.4, "text": "Wie kann ich helfen?" }
  ]
}

Notes

text returns a plain-text body. verbose_json is partially supported: segment-level timestamps are returned, word-level timestamps are not. Subtitle formats (srt, vtt) and translation (/v1/audio/translations) are not currently supported. Uploads above your plan's limit return 413 request_too_large.
POST/v1/rerank
Not part of the OpenAI API

Scores each document against a query with a cross-encoder and returns them ordered by relevance. Typical use: retrieve 20–100 candidates with embeddings, rerank, and pass the best few to a chat model.

This endpoint is not part of the OpenAI API, so the OpenAI SDKs have no method for it. The request and response shape follow the Cohere/Jina-style rerank convention; call it with any HTTP client.

Request body

Rerank request parameters
ParameterTypeRequiredDescription
model
stringRequired
ID of a reranking model, e.g. bge-reranker-v2-m3.
query
stringRequired
The search query to score documents against.
documents
string[]Required
Candidate passages, typically the top results from vector or keyword search.
top_n
integerOptional
Return only the top_n highest-scoring documents. Defaults to all documents.

Request

{
  "model": "bge-reranker-v2-m3",
  "query": "How often should the hydraulic pump be serviced?",
  "documents": [
    "The conveyor belt must be inspected weekly.",
    "Hydraulic pumps require maintenance every 2,000 operating hours.",
    "Wartungsintervall der Hydraulikpumpe: alle 2.000 Betriebsstunden."
  ],
  "top_n": 2
}

Response

200 OK
{
  "model": "bge-reranker-v2-m3",
  "results": [
    { "index": 1, "relevance_score": 0.9412 },
    { "index": 2, "relevance_score": 0.9027 }
  ]
}
  • results[].index — position of the document in your documents array.
  • results[].relevance_score — higher means more relevant. Scores are comparable within one request; do not treat them as calibrated probabilities across queries.

Notes

Rerank is available on plans that include reranking and on dedicated nodes — see Pricing.

Limits are set per plan and, on dedicated nodes, sized to your deployment. Every response reports the current request limit state:

Rate-limit response headers
HeaderDescription
x-ratelimit-limit-requestsMaximum number of requests allowed in the current window.
x-ratelimit-remaining-requestsRequests remaining in the current window.
x-ratelimit-reset-requestsTime until the window resets, as a duration string (e.g. 1s, 6m0s).
Response headers (illustrative values)
HTTP/1.1 200 OK
content-type: application/json
x-request-id: req_5c1e0b7d
x-ratelimit-limit-requests: 60
x-ratelimit-remaining-requests: 59
x-ratelimit-reset-requests: 1s

When the limit is exceeded the API returns 429 rate_limit_exceeded with a Retry-After header. See retry guidance.