Reference
API reference
All endpoints are served from https://api.lirux.ai/v1 and authenticated with a Bearer token. Request and response shapes follow the OpenAI API for supported patterns; differences are noted per endpoint.
Returns the models your API key can call, in the OpenAI list format. The list reflects your plan and, on dedicated nodes, the models deployed on your node. This endpoint takes no parameters.
Request
curl https://api.lirux.ai/v1/models \
-H "Authorization: Bearer $LIRUX_API_KEY"Response
{
"object": "list",
"data": [
{
"id": "qwen3-30b-a3b",
"object": "model",
"created": 1759276800,
"owned_by": "lirux"
},
{
"id": "qwen3-8b",
"object": "model",
"created": 1759276800,
"owned_by": "lirux"
},
{
"id": "bge-m3",
"object": "model",
"created": 1759276800,
"owned_by": "lirux"
}
]
}Generates a model response for a conversation. Works with chat, reasoning and — with image content parts — vision models. Use stream: true for token-by-token output.
Request body
| Parameter | Type | Required | Description |
|---|---|---|---|
modelstringRequired ID of the model to use, e.g. qwen3-30b-a3b. See Models. | string | Required | ID of the model to use, e.g. qwen3-30b-a3b. See Models. |
messagesarrayRequired The conversation so far. Each message has a role (system, user, assistant or tool) and content — a string, or an array of content parts (text, image_url) on vision-capable models. | array | Required | The conversation so far. Each message has a role (system, user, assistant or tool) and content — a string, or an array of content parts (text, image_url) on vision-capable models. |
temperaturenumberOptional Sampling temperature between 0 and 2. Lower is more deterministic. If omitted, the model's recommended default is used. | number | Optional | Sampling temperature between 0 and 2. Lower is more deterministic. If omitted, the model's recommended default is used. |
top_pnumberOptional Nucleus sampling: only tokens within the top top_p probability mass are considered. Adjust this or temperature, not both. | number | Optional | Nucleus sampling: only tokens within the top top_p probability mass are considered. Adjust this or temperature, not both. |
max_tokensintegerOptional Maximum number of tokens to generate. Prompt plus completion must fit within the served context window of the deployment. | integer | Optional | Maximum number of tokens to generate. Prompt plus completion must fit within the served context window of the deployment. |
streambooleanOptional If true, partial deltas are sent as server-sent events, ending with data: [DONE]. Default false. | boolean | Optional | If true, partial deltas are sent as server-sent events, ending with data: [DONE]. Default false. |
stopstring | string[]Optional Up to 4 sequences at which generation stops. The stop sequence is not included in the output. | string | string[] | Optional | Up to 4 sequences at which generation stops. The stop sequence is not included in the output. |
response_formatobjectOptional {"type": "json_object"} constrains output to valid JSON, on supported models. Also instruct the model to answer in JSON in your prompt. | object | Optional | {"type": "json_object"} constrains output to valid JSON, on supported models. Also instruct the model to answer in JSON in your prompt. |
toolsarrayOptional Function definitions the model may call, in the OpenAI {"type": "function", ...} format, on supported models. | array | Optional | Function definitions the model may call, in the OpenAI {"type": "function", ...} format, on supported models. |
tool_choicestring | objectOptional auto, none or a specific function, on models that support tool calling. | string | object | Optional | auto, none or a specific function, on models that support tool calling. |
seedintegerOptional Best effort: requests with the same seed and parameters tend to return the same result, but determinism is not guaranteed. | integer | Optional | Best effort: requests with the same seed and parameters tend to return the same result, but determinism is not guaranteed. |
logprobsbooleanOptional Return log probabilities of output tokens, where the model runtime supports it. See compatibility. | boolean | Optional | Return log probabilities of output tokens, where the model runtime supports it. See compatibility. |
Request
from openai import OpenAI
import os
client = OpenAI(
base_url="https://api.lirux.ai/v1",
api_key=os.environ["LIRUX_API_KEY"],
)
response = client.chat.completions.create(
model="qwen3-30b-a3b",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Summarise our Q3 support tickets."},
],
)
print(response.choices[0].message.content)Response
{
"id": "chatcmpl-9b2e4c1a",
"object": "chat.completion",
"created": 1759312800,
"model": "qwen3-30b-a3b",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "The customer reports failed invoice exports since the last update. A workaround was shared; a fix is scheduled."
},
"finish_reason": "stop"
}
],
"usage": { "prompt_tokens": 41, "completion_tokens": 27, "total_tokens": 68 }
}Response fields
choices[].message— the assistant message. For tool calls,contentmay benullandtool_callsis set.choices[].finish_reason—stop(natural end or stop sequence),length(hitmax_tokensor the context limit) ortool_calls.usage—prompt_tokens,completion_tokensandtotal_tokens, counted with the model's own tokenizer.
JSON mode
On supported models, response_format: {"type": "json_object"} constrains output to syntactically valid JSON. It does not enforce a schema — validate the result in your code.
response = client.chat.completions.create(
model="qwen3-30b-a3b",
response_format={"type": "json_object"},
messages=[
{"role": "system", "content": "Extract fields as JSON with keys: customer, product, issue."},
{"role": "user", "content": ticket_text},
],
)
data = json.loads(response.choices[0].message.content)Tool calling
On models that support function calling, pass tools; the model may answer with tool_calls instead of text. Execute the function in your code, append the result as a tool message with the matching tool_call_id, and call the endpoint again. Reliability varies by model — evaluate with your own tool definitions.
response = client.chat.completions.create(
model="qwen3-30b-a3b",
messages=[{"role": "user", "content": "Where is order 10442?"}],
tools=[
{
"type": "function",
"function": {
"name": "get_order_status",
"description": "Look up the shipping status of an order.",
"parameters": {
"type": "object",
"properties": {"order_id": {"type": "string"}},
"required": ["order_id"],
},
},
}
],
)
call = response.choices[0].message.tool_calls[0]
# call.function.name == "get_order_status"
# call.function.arguments == '{"order_id": "10442"}'Notes
400 invalid_request_error. Reasoning models may spend part of max_tokens on internal reasoning before the final answer — leave enough headroom.With stream: true the response is text/event-stream. Each event carries a chat.completion.chunk whose choices[].delta holds the new content. The final chunk sets finish_reason, followed by data: [DONE]. Errors that occur before the first token are returned as a normal JSON error response with the appropriate status code.
Request
curl -N https://api.lirux.ai/v1/chat/completions \
-H "Authorization: Bearer $LIRUX_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3-30b-a3b",
"messages": [{"role": "user", "content": "Write a release note."}],
"stream": true
}'Event stream
data: {"id":"chatcmpl-7f1c","object":"chat.completion.chunk","created":1759312800,"model":"qwen3-30b-a3b","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]}
data: {"id":"chatcmpl-7f1c","object":"chat.completion.chunk","created":1759312800,"model":"qwen3-30b-a3b","choices":[{"index":0,"delta":{"content":"Release"},"finish_reason":null}]}
data: {"id":"chatcmpl-7f1c","object":"chat.completion.chunk","created":1759312800,"model":"qwen3-30b-a3b","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: [DONE]Vision-capable models accept images as content parts in a user message. Pass each image as an image_url part containing a base64 data URL (data:image/png;base64,… or data:image/jpeg;base64,…). Images count towards the context window.
Vision-capable models in the catalog: gemma-3-27b-it, mistral-small-3.2-24b-instruct, qwen2.5-vl-7b-instruct. Availability depends on your plan; text-only models reject image parts with 400 invalid_request_error.
import base64
with open("invoice.png", "rb") as f:
image_b64 = base64.b64encode(f.read()).decode()
response = client.chat.completions.create(
model="gemma-3-27b-it",
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": "Extract the invoice number and total."},
{
"type": "image_url",
"image_url": {"url": f"data:image/png;base64,{image_b64}"},
},
],
}
],
)Returns a vector representation of each input. Use the same embedding model for indexing and querying — vectors from different models are not comparable.
Request body
| Parameter | Type | Required | Description |
|---|---|---|---|
modelstringRequired ID of an embedding model, e.g. bge-m3. | string | Required | ID of an embedding model, e.g. bge-m3. |
inputstring | string[]Required Text to embed. Pass an array to embed several inputs in one request; each input must fit the model's maximum input length. | string | string[] | Required | Text to embed. Pass an array to embed several inputs in one request; each input must fit the model's maximum input length. |
encoding_formatstringOptional float (default) or base64. | string | Optional | float (default) or base64. |
Request
result = client.embeddings.create(
model="bge-m3",
input=["Wartungsintervall Hydraulikpumpe", "Pump maintenance interval"],
)
vector = result.data[0].embeddingResponse
{
"object": "list",
"data": [
{ "object": "embedding", "index": 0, "embedding": [0.0187, -0.0342, 0.0051, ...] },
{ "object": "embedding", "index": 1, "embedding": [0.0193, -0.0298, 0.0064, ...] }
],
"model": "bge-m3",
"usage": { "prompt_tokens": 14, "total_tokens": 14 }
}Notes
bge-m3 returns 1024-dimensional dense vectors). The dimensions parameter is not supported. Inputs longer than the model's maximum input length return 400 invalid_request_error.Transcribes audio into text in the spoken language. The request is multipart/form-data. Long recordings are processed in windows and returned as one transcript.
Request body (multipart/form-data)
| Parameter | Type | Required | Description |
|---|---|---|---|
filefileRequired The audio file, in a common format such as mp3, m4a, wav or webm. Maximum upload size depends on your plan. | file | Required | The audio file, in a common format such as mp3, m4a, wav or webm. Maximum upload size depends on your plan. |
modelstringRequired ID of a speech model, e.g. whisper-large-v3-turbo. | string | Required | ID of a speech model, e.g. whisper-large-v3-turbo. |
languagestringOptional Language of the audio as an ISO-639-1 code (e.g. de, en). Improves accuracy and latency; detected automatically if omitted. | string | Optional | Language of the audio as an ISO-639-1 code (e.g. de, en). Improves accuracy and latency; detected automatically if omitted. |
promptstringOptional Optional text to guide spelling and style — for example product names or a previous segment. | string | Optional | Optional text to guide spelling and style — for example product names or a previous segment. |
response_formatstringOptional json (default), text or verbose_json. verbose_json adds language, duration and segment timestamps; see notes. | string | Optional | json (default), text or verbose_json. verbose_json adds language, duration and segment timestamps; see notes. |
temperaturenumberOptional Sampling temperature between 0 and 1. Default 0. | number | Optional | Sampling temperature between 0 and 1. Default 0. |
Request
with open("call-2026-10-01.mp3", "rb") as audio:
transcript = client.audio.transcriptions.create(
model="whisper-large-v3-turbo",
file=audio,
)
print(transcript.text)Response
{
"text": "Guten Tag, Sie sprechen mit dem Kundenservice. Wie kann ich helfen?"
}{
"task": "transcribe",
"language": "german",
"duration": 6.4,
"text": "Guten Tag, Sie sprechen mit dem Kundenservice. Wie kann ich helfen?",
"segments": [
{ "id": 0, "start": 0.0, "end": 3.1, "text": "Guten Tag, Sie sprechen mit dem Kundenservice." },
{ "id": 1, "start": 3.1, "end": 6.4, "text": "Wie kann ich helfen?" }
]
}Notes
text returns a plain-text body. verbose_json is partially supported: segment-level timestamps are returned, word-level timestamps are not. Subtitle formats (srt, vtt) and translation (/v1/audio/translations) are not currently supported. Uploads above your plan's limit return 413 request_too_large.Scores each document against a query with a cross-encoder and returns them ordered by relevance. Typical use: retrieve 20–100 candidates with embeddings, rerank, and pass the best few to a chat model.
This endpoint is not part of the OpenAI API, so the OpenAI SDKs have no method for it. The request and response shape follow the Cohere/Jina-style rerank convention; call it with any HTTP client.
Request body
| Parameter | Type | Required | Description |
|---|---|---|---|
modelstringRequired ID of a reranking model, e.g. bge-reranker-v2-m3. | string | Required | ID of a reranking model, e.g. bge-reranker-v2-m3. |
querystringRequired The search query to score documents against. | string | Required | The search query to score documents against. |
documentsstring[]Required Candidate passages, typically the top results from vector or keyword search. | string[] | Required | Candidate passages, typically the top results from vector or keyword search. |
top_nintegerOptional Return only the top_n highest-scoring documents. Defaults to all documents. | integer | Optional | Return only the top_n highest-scoring documents. Defaults to all documents. |
Request
{
"model": "bge-reranker-v2-m3",
"query": "How often should the hydraulic pump be serviced?",
"documents": [
"The conveyor belt must be inspected weekly.",
"Hydraulic pumps require maintenance every 2,000 operating hours.",
"Wartungsintervall der Hydraulikpumpe: alle 2.000 Betriebsstunden."
],
"top_n": 2
}Response
{
"model": "bge-reranker-v2-m3",
"results": [
{ "index": 1, "relevance_score": 0.9412 },
{ "index": 2, "relevance_score": 0.9027 }
]
}results[].index— position of the document in yourdocumentsarray.results[].relevance_score— higher means more relevant. Scores are comparable within one request; do not treat them as calibrated probabilities across queries.
Notes
Limits are set per plan and, on dedicated nodes, sized to your deployment. Every response reports the current request limit state:
| Header | Description |
|---|---|
x-ratelimit-limit-requests | Maximum number of requests allowed in the current window. |
x-ratelimit-remaining-requests | Requests remaining in the current window. |
x-ratelimit-reset-requests | Time until the window resets, as a duration string (e.g. 1s, 6m0s). |
HTTP/1.1 200 OK
content-type: application/json
x-request-id: req_5c1e0b7d
x-ratelimit-limit-requests: 60
x-ratelimit-remaining-requests: 59
x-ratelimit-reset-requests: 1sWhen the limit is exceeded the API returns 429 rate_limit_exceeded with a Retry-After header. See retry guidance.