Reference
OpenAI compatibility
Lirux provides OpenAI-compatible endpoints for supported API patterns. This page lists exactly which patterns are covered, so you can check your integration before you move it.
- Supported
Works with the OpenAI SDKs using the documented request and response shape.
- Partial
Available, but depends on the model or covers only part of the OpenAI behaviour. See the note.
- Not supported
Not currently supported. Requests to these endpoints return an error.
| Feature | OpenAI API | Status | Notes |
|---|---|---|---|
| Core endpoints | |||
| Chat completions | POST /v1/chat/completions | Supported | Documented parameters only — see the API reference. |
| Streaming (SSE) | stream: true | Supported | Server-sent chat.completion.chunk events, terminated by [DONE]. |
| Embeddings | POST /v1/embeddings | Supported | float and base64 encodings. dimensions is not supported. |
| Audio transcriptions | POST /v1/audio/transcriptions | Supported | json and text response formats. |
| Transcriptions: verbose_json | response_format | Partial | Segment timestamps are returned; word-level timestamps are not. |
| List models | GET /v1/models | Supported | Lists the models enabled for your key. |
| Model-dependent features | |||
| JSON mode | response_format: json_object | Partial | Supported models only. json_schema structured outputs are not supported; validate output in your code. |
| Tool / function calling | tools, tool_choice | Partial | Supported models only. Reliability differs between models — evaluate with your own tool definitions. |
| Vision input | image_url content parts | Partial | Vision-capable models only, with base64 data URLs. See Vision input. |
| Log probabilities | logprobs | Partial | Depends on the model runtime of the deployment. |
| Reproducible outputs | seed | Partial | Best effort; determinism is not guaranteed. |
| Other OpenAI APIs | |||
| Assistants API | — | Not supported | Not currently supported. Keep conversation state in your application. |
| Responses API | POST /v1/responses | Not supported | Not currently supported. Use chat completions. |
| Batch API | — | Not supported | Not currently supported. For bulk jobs, send concurrent requests within your rate limit or use a dedicated node. |
| Files API | — | Not supported | Not currently supported. |
| Fine-tuning API | — | Not supported | Not currently supported. Custom model deployments can be discussed for dedicated nodes. |
| Image generation | — | Not supported | Not currently supported. |
| Text-to-speech | POST /v1/audio/speech | Not supported | Not currently supported. |
| Moderation | POST /v1/moderations | Not supported | Not currently supported. |
| Realtime API | — | Not supported | Not currently supported. |
Need an endpoint that is not supported? Tell us about your use case — the matrix reflects current support, not a fixed roadmap.
Compatible request shapes do not make two platforms identical. Plan for these differences when you move a workload:
Model names
Model IDs refer to open-weight models such as qwen3-30b-a3b or bge-m3. OpenAI model names (gpt-4o, text-embedding-3-small, …) are not aliased — set the model explicitly. See Models.
Token counts
Every model family has its own tokenizer, so the same text produces different token counts than with OpenAI models — and different counts between our models. Do not reuse tiktoken estimates for context budgeting; use the usage field of responses.
Rate limits
Limits are set per plan on the shared API and sized to the deployment on dedicated nodes. They are reported in x-ratelimit-* headers; the token-based limit headers of OpenAI are not returned. See rate-limit headers.
Model behaviour
Open models respond differently from GPT models to the same prompt: system-prompt adherence, verbosity, formatting, refusal behaviour and tool-calling reliability all vary. Re-test prompts and run an evaluation set before switching production traffic. Reasoning models may also emit their reasoning before the final answer.
Context windows
The served context of a deployment can be shorter than the model's published maximum, because the KV cache shares memory with the model weights. Requests that exceed it return 400 with code context_length_exceeded.
Unsupported parameters
Parameters that are not documented in the API reference are not supported. Depending on the parameter they are ignored or rejected — test any parameter your code relies on.
For supported patterns, the change is the client configuration and the model name:
from openai import OpenAI
client = OpenAI(
base_url="https://api.lirux.ai/v1",
api_key=os.environ["LIRUX_API_KEY"],
)
client.chat.completions.create(
model="qwen3-30b-a3b",
messages=messages,
)