Documentation
Build on a private, OpenAI-compatible API
Lirux serves open-weight models for chat, embeddings, speech and reranking behind OpenAI-compatible endpoints for supported API patterns. If your code already uses an OpenAI SDK, integration is a base URL, an API key and a model name.
About API access
Base URL
https://api.lirux.ai/v1Authentication
Authorization: Bearer sk-…Example model
qwen3-30b-a3bfrom openai import OpenAI
import os
client = OpenAI(
base_url="https://api.lirux.ai/v1",
api_key=os.environ["LIRUX_API_KEY"],
)
response = client.chat.completions.create(
model="qwen3-30b-a3b",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Summarise our Q3 support tickets."},
],
)
print(response.choices[0].message.content)- POST/v1/chat/completionsChat completions — Generate responses from chat, reasoning and vision models. Streaming, JSON mode and tool calling on supported models.
- POST/v1/embeddingsEmbeddings — Turn text into dense vectors for semantic search, clustering, deduplication and RAG retrieval.
- POST/v1/audio/transcriptionsAudio transcriptions — Whisper-compatible speech-to-text for calls, meetings and voice input. Multipart upload.
- POST/v1/rerankRerank — Score candidate passages against a query to reorder retrieval results. Cohere/Jina-style request shape.
- GET/v1/modelsList models — List the model IDs your API key can call, in the OpenAI list format.
The API version is part of the path: every endpoint lives under /v1. Within a version we only make additive changes — new endpoints, new optional parameters, new response fields and new models. Your client should ignore fields it does not recognise.
A change that would break existing integrations ships under a new version path, and the previous version keeps working during a migration period agreed with customers. Model IDs are versioned separately; see pinning model versions.
- All requests use HTTPS. Request and response bodies are JSON (
application/json), except audio uploads, which usemultipart/form-data. - Timestamps are Unix seconds. Token counts are reported by the model's own tokenizer.
- Every response carries an
x-request-idheader. Include it when you contact support about a specific request. - Rate-limit state is returned in
x-ratelimit-*headers — see rate limits.
Questions about an integration, a parameter you need, or a model you want to run: write to [email protected] or talk to an engineer. Moving an existing OpenAI integration? Start with Migrate from OpenAI.
Already integrated? The /v1/models endpoint always shows which model IDs your key can call.