Developers
Private models behind the API you already know.
Use the official OpenAI SDKs against private open-weight models. One base URL, one API key, one model name — and your existing request code keeps working for supported patterns.
Integration at a glance
OpenAI SDK- Base URL
- https://api.lirux.ai/v1
- API key
- LIRUX_API_KEY=sk-…
- Model
- qwen3-30b-a3b
Everything else — messages, streaming, response parsing — stays as it is.
First request
Three lines of configuration.
Point the client at your endpoint, read the key from the environment, choose a model. That is the integration.
- Works with the official openai packages for Python and Node.js
- Chat, embeddings and transcription through the standard SDK methods
- Server-sent event streaming
- Errors in the OpenAI format, so existing handling keeps working
from openai import OpenAI
import os
client = OpenAI(
base_url="https://api.lirux.ai/v1",
api_key=os.environ["LIRUX_API_KEY"],
)
response = client.chat.completions.create(
model="qwen3-30b-a3b",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Summarise our Q3 support tickets."},
],
)
print(response.choices[0].message.content)Request path
Your code calls an API. We run everything behind it.
Runtime, quantization, gateway, authentication, rate limits and monitoring are managed for you — on shared capacity or on your own dedicated node.
- Your application
- OpenAI-compatible API
- Private AI infrastructure
- Open-weight model
- Response
Endpoints
Chat, embeddings, speech and reranking.
The endpoints a production AI feature needs — from RAG retrieval to call transcription.
Chat completions
POST/v1/chat/completions
Generate responses from chat, reasoning and vision models. Streaming, JSON mode and tool calling on supported models.
Example model: qwen3-30b-a3b
Embeddings
POST/v1/embeddings
Turn text into dense vectors for semantic search, clustering, deduplication and RAG retrieval.
Example model: bge-m3
Audio transcriptions
POST/v1/audio/transcriptions
Whisper-compatible speech-to-text for calls, meetings and voice input. Multipart upload.
Example model: whisper-large-v3-turbo
Rerank
POST/v1/rerank
Score candidate passages against a query to reorder retrieval results. Cohere/Jina-style request shape.
Example model: bge-reranker-v2-m3
List models
GET/v1/models
List the model IDs your API key can call, in the OpenAI list format.
Looking for something else — Assistants, Batch, image generation?
See the compatibility matrixSDK compatibility
No proprietary SDK.
OpenAI-compatible endpoints for supported API patterns. If a tool lets you set an OpenAI base URL, it can usually talk to your endpoint.
OpenAI Python SDK
openai
Set base_url and api_key on the client.
OpenAI Node.js SDK
openai
Set baseURL and apiKey on the client.
Any HTTP client
curl · requests · fetch
Plain JSON over HTTPS with a Bearer token.
Frameworks with a custom base URL
e.g. LangChain, LlamaIndex, Vercel AI SDK
Generally work through their OpenAI-compatible provider options. Check the features you use against the compatibility matrix.
Documentation
Everything you need to integrate.
Migrate from OpenAI
Move compatible workloads without a rewrite.
Change the base URL, the key and the model name. Then evaluate prompt behaviour, tool calling and context length on the open model before you switch production traffic.
from openai import OpenAI
client = OpenAI(
api_key=os.environ["OPENAI_API_KEY"],
)
client.chat.completions.create(
model="gpt-4o-mini",
messages=messages,
)from openai import OpenAI
client = OpenAI(
base_url="https://api.lirux.ai/v1",
api_key=os.environ["LIRUX_API_KEY"],
)
client.chat.completions.create(
model="qwen3-30b-a3b",
messages=messages,
)Get an endpoint for your workload.
Tell us which models and endpoints you need. We deploy, configure and hand over a working API key.