Managed AI API
Open models behind an OpenAI-compatible API. Nothing to run.
Chat, embeddings, speech-to-text and reranking on managed, shared infrastructure. Change the base URL, the API key and the model name in the SDK you already use — and ship. Plans from €49 per month.
OpenAI-compatible endpoints for supported API patterns · Fixed monthly plans · No servers to size
from openai import OpenAI
import os
client = OpenAI(
base_url="https://api.lirux.ai/v1",
api_key=os.environ["LIRUX_API_KEY"],
)
response = client.chat.completions.create(
model="qwen3-30b-a3b",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Summarise our Q3 support tickets."},
],
)
print(response.choices[0].message.content)Endpoints
Four endpoints that cover most AI product features.
One base URL — https://api.lirux.ai/v1 — and the request format your OpenAI client library already speaks. Not every OpenAI parameter or endpoint is supported; the compatibility page lists exactly what is.
Chat completions
POST /v1/chat/completions
Assistants, RAG answers, extraction, classification and drafting. Streaming responses supported.
Included in: All plans, from Developer
Embeddings
POST /v1/embeddings
Multilingual vectors for semantic search, retrieval and deduplication — the backbone of any RAG pipeline.
Included in: All plans, from Developer
Speech-to-text
POST /v1/audio/transcriptions
Whisper-compatible transcription for calls, meetings and voice input, in many languages.
Included in: Growth and Scale
Reranking
POST /v1/rerank
Cross-encoder reranking that reorders retrieved passages before they reach the chat model.
Included in: Scale
Model listing is available at GET /v1/models. See OpenAI compatibility.
Models on the shared API
Served today on shared infrastructure.
Selected for quality, licensing and efficient serving. Larger and specialised models run on dedicated Private AI Nodes.
Qwen3 30B-A3B
Availableqwen3-30b-a3b
Mixture-of-experts model with a small active parameter count — a strong default for assistants, RAG and multilingual chat.
- Type
- Chat · Reasoning
- Context
- 32K native (longer with RoPE scaling)
- License
- Apache 2.0
Qwen3 8B
Availableqwen3-8b
Compact, fast model for high-volume tasks such as tagging, translation drafts and short replies.
- Type
- Chat
- Context
- 32K native (longer with RoPE scaling)
- License
- Apache 2.0
BGE-M3
Availablebge-m3
Multilingual embedding model for semantic search and RAG across 100+ languages.
- Type
- Embedding
- Context
- 8K
- License
- MIT
BGE Reranker v2 M3
Availablebge-reranker-v2-m3
Cross-encoder reranker that reorders retrieved passages to improve RAG answer quality.
- Type
- Reranking
- Context
- 8K
- License
- Apache 2.0
Whisper Large v3 Turbo
Availablewhisper-large-v3-turbo
Speech-to-text model for transcription in many languages, served via a Whisper-compatible endpoint.
- Type
- Speech
- Context
- Audio (30 s windows, chunked)
- License
- MIT
Model availability depends on licensing, memory requirements and deployment configuration. Context windows are the model's published maximum; served context is set per deployment.
Plans
Fixed monthly plans. No per-token invoices.
Pick the plan that matches the endpoints you need and the volume you expect. Upgrade when your product grows.
Developer
For testing and small applications.
€49/ month
Launch pricing
Fair-use compute allocation on shared infrastructure
- Shared AI infrastructure
- Chat completions API
- Embeddings API
- API keys
- Usage dashboard
- Standard support
- Fair-use compute allocation
Growth
Most popularFor production SaaS applications.
€149/ month
Launch pricing
Higher compute allowance, tailored to the selected model
- Higher compute allocation
- Chat completions
- Embeddings
- Speech-to-text
- Priority queue
- Usage analytics
- Email support
Scale
For growing production workloads.
€349/ month
Launch pricing
High compute allowance, sized with you per model
- High compute allocation
- Priority inference
- Chat, embeddings and speech
- Reranking
- Higher concurrency
- Priority support
- Custom model consultation
All prices exclude VAT. Annual billing: 2 months free, invoiced yearly in advance.
Compute allowance
Fair use, sized to the model you call.
We don't publish token quotas, because a token on an 8B model and a token on a 30B model cost very different amounts of compute. Each plan comes with a compute allowance tailored to the selected model.
Developer
Fair-use compute allocation on shared infrastructure
Growth
Higher compute allowance, tailored to the selected model
Scale
High compute allowance, sized with you per model
How the allowance works in practice
- 1
You choose the models you use
Your allowance is set for the models and endpoints in your plan — smaller models stretch further.
- 2
Requests are rate-limited, not billed per token
Each key has rate limits that protect you and other tenants. Bursts above the limit receive an HTTP 429 and can be retried.
- 3
You see your usage
The usage dashboard shows requests and token counts per key and model, so there are no surprises at month end.
- 4
We talk before anything changes
If your usage regularly exceeds the plan's allocation, we contact you with a recommendation — a larger plan or a dedicated node.
Growing out of shared
When to move to a dedicated node.
The shared API is the right start for most teams. These are the signals that a Managed AI Node will serve you better.
- Your volume is steady enough that a fixed node is simpler to plan than a shared allowance.
- You need a model that only runs on dedicated capacity — dense 32B, quantized 70B-class, vision or gpt-oss-120b.
- You want consistent latency that isn't affected by other tenants.
- You need content logging switched off, or a retention period agreed in writing.
- You need IP allowlisting, VPN connectivity or a pinned model version.
- You want to deploy your own fine-tuned, compatible model.
FAQ
Managed AI API questions.
Can I use an OpenAI SDK?
Yes, for supported API patterns. Our endpoints follow the OpenAI request and response format for chat completions, embeddings, audio transcription and model listing. Most applications switch by changing the base URL, the API key and the model name. Not every OpenAI parameter or endpoint is supported — see OpenAI compatibility in the docs.
What models can I deploy?
Popular open-weight families including Qwen, Llama, Gemma, Mistral and gpt-oss for chat and reasoning, plus embedding, reranking and Whisper-compatible speech models. The model catalog shows the current status of each. Availability depends on licensing, memory requirements and deployment configuration.
Is my data used for model training?
No. Lirux does not train models on customer prompts, responses or files. We serve open-weight models; we do not build our own foundation models from customer data.
What happens to prompts and responses?
Requests are processed in memory to produce a response. Request metadata (timestamp, model, token counts, status code) is logged for billing, rate limiting and operations. Whether prompt and response content is stored at all — and for how long — is configurable on dedicated deployments and defined in your agreement.
Is the hardware dedicated?
On Private AI Node and Mac Cloud plans, yes — the physical Apple Mac Studio (Apple M3 Ultra, 96 GB) is assigned to you alone. Managed AI API plans run on shared infrastructure with per-customer authentication, rate limits and isolation at the API layer.
Where is the infrastructure located?
Our current compute region is Georgia. Commercial operations, sales and customer communication are run from Germany / Europe. Georgia is outside the European Economic Area — see our Data Processing page for what that means for personal data.
How fast can deployment happen?
Managed AI API access can be issued shortly after sign-up. Dedicated nodes are deployed after a short technical call to confirm model, context length and connectivity; timing depends on current capacity and your requirements, and we confirm it before you commit.
Can I cancel monthly?
Monthly API plans can be cancelled at the end of any billing period. Dedicated nodes are offered on monthly or longer terms, with notice periods stated in your order. [POLICY: confirm minimum terms before launch]
Get an API key. Make your first request. Ship.
Start on the shared API, measure what your product actually needs, and move to dedicated capacity when the numbers say so.