Skip to content

Managed AI API

Open models behind an OpenAI-compatible API. Nothing to run.

Chat, embeddings, speech-to-text and reranking on managed, shared infrastructure. Change the base URL, the API key and the model name in the SDK you already use — and ship. Plans from €49 per month.

OpenAI-compatible endpoints for supported API patterns · Fixed monthly plans · No servers to size

from openai import OpenAI
import os

client = OpenAI(
    base_url="https://api.lirux.ai/v1",
    api_key=os.environ["LIRUX_API_KEY"],
)

response = client.chat.completions.create(
    model="qwen3-30b-a3b",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Summarise our Q3 support tickets."},
    ],
)

print(response.choices[0].message.content)

Endpoints

Four endpoints that cover most AI product features.

One base URL — https://api.lirux.ai/v1 — and the request format your OpenAI client library already speaks. Not every OpenAI parameter or endpoint is supported; the compatibility page lists exactly what is.

Chat completions

POST /v1/chat/completions

Assistants, RAG answers, extraction, classification and drafting. Streaming responses supported.

Included in: All plans, from Developer

Embeddings

POST /v1/embeddings

Multilingual vectors for semantic search, retrieval and deduplication — the backbone of any RAG pipeline.

Included in: All plans, from Developer

Speech-to-text

POST /v1/audio/transcriptions

Whisper-compatible transcription for calls, meetings and voice input, in many languages.

Included in: Growth and Scale

Reranking

POST /v1/rerank

Cross-encoder reranking that reorders retrieved passages before they reach the chat model.

Included in: Scale

Model listing is available at GET /v1/models. See OpenAI compatibility.

Models on the shared API

Served today on shared infrastructure.

Selected for quality, licensing and efficient serving. Larger and specialised models run on dedicated Private AI Nodes.

Full Model Catalog
  • Qwen3 30B-A3B

    Available

    qwen3-30b-a3b

    Mixture-of-experts model with a small active parameter count — a strong default for assistants, RAG and multilingual chat.

    Type
    Chat · Reasoning
    Context
    32K native (longer with RoPE scaling)
    License
    Apache 2.0
  • Qwen3 8B

    Available

    qwen3-8b

    Compact, fast model for high-volume tasks such as tagging, translation drafts and short replies.

    Type
    Chat
    Context
    32K native (longer with RoPE scaling)
    License
    Apache 2.0
  • BGE-M3

    Available

    bge-m3

    Multilingual embedding model for semantic search and RAG across 100+ languages.

    Type
    Embedding
    Context
    8K
    License
    MIT
  • BGE Reranker v2 M3

    Available

    bge-reranker-v2-m3

    Cross-encoder reranker that reorders retrieved passages to improve RAG answer quality.

    Type
    Reranking
    Context
    8K
    License
    Apache 2.0
  • Whisper Large v3 Turbo

    Available

    whisper-large-v3-turbo

    Speech-to-text model for transcription in many languages, served via a Whisper-compatible endpoint.

    Type
    Speech
    Context
    Audio (30 s windows, chunked)
    License
    MIT

Model availability depends on licensing, memory requirements and deployment configuration. Context windows are the model's published maximum; served context is set per deployment.

Plans

Fixed monthly plans. No per-token invoices.

Pick the plan that matches the endpoints you need and the volume you expect. Upgrade when your product grows.

Developer

For testing and small applications.

€49/ month

Launch pricing

Start Developer

Fair-use compute allocation on shared infrastructure

  • Shared AI infrastructure
  • Chat completions API
  • Embeddings API
  • API keys
  • Usage dashboard
  • Standard support
  • Fair-use compute allocation

Growth

Most popular

For production SaaS applications.

€149/ month

Launch pricing

Start Growth

Higher compute allowance, tailored to the selected model

  • Higher compute allocation
  • Chat completions
  • Embeddings
  • Speech-to-text
  • Priority queue
  • Usage analytics
  • Email support

Scale

For growing production workloads.

€349/ month

Launch pricing

Talk to Sales

High compute allowance, sized with you per model

  • High compute allocation
  • Priority inference
  • Chat, embeddings and speech
  • Reranking
  • Higher concurrency
  • Priority support
  • Custom model consultation

All prices exclude VAT. Annual billing: 2 months free, invoiced yearly in advance.

Compare every feature on the pricing page

Compute allowance

Fair use, sized to the model you call.

We don't publish token quotas, because a token on an 8B model and a token on a 30B model cost very different amounts of compute. Each plan comes with a compute allowance tailored to the selected model.

Developer

Fair-use compute allocation on shared infrastructure

Growth

Higher compute allowance, tailored to the selected model

Scale

High compute allowance, sized with you per model

How the allowance works in practice

  1. 1

    You choose the models you use

    Your allowance is set for the models and endpoints in your plan — smaller models stretch further.

  2. 2

    Requests are rate-limited, not billed per token

    Each key has rate limits that protect you and other tenants. Bursts above the limit receive an HTTP 429 and can be retried.

  3. 3

    You see your usage

    The usage dashboard shows requests and token counts per key and model, so there are no surprises at month end.

  4. 4

    We talk before anything changes

    If your usage regularly exceeds the plan's allocation, we contact you with a recommendation — a larger plan or a dedicated node.

Growing out of shared

When to move to a dedicated node.

The shared API is the right start for most teams. These are the signals that a Managed AI Node will serve you better.

  • Your volume is steady enough that a fixed node is simpler to plan than a shared allowance.
  • You need a model that only runs on dedicated capacity — dense 32B, quantized 70B-class, vision or gpt-oss-120b.
  • You want consistent latency that isn't affected by other tenants.
  • You need content logging switched off, or a retention period agreed in writing.
  • You need IP allowlisting, VPN connectivity or a pinned model version.
  • You want to deploy your own fine-tuned, compatible model.
Shared and dedicated endpoints use the same OpenAI-compatible request format. Moving up usually means a new base URL and key — your application code stays as it is.

FAQ

Managed AI API questions.

Can I use an OpenAI SDK?

Yes, for supported API patterns. Our endpoints follow the OpenAI request and response format for chat completions, embeddings, audio transcription and model listing. Most applications switch by changing the base URL, the API key and the model name. Not every OpenAI parameter or endpoint is supported — see OpenAI compatibility in the docs.

What models can I deploy?

Popular open-weight families including Qwen, Llama, Gemma, Mistral and gpt-oss for chat and reasoning, plus embedding, reranking and Whisper-compatible speech models. The model catalog shows the current status of each. Availability depends on licensing, memory requirements and deployment configuration.

Is my data used for model training?

No. Lirux does not train models on customer prompts, responses or files. We serve open-weight models; we do not build our own foundation models from customer data.

What happens to prompts and responses?

Requests are processed in memory to produce a response. Request metadata (timestamp, model, token counts, status code) is logged for billing, rate limiting and operations. Whether prompt and response content is stored at all — and for how long — is configurable on dedicated deployments and defined in your agreement.

Is the hardware dedicated?

On Private AI Node and Mac Cloud plans, yes — the physical Apple Mac Studio (Apple M3 Ultra, 96 GB) is assigned to you alone. Managed AI API plans run on shared infrastructure with per-customer authentication, rate limits and isolation at the API layer.

Where is the infrastructure located?

Our current compute region is Georgia. Commercial operations, sales and customer communication are run from Germany / Europe. Georgia is outside the European Economic Area — see our Data Processing page for what that means for personal data.

How fast can deployment happen?

Managed AI API access can be issued shortly after sign-up. Dedicated nodes are deployed after a short technical call to confirm model, context length and connectivity; timing depends on current capacity and your requirements, and we confirm it before you commit.

Can I cancel monthly?

Monthly API plans can be cancelled at the end of any billing period. Dedicated nodes are offered on monthly or longer terms, with notice periods stated in your order. [POLICY: confirm minimum terms before launch]

Get an API key. Make your first request. Ship.

Start on the shared API, measure what your product actually needs, and move to dedicated capacity when the numbers say so.