Skip to content
Private AI infrastructure

Deploy private open-weight models without managing servers, GPUs or inference infrastructure. We run the stack; you get an OpenAI-compatible endpoint.

Find your plan

Pick a model. See the plan and the code.

Deploy a private 32B model behind an OpenAI-compatible API without managing the server. This is what the integration looks like.

2 · Select deployment

Recommended plan

Managed AI Node

From €699 / month

Qwen3 32B on a dedicated, fully managed 96 GB node.

1 × dedicated 96 GB node · private endpoint · managed deployment

Deploy Managed AI
from openai import OpenAI
import os

client = OpenAI(
    base_url="https://api.lirux.ai/v1",
    api_key=os.environ["LIRUX_API_KEY"],
)

response = client.chat.completions.create(
    model="qwen3-32b",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Summarise our Q3 support tickets."},
    ],
)

print(response.choices[0].message.content)

The same code works for shared and dedicated deployments — only the base URL and key differ.

Models

Open-weight models for chat, reasoning, search and speech.

Qwen, Llama, Gemma, Mistral, gpt-oss and Whisper-compatible speech — selected for quality, licensing and fit on 96 GB unified memory.

Why teams choose us

Built for software teams that want to ship AI features without becoming an infrastructure company.

Easy to deploy

Pick a model and a plan. We install, configure and expose it behind an OpenAI-compatible endpoint. Your team writes application code, not infrastructure.

Private infrastructure

Dedicated nodes are single-tenant. Your prompts are not used for training, and content logging is configurable on private deployments.

Managed for you

Model installation, runtime updates, monitoring, API gateway, backups of configuration — handled by engineers who run this every day.

Predictable costs

Fixed monthly plans instead of per-request surprises. You know what AI costs before the month starts, and so does your CFO.

Developer friendly

Use the OpenAI SDKs you already know. Change the base URL, the key and the model name — and keep shipping.

Engineer-led support

Support comes from the engineers who run the platform — in English and German, without a ticket maze.

Compare

Three ways to run open models. One of them is someone else's job.

A qualitative comparison. The right choice depends on your team, workload and compliance needs.

Infrastructure setup

Traditional hyperscaler
You configure services, networking and quotas
Self-hosted hardware
You buy, rack and configure hardware
Managed private AI
Done for you

Model deployment

Traditional hyperscaler
You choose runtimes and deploy
Self-hosted hardware
You build the serving stack
Managed private AI
Installed and configured for you

Monthly predictability

Traditional hyperscaler
Usage based
Self-hosted hardware
Predictable after upfront cost
Managed private AI
Fixed monthly options

Hardware management

Traditional hyperscaler
Abstracted, but GPU capacity varies
Self-hosted hardware
Your team
Managed private AI
Included

API configuration

Traditional hyperscaler
You build gateway, keys and limits
Self-hosted hardware
You build gateway, keys and limits
Managed private AI
OpenAI-compatible endpoint included

Monitoring

Traditional hyperscaler
Available, you configure it
Self-hosted hardware
You build it
Managed private AI
Included

Initial investment

Traditional hyperscaler
Low
Self-hosted hardware
High upfront investment
Managed private AI
None beyond the first month

Support

Traditional hyperscaler
Tiered, often paid
Self-hosted hardware
Internal only
Managed private AI
Engineers who know your deployment

Cost

Stop paying for infrastructure complexity.

Buy a working AI endpoint, not weeks of assembling an AI stack. Built to deliver private AI infrastructure at a fraction of traditional managed compute costs — with fixed monthly plans your finance team can plan around.

You do

  • Choose a model
  • Choose a plan
  • Insert the endpoint into your application
  • Build your product

We do

  • Hardware and networking
  • Model downloads and quantization
  • Inference runtime (e.g. MLX, llama.cpp)
  • OpenAI-compatible API gateway
  • Authentication and rate limits
  • Monitoring and alerting
  • Runtime and model updates
  • Configuration backups

Partner program

Build AI for your clients. We run the infrastructure.

For software agencies that want to offer AI features — and recurring AI revenue — without an infrastructure team. Resell private nodes and managed endpoints under your own client relationships.

  • Private client nodes
  • Agency pricing
  • Deployment support
  • Centralised management
  • White-label architecture
  • Custom endpoints
  1. Your agency
  2. Client application
  3. White-label endpoint
  4. Lirux infrastructure

Infrastructure

Transparent about what runs where.

Large unified-memory configurations enable efficient local deployment of many open-weight models without the cost profile of traditional multi-GPU systems.

Engineer-led sales & support

Commercial operations from Germany / Europe.

Infrastructure region: Georgia

Current compute region, outside the EEA. Transfer safeguards available.

Apple Silicon compute

Apple M3 Ultra nodes with 96 GB unified memory each.

Private networking

Nodes on a private high-speed network; VPN for eligible plans.

NVMe storage

Local NVMe for fast model loading, plus external NVMe where needed.

Managed monitoring

Hardware, runtime and endpoint health with alerting.

Security

Security practices you can inspect, not badges you can't.

We list what we actually do. Certifications will be shown here only once they are achieved.

Security & data processing

TLS encryption

All API traffic is encrypted in transit with TLS. Plain HTTP is not accepted.

API key authentication

Bearer-token authentication with per-key scoping and revocation.

Private endpoints

Dedicated endpoints, IP allowlisting and VPN connectivity on eligible plans.

Network isolation

Dedicated nodes are isolated from other tenants at the network level.

Access control

Administrative access is restricted to named engineers, with key-based authentication and logging.

Monitoring

Hardware, runtime and endpoint health are monitored continuously, with alerting to on-call engineers.

Configurable logs & retention

Choose whether prompt content is logged at all, and for how long, on private deployments.

DPA & SCC support

Data Processing Agreement and Standard Contractual Clauses available where applicable.

Console

Nodes, keys, models and usage in one place.

Manage API keys, see which model runs on which node, track requests and latency, and download invoices. The customer console is rolling out to pilot customers.

Illustrative console preview · synthetic demo data, not a real customer

Your private AI endpoint can be simpler than your cloud bill.

Tell us what you are building. An engineer — not a sales script — will come back with a concrete deployment proposal.