Choose a model
Qwen, Llama, Gemma, Mistral, gpt-oss or Whisper — or a compatible model of your own.
Deploy private open-weight models without managing servers, GPUs or inference infrastructure. We run the stack; you get an OpenAI-compatible endpoint.
You buy an endpoint, capacity, privacy and a predictable monthly cost. The hardware is our job.
Explore the productsQwen, Llama, Gemma, Mistral, gpt-oss or Whisper — or a compatible model of your own.
Shared API from €49 / month, or a managed 96 GB node from €699.
Runtime, quantization, gateway, keys, monitoring and updates are handled by our engineers.
Point any OpenAI SDK at your base URL. That is the integration.
Products
The hardware is the implementation. What you buy is an endpoint, capacity, privacy, a predictable monthly cost — and engineers who keep it running.
Hosted open models behind an OpenAI-compatible API. No servers, no GPUs — just an endpoint and a key.
A dedicated 96 GB Apple Silicon node running the models you choose — deployed, monitored and updated by us.
Dedicated Apple Silicon for Xcode builds, CI/CD runners and automated testing, with maintenance included.
Request path
Find your plan
Deploy a private 32B model behind an OpenAI-compatible API without managing the server. This is what the integration looks like.
Recommended plan
Managed AI Node
From €699 / month
Qwen3 32B on a dedicated, fully managed 96 GB node.
1 × dedicated 96 GB node · private endpoint · managed deployment
Deploy Managed AIfrom openai import OpenAI
import os
client = OpenAI(
base_url="https://api.lirux.ai/v1",
api_key=os.environ["LIRUX_API_KEY"],
)
response = client.chat.completions.create(
model="qwen3-32b",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Summarise our Q3 support tickets."},
],
)
print(response.choices[0].message.content)The same code works for shared and dedicated deployments — only the base URL and key differ.
Models
Qwen, Llama, Gemma, Mistral, gpt-oss and Whisper-compatible speech — selected for quality, licensing and fit on 96 GB unified memory.
Why teams choose us
Compare
A qualitative comparison. The right choice depends on your team, workload and compliance needs.
| Traditional hyperscaler | Self-hosted hardware | Managed private AI | |
|---|---|---|---|
| Infrastructure setup | You configure services, networking and quotas | You buy, rack and configure hardware | Done for you |
| Model deployment | You choose runtimes and deploy | You build the serving stack | Installed and configured for you |
| Monthly predictability | Usage based | Predictable after upfront cost | Fixed monthly options |
| Hardware management | Abstracted, but GPU capacity varies | Your team | Included |
| API configuration | You build gateway, keys and limits | You build gateway, keys and limits | OpenAI-compatible endpoint included |
| Monitoring | Available, you configure it | You build it | Included |
| Initial investment | Low | High upfront investment | None beyond the first month |
| Support | Tiered, often paid | Internal only | Engineers who know your deployment |
Infrastructure setup
Model deployment
Monthly predictability
Hardware management
API configuration
Monitoring
Initial investment
Support
Cost
Buy a working AI endpoint, not weeks of assembling an AI stack. Built to deliver private AI infrastructure at a fraction of traditional managed compute costs — with fixed monthly plans your finance team can plan around.
You do
We do
Solutions
Ship AI features for clients without running the stack.
Learn morePredictable inference for product features.
Learn moreSearch, product content and translation.
Learn morePrivate RAG over maintenance and quality docs.
Learn moreExtraction, classification, summarisation.
Learn moreTranscribe, summarise and classify calls.
Learn moreDedicated Apple Silicon build runners.
Learn moreSomething else?
Get an architecture recommendationPartner program
For software agencies that want to offer AI features — and recurring AI revenue — without an infrastructure team. Resell private nodes and managed endpoints under your own client relationships.
Infrastructure
Large unified-memory configurations enable efficient local deployment of many open-weight models without the cost profile of traditional multi-GPU systems.
Commercial operations from Germany / Europe.
Current compute region, outside the EEA. Transfer safeguards available.
Apple M3 Ultra nodes with 96 GB unified memory each.
Nodes on a private high-speed network; VPN for eligible plans.
Local NVMe for fast model loading, plus external NVMe where needed.
Hardware, runtime and endpoint health with alerting.
Security
We list what we actually do. Certifications will be shown here only once they are achieved.
Console
Manage API keys, see which model runs on which node, track requests and latency, and download invoices. The customer console is rolling out to pilot customers.
Pricing
Launch pricing. Fixed monthly plans, excluding VAT.
Tell us what you are building. An engineer — not a sales script — will come back with a concrete deployment proposal.