For software agencies
Built for software teams that want to ship AI features without becoming an infrastructure company.
Your clients are asking for AI features. You know how to build the product — but running GPUs, inference servers, model updates and monitoring for every client is a different business. We run that layer for you, behind an endpoint your developers already know how to use.
The problem
What gets in the way today.
- Every client project reinvents the AI stack
- Hyperscaler bills are hard to quote in a fixed-price project
- Clients ask where their data is processed — and expect a clear answer
- Nobody on the team wants to be on call for an inference server
Use cases
What teams build with it.
Client chatbots and assistants
RAG and semantic search
Document analysis
Translation and multilingual content
Speech transcription
Image understanding
How it fits
Where it sits in your architecture.
Your application keeps running where it runs today. It calls a managed endpoint; we operate everything behind it.
- Agency
- Client application
- Client-specific endpoint
- Managed private AI
Recommended plans
Where most teams start.
Fixed monthly plans. All prices exclude VAT.
Growth
Most popularFor production SaaS applications.
€149/ month
Launch pricing
Higher compute allowance, tailored to the selected model
- Higher compute allocation
- Chat completions
- Embeddings
- Speech-to-text
- Priority queue
- Usage analytics
- Email support
Managed AI Node
RecommendedA dedicated, fully managed AI endpoint. Our flagship.
From€699/ month
Launch pricing
- Dedicated 96 GB Apple Silicon environment
- Model installation and configuration
- OpenAI-compatible endpoint
- Monitoring and alerting
- Model and system updates
- Secure API gateway
- Backup configuration
- Technical support
- Deployment assistance
Recommended models
Models that suit this workload.
Model availability depends on licensing, memory requirements and deployment configuration.
Qwen3 30B-A3B
AvailableAlibaba Qwen · 30B total · 3B active (MoE)
Mixture-of-experts model with a small active parameter count — a strong default for assistants, RAG and multilingual chat.
BGE-M3
AvailableBAAI · 568M
Multilingual embedding model for semantic search and RAG across 100+ languages.
BGE Reranker v2 M3
AvailableBAAI · 568M
Cross-encoder reranker that reorders retrieved passages to improve RAG answer quality.
Whisper Large v3 Turbo
AvailableOpenAI (open-weight) · 809M
Speech-to-text model for transcription in many languages, served via a Whisper-compatible endpoint.
Considerations
Worth knowing before you start.
One node per client — or shared across clients
Quote fixed prices with confidence
Resell under your brand
Data & location
Transparent about where your data is processed.
Designed to support GDPR-conscious deployments. Customers remain responsible for determining the appropriate legal basis for their workloads.
- Infrastructure region
- Georgia
- Outside the EEA. All compute and model storage currently run here.
- Commercial operations
- Germany / Europe
- Sales, contracts, onboarding and customer communication.
- International data transfer
- Safeguards available
- For workloads involving EEA personal data, appropriate contractual and technical safeguards may be required.
Tell us what you are building.
An engineer reviews your requirements and comes back with a concrete proposal — model, plan and architecture.