For SaaS companies
Add AI features to your product with inference costs you can put in a spreadsheet.
Whether you build HR software, a CRM, a support desk or a vertical SaaS product, AI features need reliable inference — and gross margins need predictable costs. Run open-weight models on managed infrastructure with a fixed monthly price and an API your engineers already know.
The problem
What gets in the way today.
- Per-token pricing makes AI features hard to price into plans
- Customers ask where their data goes and whether it trains models
- Building a serving stack distracts the product team
- Latency and capacity vary on shared consumer-grade APIs
Use cases
What teams build with it.
In-product assistants
Summaries and drafting
Classification and routing
Semantic search
Analytics narratives
Multilingual features
How it fits
Where it sits in your architecture.
Your application keeps running where it runs today. It calls a managed endpoint; we operate everything behind it.
- Your SaaS backend
- OpenAI-compatible API
- Managed inference
- Open-weight model
Recommended plans
Where most teams start.
Fixed monthly plans. All prices exclude VAT.
Growth
Most popularFor production SaaS applications.
€149/ month
Launch pricing
Higher compute allowance, tailored to the selected model
- Higher compute allocation
- Chat completions
- Embeddings
- Speech-to-text
- Priority queue
- Usage analytics
- Email support
Scale
For growing production workloads.
€349/ month
Launch pricing
High compute allowance, sized with you per model
- High compute allocation
- Priority inference
- Chat, embeddings and speech
- Reranking
- Higher concurrency
- Priority support
- Custom model consultation
Managed AI Node
RecommendedA dedicated, fully managed AI endpoint. Our flagship.
From€699/ month
Launch pricing
- Dedicated 96 GB Apple Silicon environment
- Model installation and configuration
- OpenAI-compatible endpoint
- Monitoring and alerting
- Model and system updates
- Secure API gateway
- Backup configuration
- Technical support
- Deployment assistance
Recommended models
Models that suit this workload.
Model availability depends on licensing, memory requirements and deployment configuration.
Qwen3 30B-A3B
AvailableAlibaba Qwen · 30B total · 3B active (MoE)
Mixture-of-experts model with a small active parameter count — a strong default for assistants, RAG and multilingual chat.
Mistral Small 3.2 24B
Private NodeMistral AI · 24B (dense)
European-developed model with strong instruction following, function calling and image input.
BGE-M3
AvailableBAAI · 568M
Multilingual embedding model for semantic search and RAG across 100+ languages.
Considerations
Worth knowing before you start.
Start shared, move dedicated
Tell your customers a clear data story
Keep a fallback
Data & location
Transparent about where your data is processed.
Designed to support GDPR-conscious deployments. Customers remain responsible for determining the appropriate legal basis for their workloads.
- Infrastructure region
- Georgia
- Outside the EEA. All compute and model storage currently run here.
- Commercial operations
- Germany / Europe
- Sales, contracts, onboarding and customer communication.
- International data transfer
- Safeguards available
- For workloads involving EEA personal data, appropriate contractual and technical safeguards may be required.
Other solutions
See how other teams use the platform.
Tell us what you are building.
An engineer reviews your requirements and comes back with a concrete proposal — model, plan and architecture.