Document AI
Search, extract and summarise your documents on private infrastructure.
Most business AI starts with documents. Combine embeddings, reranking, vision and chat models to build RAG systems, extract structured data from forms and invoices, and classify incoming documents — all through one OpenAI-compatible API.
The problem
What gets in the way today.
- Documents contain confidential or personal information
- Extraction pipelines need reliable structured output
- Retrieval quality depends on more than one model
- Volume makes per-page pricing expensive
Use cases
What teams build with it.
Retrieval-augmented generation
Structured extraction
Classification
Summarisation
Scanned documents
Multilingual corpora
How it fits
Where it sits in your architecture.
Your application keeps running where it runs today. It calls a managed endpoint; we operate everything behind it.
- Document store
- Embeddings & rerank
- Chat model
- Grounded answer
Recommended plans
Where most teams start.
Fixed monthly plans. All prices exclude VAT.
Scale
For growing production workloads.
€349/ month
Launch pricing
High compute allowance, sized with you per model
- High compute allocation
- Priority inference
- Chat, embeddings and speech
- Reranking
- Higher concurrency
- Priority support
- Custom model consultation
Managed AI Node
RecommendedA dedicated, fully managed AI endpoint. Our flagship.
From€699/ month
Launch pricing
- Dedicated 96 GB Apple Silicon environment
- Model installation and configuration
- OpenAI-compatible endpoint
- Monitoring and alerting
- Model and system updates
- Secure API gateway
- Backup configuration
- Technical support
- Deployment assistance
Recommended models
Models that suit this workload.
Model availability depends on licensing, memory requirements and deployment configuration.
Qwen3 32B
Private NodeAlibaba Qwen · 32B (dense)
Dense 32B model for higher-quality reasoning, structured output and document analysis on a dedicated node.
BGE-M3
AvailableBAAI · 568M
Multilingual embedding model for semantic search and RAG across 100+ languages.
BGE Reranker v2 M3
AvailableBAAI · 568M
Cross-encoder reranker that reorders retrieved passages to improve RAG answer quality.
Qwen2.5-VL 7B
On RequestAlibaba Qwen · 7B (dense)
Vision-language model for reading documents, screenshots, charts and product images.
Considerations
Worth knowing before you start.
Keep the pipeline together
Choose your logging
Plan memory for context
Data & location
Transparent about where your data is processed.
Designed to support GDPR-conscious deployments. Customers remain responsible for determining the appropriate legal basis for their workloads.
- Infrastructure region
- Georgia
- Outside the EEA. All compute and model storage currently run here.
- Commercial operations
- Germany / Europe
- Sales, contracts, onboarding and customer communication.
- International data transfer
- Safeguards available
- For workloads involving EEA personal data, appropriate contractual and technical safeguards may be required.
Other solutions
See how other teams use the platform.
Tell us what you are building.
An engineer reviews your requirements and comes back with a concrete proposal — model, plan and architecture.