Private AI Node · Flagship
Your models. Your endpoint. Your private AI infrastructure.
A dedicated Apple M3 Ultra node with 96 GB of unified memory, running the open-weight models you choose behind a private, OpenAI-compatible endpoint. We deploy it, monitor it and keep it updated — you call it from your application.
- Managed AI Node
- From €699 / month
- Dedicated Node
- From €449 / month
Private AI Node · spec sheet
Single-tenant- Node
- Apple Mac Studio
- Chip
- Apple M3 Ultra
- Unified memory
- 96 GB, addressable by CPU and GPU
- Memory bandwidth
- 819 GB/s (Apple specification)
- Storage
- Local NVMe SSD per node, with additional external NVMe where required
- Network
- Private high-speed local network between nodes
Apple Silicon (arm64), unified memory shared by CPU and GPU · Compute region: Georgia
What's included in Managed AI Node
A production AI endpoint, not a server you have to babysit.
Everything between the hardware and your HTTP request is our job. That is the difference between renting a machine and buying a working private AI service.
Dedicated hardware
Model installation
OpenAI-compatible endpoint
Secure API gateway
Monitoring and alerting
Updates and backups
Your part
- Choose a model
- Choose a plan
- Insert the endpoint into your application
- Build your product
Our part
- Hardware and networking
- Model downloads and quantization
- Inference runtime (e.g. MLX, llama.cpp)
- OpenAI-compatible API gateway
- Authentication and rate limits
- Monitoring and alerting
- Runtime and model updates
- Configuration backups
What fits on one node
96 GB of unified memory changes which models a single machine can hold.
From compact 8B models running side by side to quantized 70B-class models and gpt-oss-120b. Fit depends on parameter count, quantization and the context length you need.
- 01S · ≤ 24 GB
7B–8B
Can fit several times over — for example a chat model next to embedding, reranking and speech models on the same node.
- Qwen3 8BAvailable
- Llama 3.1 8B InstructOn Request
- Qwen2.5-VL 7BOn Request
- 02S · ≤ 24 GB
14B
Can fit at 8-bit with generous memory left for long context and concurrent requests.
- Qwen3 14BPrivate Node
- 03S/M · ≤ 48 GB
24B–32B
A common choice for one dedicated node: 4- or 8-bit weights with room for the KV cache.
- Mistral Small 3.2 24BPrivate Node
- Gemma 3 27BPrivate Node
- Qwen3 30B-A3BAvailable
- Qwen3 32BPrivate Node
- DeepSeek-R1 Distill Qwen 32BOn Request
- 04M · ≤ 48 GB
70B-class (quantized)
Can fit at 4-bit. Usable context is constrained by the memory left for the KV cache — validated on your workload first.
- Llama 3.3 70B InstructPrivate Node
- 05L · ≤ 80 GB
Large MoE · ~120B
Native MXFP4 weights can fit on a single node but occupy most of it. Plan for one model per node.
- gpt-oss 120BOn Request
XS · < 8 GB
Small embedding, reranking and speech models
S · 8–24 GB
7B–14B models at 4–8-bit, or ~30B at 4-bit
M · 24–48 GB
~30B models at 8-bit or 70B-class at 4-bit
L · 48–80 GB
Large models — uses most of a 96 GB node
Memory classes are approximate weight footprints and exclude the KV cache. On a 96 GB node, the memory not used by weights is what holds context: longer context windows and more concurrent requests need more of it. We confirm fit for your model, quantization and context length before deployment.
Architecture
How a Private AI Node sits behind your application.
Your application calls a private endpoint over HTTPS or VPN. The gateway authenticates and rate-limits every request before it reaches the model runtime on your node. Model weights load from local NVMe.
Private by design
Your node serves only your keys. Content logging is configurable and can be switched off.
Your connectivity
Public HTTPS with IP allowlisting, or private connectivity over VPN on eligible deployments.
Pinned versions
Model and runtime versions change only when you agree — no silent upgrades under your product.
Your application
Anywhere in the world
Secure API gateway
TLS termination · OpenAI-compatible routes
Authentication & rate limits
API keys · per-key quotas · allowlists
Load balancer
Model-aware routing
AI Node 01
96 GB unified
AI Node 02
96 GB unified
AI Node 03
96 GB unified
AI Node 04
96 GB unified
NVMe storage & model repository
Local NVMe · versioned model weights
Compute region: Georgia · Apple M3 Ultra
Why Apple Silicon
Large unified memory, used deliberately.
Large unified-memory configurations enable efficient local deployment of many open-weight models without the cost profile of traditional multi-GPU systems.
On Apple M3 Ultra, CPU and GPU share one pool of 96 GB. The GPU can address all of it, so a model that would otherwise be split across several accelerator cards can be held by a single node. Memory bandwidth is 819 GB/s (Apple specification).
For inference on open-weight models — chat, RAG, extraction, embeddings and transcription — that is a practical, predictable platform. It is not the right platform for everything, and we would rather tell you before you sign.
Know the limits
- No NVIDIA CUDA: workloads that depend on CUDA kernels, TensorRT or CUDA-only libraries need a different architecture.
- Model fit depends on parameter count, quantization and context length — the KV cache needs memory too.
- Throughput depends on model, quantization, context length and concurrency, and is validated per workload.
Running CUDA workloads?
Benchmarks
Measured numbers, published only when they are real.
Benchmark data coming after production validation. We will publish throughput, latency and memory figures with methodology — never estimates. Until then, a pilot benchmarks your actual workload.
Reference: Qwen3 30B-A3B · 4-bit · M3 Ultra · 96 GB
Benchmark methodologyTokens / second
Benchmark data coming after production validation.
Generation throughput for a single stream.
Time to first token
Benchmark data coming after production validation.
Latency until the first streamed token, at a fixed prompt length.
Concurrent requests
Benchmark data coming after production validation.
Concurrent streams sustained before latency targets are exceeded.
Memory usage
Benchmark data coming after production validation.
Resident memory including KV cache at the tested context length.
Plans
Dedicated capacity at a fixed monthly price.
Start with a Dedicated Node if your team runs its own stack. Choose Managed AI Node if you want a working endpoint we operate for you. Add a second node with High Availability.
Dedicated Node
Entry dedicated compute for teams that run their own stack.
From€449/ month
Launch pricing
- Dedicated Apple Silicon node
- Up to 96 GB unified memory
- Private, single-tenant environment
- Private API endpoint
- Monitoring
- Basic managed deployment
Managed AI Node
RecommendedA dedicated, fully managed AI endpoint. Our flagship.
From€699/ month
Launch pricing
- Dedicated 96 GB Apple Silicon environment
- Model installation and configuration
- OpenAI-compatible endpoint
- Monitoring and alerting
- Model and system updates
- Secure API gateway
- Backup configuration
- Technical support
- Deployment assistance
High Availability
Two-node architecture for production workloads that must stay up.
From€1,299/ month
Launch pricing
- Two-node architecture
- Load balancing across nodes
- Redundant capacity
- Monitoring
- Failover design
- Priority support
- Custom deployment
- Private networking
All prices exclude VAT. Annual billing: 2 months free, invoiced yearly in advance.
High Availability
When one node is not enough, add a second one.
High-availability architecture available from From €1,299 / month. Two dedicated nodes behind a load balancer, with redundant capacity and a failover design agreed with you.
Two dedicated nodes
The same model deployed on both, or workloads split by role.
Load balancing
Requests are distributed across healthy nodes.
Redundant capacity
Maintenance and updates can be done one node at a time.
Private networking
Nodes communicate over a private network.
Failover behaviour — what happens to in-flight requests, how capacity is reserved, how quickly traffic shifts — is designed and tested with you during deployment rather than promised in a brochure. Production SLA options are available for eligible managed deployments.
Two-node reference design
- Your application
- Secure gateway
- Load balancer
- Node A · Node B
Security & data location
Clear about where your node runs and who can touch it.
Designed to support GDPR-conscious deployments. Contractual and technical safeguards are available for European customers; customers remain responsible for determining the appropriate legal basis for their workloads.
- Infrastructure region
- Georgia
- Outside the EEA. All compute and model storage currently run here.
- Commercial operations
- Germany / Europe
- Sales, contracts, onboarding and customer communication.
- International data transfer
- Safeguards available
- For workloads involving EEA personal data, appropriate contractual and technical safeguards may be required.
No training on your prompts, responses or files
Content logging configurable — or off — on your node
DPA and Standard Contractual Clauses where applicable
Read how we handle data processing · An EU region is not currently available.
Console
See your node, your models and your keys in one place.
Which model runs where, memory in use, request volume and latency, API keys and invoices. The customer console is rolling out to pilot customers; the view below is illustrative.
FAQ
Private AI Node questions.
Why use a dedicated node?
A dedicated node gives you predictable capacity that no other customer shares, a private endpoint, freedom to choose and pin model versions, and a fixed monthly cost. It suits production workloads, sensitive data and teams that want consistent latency.
Can you deploy 70B models?
Some 70B-class models can run in quantized form (typically 4-bit) on a single 96 GB node. Usable context length is constrained by the memory left for the KV cache. We validate quality and performance on your workload before committing.
Can I bring or run my own model?
On Private AI Nodes, yes — including fine-tuned variants of supported architectures — provided the model is technically compatible with Apple Silicon inference runtimes, fits in memory with your required context length, and you have the rights to deploy it. We check compatibility before deployment.
Can I connect over VPN?
Private connectivity (for example a site-to-site VPN or WireGuard tunnel) is available on eligible Private AI Node and High Availability deployments. Contact us for deployment requirements.
Do you support static IP allowlisting?
Yes, on dedicated deployments you can restrict your endpoint to a list of source IP addresses. Combined with API keys and TLS, this is a common setup for server-to-server integrations.
Do you support CUDA workloads?
No. Our current infrastructure is Apple Silicon, which does not run NVIDIA CUDA. Workloads that depend on CUDA kernels, TensorRT or CUDA-only libraries — such as many training pipelines — need a different architecture, and we will tell you honestly if that applies to you.
Why Apple Silicon?
Large unified-memory configurations enable efficient local deployment of many open-weight models without the cost profile of traditional multi-GPU systems. Unified memory means the GPU can address most of the 96 GB, which lets a single node hold models that would otherwise require several GPUs. The right hardware always depends on the workload, and we will recommend accordingly.
Can I scale to multiple nodes?
Yes. You can add dedicated nodes behind a load balancer, separate workloads per node (for example chat on one, embeddings and speech on another), or move to a High Availability architecture. Capacity is finite, so larger expansions are planned with you in advance.
Private AI Pilot
Not sure yet? Test your workload first.
Validate your workload on dedicated infrastructure before committing to a larger deployment.
Pricing: Contact sales · scoped to your workload
- Model deployment on dedicated capacity
- Private OpenAI-compatible endpoint
- Benchmarking on your real workload
- Architecture recommendation
- Migration plan
A private AI endpoint, running on hardware that is yours alone.
Managed AI Node from €699 / month. Tell us the model and the workload; an engineer confirms fit, context length and connectivity before you commit.