Skip to content

Private AI Node · Flagship

Your models. Your endpoint. Your private AI infrastructure.

A dedicated Apple M3 Ultra node with 96 GB of unified memory, running the open-weight models you choose behind a private, OpenAI-compatible endpoint. We deploy it, monitor it and keep it updated — you call it from your application.

Managed AI Node
From €699 / month
Dedicated Node
From €449 / month

Private AI Node · spec sheet

Single-tenant
Node
Apple Mac Studio
Chip
Apple M3 Ultra
Unified memory
96 GB, addressable by CPU and GPU
Memory bandwidth
819 GB/s (Apple specification)
Storage
Local NVMe SSD per node, with additional external NVMe where required
Network
Private high-speed local network between nodes

Apple Silicon (arm64), unified memory shared by CPU and GPU · Compute region: Georgia

What's included in Managed AI Node

A production AI endpoint, not a server you have to babysit.

Everything between the hardware and your HTTP request is our job. That is the difference between renting a machine and buying a working private AI service.

Dedicated hardware

A full Apple Mac Studio with 96 GB unified memory, assigned to you alone. No noisy neighbours, no shared queue.

Model installation

We download, quantize and configure the models you choose, and tune context length to the memory available.

OpenAI-compatible endpoint

A private endpoint for supported OpenAI API patterns — chat, embeddings, transcription and model listing.

Secure API gateway

TLS, API keys with per-key scoping and revocation, rate limits and optional IP allowlisting.

Monitoring and alerting

Hardware, memory pressure, runtime and endpoint health watched continuously, with alerts to on-call engineers.

Updates and backups

Runtime and model updates planned with you, and configuration backups so a node can be rebuilt predictably.

Your part

  • Choose a model
  • Choose a plan
  • Insert the endpoint into your application
  • Build your product

Our part

  • Hardware and networking
  • Model downloads and quantization
  • Inference runtime (e.g. MLX, llama.cpp)
  • OpenAI-compatible API gateway
  • Authentication and rate limits
  • Monitoring and alerting
  • Runtime and model updates
  • Configuration backups

What fits on one node

96 GB of unified memory changes which models a single machine can hold.

From compact 8B models running side by side to quantized 70B-class models and gpt-oss-120b. Fit depends on parameter count, quantization and the context length you need.

  1. 01S · ≤ 24 GB

    7B–8B

    Can fit several times over — for example a chat model next to embedding, reranking and speech models on the same node.

  2. 02S · ≤ 24 GB

    14B

    Can fit at 8-bit with generous memory left for long context and concurrent requests.

  3. 03S/M · ≤ 48 GB

    24B–32B

    A common choice for one dedicated node: 4- or 8-bit weights with room for the KV cache.

  4. 04M · ≤ 48 GB

    70B-class (quantized)

    Can fit at 4-bit. Usable context is constrained by the memory left for the KV cache — validated on your workload first.

  5. 05L · ≤ 80 GB

    Large MoE · ~120B

    Native MXFP4 weights can fit on a single node but occupy most of it. Plan for one model per node.

XS · < 8 GB
Small embedding, reranking and speech models

S · 8–24 GB
7B–14B models at 4–8-bit, or ~30B at 4-bit

M · 24–48 GB
~30B models at 8-bit or 70B-class at 4-bit

L · 48–80 GB
Large models — uses most of a 96 GB node

Memory classes are approximate weight footprints and exclude the KV cache. On a 96 GB node, the memory not used by weights is what holds context: longer context windows and more concurrent requests need more of it. We confirm fit for your model, quantization and context length before deployment.

Architecture

How a Private AI Node sits behind your application.

Your application calls a private endpoint over HTTPS or VPN. The gateway authenticates and rate-limits every request before it reaches the model runtime on your node. Model weights load from local NVMe.

  • Private by design

    Your node serves only your keys. Content logging is configurable and can be switched off.

  • Your connectivity

    Public HTTPS with IP allowlisting, or private connectivity over VPN on eligible deployments.

  • Pinned versions

    Model and runtime versions change only when you agree — no silent upgrades under your product.

Your application

Anywhere in the world

Secure API gateway

TLS termination · OpenAI-compatible routes

Authentication & rate limits

API keys · per-key quotas · allowlists

Load balancer

Model-aware routing

AI Node 01

96 GB unified

AI Node 02

96 GB unified

AI Node 03

96 GB unified

AI Node 04

96 GB unified

NVMe storage & model repository

Local NVMe · versioned model weights

Compute region: Georgia · Apple M3 Ultra

Why Apple Silicon

Large unified memory, used deliberately.

Large unified-memory configurations enable efficient local deployment of many open-weight models without the cost profile of traditional multi-GPU systems.

On Apple M3 Ultra, CPU and GPU share one pool of 96 GB. The GPU can address all of it, so a model that would otherwise be split across several accelerator cards can be held by a single node. Memory bandwidth is 819 GB/s (Apple specification).

For inference on open-weight models — chat, RAG, extraction, embeddings and transcription — that is a practical, predictable platform. It is not the right platform for everything, and we would rather tell you before you sign.

Know the limits

  • No NVIDIA CUDA: workloads that depend on CUDA kernels, TensorRT or CUDA-only libraries need a different architecture.
  • Model fit depends on parameter count, quantization and context length — the KV cache needs memory too.
  • Throughput depends on model, quantization, context length and concurrency, and is validated per workload.

Running CUDA workloads?

Training pipelines and anything built on CUDA kernels need a different architecture. Tell us what you run — if a Private AI Node is the wrong fit, we will say so.

Benchmarks

Measured numbers, published only when they are real.

Benchmark data coming after production validation. We will publish throughput, latency and memory figures with methodology — never estimates. Until then, a pilot benchmarks your actual workload.

Reference: Qwen3 30B-A3B · 4-bit · M3 Ultra · 96 GB

Benchmark methodology

Tokens / second

Benchmark data coming after production validation.

Generation throughput for a single stream.

Time to first token

Benchmark data coming after production validation.

Latency until the first streamed token, at a fixed prompt length.

Concurrent requests

Benchmark data coming after production validation.

Concurrent streams sustained before latency targets are exceeded.

Memory usage

Benchmark data coming after production validation.

Resident memory including KV cache at the tested context length.

Plans

Dedicated capacity at a fixed monthly price.

Start with a Dedicated Node if your team runs its own stack. Choose Managed AI Node if you want a working endpoint we operate for you. Add a second node with High Availability.

Dedicated Node

Entry dedicated compute for teams that run their own stack.

From€449/ month

Launch pricing

Deploy a Private Node
  • Dedicated Apple Silicon node
  • Up to 96 GB unified memory
  • Private, single-tenant environment
  • Private API endpoint
  • Monitoring
  • Basic managed deployment

Managed AI Node

Recommended

A dedicated, fully managed AI endpoint. Our flagship.

From€699/ month

Launch pricing

Deploy Managed AI
  • Dedicated 96 GB Apple Silicon environment
  • Model installation and configuration
  • OpenAI-compatible endpoint
  • Monitoring and alerting
  • Model and system updates
  • Secure API gateway
  • Backup configuration
  • Technical support
  • Deployment assistance

High Availability

Two-node architecture for production workloads that must stay up.

From€1,299/ month

Launch pricing

Talk to an Engineer
  • Two-node architecture
  • Load balancing across nodes
  • Redundant capacity
  • Monitoring
  • Failover design
  • Priority support
  • Custom deployment
  • Private networking

All prices exclude VAT. Annual billing: 2 months free, invoiced yearly in advance.

Full feature comparison and support levels

High Availability

When one node is not enough, add a second one.

High-availability architecture available from From €1,299 / month. Two dedicated nodes behind a load balancer, with redundant capacity and a failover design agreed with you.

  • Two dedicated nodes

    The same model deployed on both, or workloads split by role.

  • Load balancing

    Requests are distributed across healthy nodes.

  • Redundant capacity

    Maintenance and updates can be done one node at a time.

  • Private networking

    Nodes communicate over a private network.

Failover behaviour — what happens to in-flight requests, how capacity is reserved, how quickly traffic shifts — is designed and tested with you during deployment rather than promised in a brochure. Production SLA options are available for eligible managed deployments.

Two-node reference design

  1. Your application
  2. Secure gateway
  3. Load balancer
  4. Node A · Node B

Security & data location

Clear about where your node runs and who can touch it.

Designed to support GDPR-conscious deployments. Contractual and technical safeguards are available for European customers; customers remain responsible for determining the appropriate legal basis for their workloads.

Security Practices
Infrastructure region
Georgia
Outside the EEA. All compute and model storage currently run here.
Commercial operations
Germany / Europe
Sales, contracts, onboarding and customer communication.
International data transfer
Safeguards available
For workloads involving EEA personal data, appropriate contractual and technical safeguards may be required.

No training on your prompts, responses or files

Content logging configurable — or off — on your node

DPA and Standard Contractual Clauses where applicable

Read how we handle data processing · An EU region is not currently available.

Console

See your node, your models and your keys in one place.

Which model runs where, memory in use, request volume and latency, API keys and invoices. The customer console is rolling out to pilot customers; the view below is illustrative.

Illustrative console preview · synthetic demo data, not a real customer

FAQ

Private AI Node questions.

Why use a dedicated node?

A dedicated node gives you predictable capacity that no other customer shares, a private endpoint, freedom to choose and pin model versions, and a fixed monthly cost. It suits production workloads, sensitive data and teams that want consistent latency.

Can you deploy 70B models?

Some 70B-class models can run in quantized form (typically 4-bit) on a single 96 GB node. Usable context length is constrained by the memory left for the KV cache. We validate quality and performance on your workload before committing.

Can I bring or run my own model?

On Private AI Nodes, yes — including fine-tuned variants of supported architectures — provided the model is technically compatible with Apple Silicon inference runtimes, fits in memory with your required context length, and you have the rights to deploy it. We check compatibility before deployment.

Can I connect over VPN?

Private connectivity (for example a site-to-site VPN or WireGuard tunnel) is available on eligible Private AI Node and High Availability deployments. Contact us for deployment requirements.

Do you support static IP allowlisting?

Yes, on dedicated deployments you can restrict your endpoint to a list of source IP addresses. Combined with API keys and TLS, this is a common setup for server-to-server integrations.

Do you support CUDA workloads?

No. Our current infrastructure is Apple Silicon, which does not run NVIDIA CUDA. Workloads that depend on CUDA kernels, TensorRT or CUDA-only libraries — such as many training pipelines — need a different architecture, and we will tell you honestly if that applies to you.

Why Apple Silicon?

Large unified-memory configurations enable efficient local deployment of many open-weight models without the cost profile of traditional multi-GPU systems. Unified memory means the GPU can address most of the 96 GB, which lets a single node hold models that would otherwise require several GPUs. The right hardware always depends on the workload, and we will recommend accordingly.

Can I scale to multiple nodes?

Yes. You can add dedicated nodes behind a load balancer, separate workloads per node (for example chat on one, embeddings and speech on another), or move to a High Availability architecture. Capacity is finite, so larger expansions are planned with you in advance.

Private AI Pilot

Not sure yet? Test your workload first.

Validate your workload on dedicated infrastructure before committing to a larger deployment.

Pricing: Contact sales · scoped to your workload

  • Model deployment on dedicated capacity
  • Private OpenAI-compatible endpoint
  • Benchmarking on your real workload
  • Architecture recommendation
  • Migration plan

A private AI endpoint, running on hardware that is yours alone.

Managed AI Node from €699 / month. Tell us the model and the workload; an engineer confirms fit, context length and connectivity before you commit.