Infrastructure
Transparent about what runs where.
4 Apple Mac Studio nodes with Apple M3 Ultra and 96 GB of unified memory each, on a private network in Georgia. Operated and supported from Germany / Europe. Here is what we run, why, and where its limits are.
Hardware class
Apple Silicon nodes with large unified memory.
These are the real characteristics of the hardware we operate. Performance figures are published separately, and only once measured.
- Compute nodes
- 4 × Apple Mac Studio
- Chip
- Apple M3 Ultra
- Unified memory
- 96 GB per node
- Memory bandwidth
- 819 GB/s (Apple specification)
- Architecture
- Apple Silicon (arm64), unified memory shared by CPU and GPU
- Storage
- Local NVMe SSD per node, with additional external NVMe where required
- Node network
- Private high-speed local network between nodes
- Compute region
- Georgia (outside the EEA)
Why Apple Silicon
A deliberate choice for inference — not a universal one.
Large unified-memory configurations enable efficient local deployment of many open-weight models without the cost profile of traditional multi-GPU systems.
Unified memory
CPU and GPU share one pool of 96 GB. The GPU can address most of it, which lets a single node hold models that would otherwise need to be split across several GPUs.
Fits the workload
Infrastructure choice depends on the workload. For serving open-weight models to a team or a product, memory capacity per node is often the deciding constraint — that is where this hardware class is strong.
Where it is not the right fit
- No NVIDIA CUDA: workloads that depend on CUDA kernels, TensorRT or CUDA-only libraries need a different architecture.
- Model fit depends on parameter count, quantization and context length — the KV cache needs memory too.
- Throughput depends on model, quantization, context length and concurrency, and is validated per workload.
Model memory classes (weights only, excluding KV cache)
| Class | Approx. footprint | Typical models |
|---|---|---|
| XS | < 8 GB | Small embedding, reranking and speech models |
| S | 8–24 GB | 7B–14B models at 4–8-bit, or ~30B at 4-bit |
| M | 24–48 GB | ~30B models at 8-bit or 70B-class at 4-bit |
| L | 48–80 GB | Large models — uses most of a 96 GB node |
We do not make GPU comparisons.
Reference architecture
From your application to a model runtime.
Every request passes the same layers: encrypted ingress, an authenticating gateway, model-aware routing and a dedicated or shared node.
Your application
Anywhere in the world
Secure API gateway
TLS termination · OpenAI-compatible routes
Authentication & rate limits
API keys · per-key quotas · allowlists
Load balancer
Model-aware routing
AI Node 01
96 GB unified
AI Node 02
96 GB unified
AI Node 03
96 GB unified
AI Node 04
96 GB unified
NVMe storage & model repository
Local NVMe · versioned model weights
Compute region: Georgia · Apple M3 Ultra
Network
Network architecture at a high level.
This overview is intentionally high level. Detailed network documentation is shared under NDA during procurement.
TLS ingress
All public API traffic terminates on TLS at the edge. Plain HTTP is not accepted for API requests.
API gateway
A gateway in front of every model enforces API-key authentication, per-key rate limits and, on dedicated deployments, IP allowlists before a request reaches a node.
Private inter-node network
Nodes communicate over a private network that is not reachable from the internet. Model runtimes are never exposed directly.
VPN option
Eligible Private AI Node and High Availability deployments can connect over a private tunnel (for example site-to-site VPN or WireGuard) instead of the public endpoint.
Storage
Fast local storage for models, backups for configuration.
Inference requests are processed in memory. Whether any prompt or response content is stored at all is configurable on dedicated deployments.
Local NVMe
Each node keeps the weights it serves on local NVMe, so models load into unified memory without network round-trips.
External NVMe
Additional external NVMe is attached where a deployment needs more model variants or larger artefacts than local storage holds.
Model repository
Versioned model weights and quantizations are kept in a repository, so a deployment can be pinned to a version and rolled back.
Configuration backups
Gateway, runtime and deployment configuration is backed up so a node can be rebuilt to a known state. Backup of customer data, where any is stored, is agreed per deployment.
Monitoring
Watched continuously, by people who can act.
Monitoring covers hardware, runtimes and endpoints. Alerts go to engineers who operate the platform.
What we monitor
- Node health, temperature and memory pressure
- Model runtime status and restarts
- Endpoint availability, latency and error rates
- Gateway authentication failures and rate-limit events
- Storage capacity on local and external NVMe
- Alerting to on-call engineers
Current service state is published on the status page.
Deployment process
From request to live endpoint.
- 1
Request
You tell us the model, expected traffic, context length and how you want to connect.
- 2
Technical review
An engineer confirms memory fit, quantization, connectivity and your data classification.
- 3
Capacity confirmed
We confirm the node(s), the plan and the timing before you commit.
- 4
Install & configure
We install the runtime and model, configure the gateway, keys, limits and monitoring.
- 5
Validation
We test the endpoint with representative requests and, if agreed, benchmark your workload.
- 6
Live endpoint
You receive the base URL and API keys. Point your OpenAI SDK at it and ship.
Timing depends on current capacity and your requirements; we confirm it before you commit.
Regions
Where compute runs today — and what is planned.
Compute currently runs in Georgia, which is outside the European Economic Area. Commercial operations, contracts and support are run from Germany / Europe.
Georgia
Georgia· Best value- Country
- Georgia
- Inside the EEA
- No
- Status
- Available — production compute region
Current production compute region. Outside the EEA — see Data Processing for transfer safeguards.
EU region
EU region· Not yet available- Country
- Germany (planned)
- Inside the EEA
- Yes
- Status
- Not currently available
An EU-based region is being evaluated. Not currently available — contact sales to register interest.
- Infrastructure region
- Georgia
- Outside the EEA. All compute and model storage currently run here.
- Commercial operations
- Germany / Europe
- Sales, contracts, onboarding and customer communication.
- International data transfer
- Safeguards available
- For workloads involving EEA personal data, appropriate contractual and technical safeguards may be required.
What we don't publish
What we don't publish, and why.
Some details make an attacker's job easier and a buyer's decision no better. We keep them off the public website and share them under NDA with prospective customers during procurement and security reviews.
- IP addresses, network ranges and DNS internals
- Firewall rules and segmentation details
- Physical address and rack location of the facility
- Administrative access architecture and tooling
- Specific software versions on internal systems
Need them for a security questionnaire or vendor assessment? Write to [email protected].
Benchmarks
Performance numbers, measured — not estimated.
We will publish tokens / second, time to first token, concurrent requests, model load time, memory usage, power efficiency per model, with the hardware, quantization, context length and runtime version used. Benchmark data coming after production validation.
Tell us the workload. We'll tell you whether this infrastructure fits.
An engineer reviews model, context length, traffic and data classification — and says plainly if a different architecture would serve you better.