Skip to content

Infrastructure

Transparent about what runs where.

4 Apple Mac Studio nodes with Apple M3 Ultra and 96 GB of unified memory each, on a private network in Georgia. Operated and supported from Germany / Europe. Here is what we run, why, and where its limits are.

Hardware class

Apple Silicon nodes with large unified memory.

These are the real characteristics of the hardware we operate. Performance figures are published separately, and only once measured.

Compute nodes
4 × Apple Mac Studio
Chip
Apple M3 Ultra
Unified memory
96 GB per node
Memory bandwidth
819 GB/s (Apple specification)
Architecture
Apple Silicon (arm64), unified memory shared by CPU and GPU
Storage
Local NVMe SSD per node, with additional external NVMe where required
Node network
Private high-speed local network between nodes
Compute region
Georgia (outside the EEA)

Why Apple Silicon

A deliberate choice for inference — not a universal one.

Large unified-memory configurations enable efficient local deployment of many open-weight models without the cost profile of traditional multi-GPU systems.

Unified memory

CPU and GPU share one pool of 96 GB. The GPU can address most of it, which lets a single node hold models that would otherwise need to be split across several GPUs.

Fits the workload

Infrastructure choice depends on the workload. For serving open-weight models to a team or a product, memory capacity per node is often the deciding constraint — that is where this hardware class is strong.

Where it is not the right fit

  • No NVIDIA CUDA: workloads that depend on CUDA kernels, TensorRT or CUDA-only libraries need a different architecture.
  • Model fit depends on parameter count, quantization and context length — the KV cache needs memory too.
  • Throughput depends on model, quantization, context length and concurrency, and is validated per workload.

Model memory classes (weights only, excluding KV cache)

ClassApprox. footprintTypical models
XS< 8 GBSmall embedding, reranking and speech models
S8–24 GB7B–14B models at 4–8-bit, or ~30B at 4-bit
M24–48 GB~30B models at 8-bit or 70B-class at 4-bit
L48–80 GBLarge models — uses most of a 96 GB node

We do not make GPU comparisons.

We do not claim this hardware is faster than any specific GPU. Workloads built on CUDA, TensorRT or other CUDA-only libraries — including most training pipelines — need a different architecture, and we will say so during the technical review.

Reference architecture

From your application to a model runtime.

Every request passes the same layers: encrypted ingress, an authenticating gateway, model-aware routing and a dedicated or shared node.

Your application

Anywhere in the world

Secure API gateway

TLS termination · OpenAI-compatible routes

Authentication & rate limits

API keys · per-key quotas · allowlists

Load balancer

Model-aware routing

AI Node 01

96 GB unified

AI Node 02

96 GB unified

AI Node 03

96 GB unified

AI Node 04

96 GB unified

NVMe storage & model repository

Local NVMe · versioned model weights

Compute region: Georgia · Apple M3 Ultra

Network

Network architecture at a high level.

This overview is intentionally high level. Detailed network documentation is shared under NDA during procurement.

TLS ingress

All public API traffic terminates on TLS at the edge. Plain HTTP is not accepted for API requests.

API gateway

A gateway in front of every model enforces API-key authentication, per-key rate limits and, on dedicated deployments, IP allowlists before a request reaches a node.

Private inter-node network

Nodes communicate over a private network that is not reachable from the internet. Model runtimes are never exposed directly.

VPN option

Eligible Private AI Node and High Availability deployments can connect over a private tunnel (for example site-to-site VPN or WireGuard) instead of the public endpoint.

Storage

Fast local storage for models, backups for configuration.

Inference requests are processed in memory. Whether any prompt or response content is stored at all is configurable on dedicated deployments.

Local NVMe

Each node keeps the weights it serves on local NVMe, so models load into unified memory without network round-trips.

External NVMe

Additional external NVMe is attached where a deployment needs more model variants or larger artefacts than local storage holds.

Model repository

Versioned model weights and quantizations are kept in a repository, so a deployment can be pinned to a version and rolled back.

Configuration backups

Gateway, runtime and deployment configuration is backed up so a node can be rebuilt to a known state. Backup of customer data, where any is stored, is agreed per deployment.

Monitoring

Watched continuously, by people who can act.

Monitoring covers hardware, runtimes and endpoints. Alerts go to engineers who operate the platform.

What we monitor

  • Node health, temperature and memory pressure
  • Model runtime status and restarts
  • Endpoint availability, latency and error rates
  • Gateway authentication failures and rate-limit events
  • Storage capacity on local and external NVMe
  • Alerting to on-call engineers

Current service state is published on the status page.

Deployment process

From request to live endpoint.

  1. 1

    Request

    You tell us the model, expected traffic, context length and how you want to connect.

  2. 2

    Technical review

    An engineer confirms memory fit, quantization, connectivity and your data classification.

  3. 3

    Capacity confirmed

    We confirm the node(s), the plan and the timing before you commit.

  4. 4

    Install & configure

    We install the runtime and model, configure the gateway, keys, limits and monitoring.

  5. 5

    Validation

    We test the endpoint with representative requests and, if agreed, benchmark your workload.

  6. 6

    Live endpoint

    You receive the base URL and API keys. Point your OpenAI SDK at it and ship.

Timing depends on current capacity and your requirements; we confirm it before you commit.

Regions

Where compute runs today — and what is planned.

Compute currently runs in Georgia, which is outside the European Economic Area. Commercial operations, contracts and support are run from Germany / Europe.

  • Georgia

    Georgia· Best value
    Country
    Georgia
    Inside the EEA
    No
    Status
    Available — production compute region

    Current production compute region. Outside the EEA — see Data Processing for transfer safeguards.

  • EU region

    EU region· Not yet available
    Country
    Germany (planned)
    Inside the EEA
    Yes
    Status
    Not currently available

    An EU-based region is being evaluated. Not currently available — contact sales to register interest.

Infrastructure region
Georgia
Outside the EEA. All compute and model storage currently run here.
Commercial operations
Germany / Europe
Sales, contracts, onboarding and customer communication.
International data transfer
Safeguards available
For workloads involving EEA personal data, appropriate contractual and technical safeguards may be required.
EU region: not currently available. Read how we handle data processing — and speak with us about your data classification before deployment.

What we don't publish

What we don't publish, and why.

Some details make an attacker's job easier and a buyer's decision no better. We keep them off the public website and share them under NDA with prospective customers during procurement and security reviews.

  • IP addresses, network ranges and DNS internals
  • Firewall rules and segmentation details
  • Physical address and rack location of the facility
  • Administrative access architecture and tooling
  • Specific software versions on internal systems

Need them for a security questionnaire or vendor assessment? Write to [email protected].

Benchmarks

Performance numbers, measured — not estimated.

We will publish tokens / second, time to first token, concurrent requests, model load time, memory usage, power efficiency per model, with the hardware, quantization, context length and runtime version used. Benchmark data coming after production validation.

Tell us the workload. We'll tell you whether this infrastructure fits.

An engineer reviews model, context length, traffic and data classification — and says plainly if a different architecture would serve you better.