Skip to content

Architecture recommendation

Tell us the workload. We'll propose the architecture.

Describe what you are building. An engineer reviews it and sends back a concrete recommendation — model, quantization, node count, connectivity and an indicative price. No obligation, no automated quote.

What you get back

A recommendation an engineer would sign off on.

Proposed architecture

Shared API, dedicated node or multi-node setup — with the reasoning, not just the answer.

Model & quantization

Which model fits your quality needs, at which quantization, and how much context length that leaves in memory.

Node count

How many 96 GB nodes your volume and concurrency need, and whether workloads should be split across nodes.

Connectivity

Public HTTPS with API keys, IP allowlisting or private connectivity over VPN on eligible plans.

Data processing considerations

What processing in Georgia means for your data classification, logging and retention settings.

Indicative monthly price

A monthly figure based on current plans, so you can budget before a technical call.

Your workload

Describe your project.

The more you share, the more specific the recommendation. If you are unsure about a question, choose “Not sure” — that is a useful answer too.

What we look at

  • Request volume and peak concurrency
  • Prompt and document length
  • Latency expectations
  • Personal or special-category data
  • High-availability requirements
  • Integration and network constraints

Honest about fit

Our infrastructure is Apple Silicon. Workloads that depend on NVIDIA CUDA need a different architecture — if that applies to you, we will tell you.
Do you need a dedicated node?
Do you process personal data?
Preferred region
Need VPN / private connectivity?
Need custom model deployment?
Need high availability?