Skip to content

Model catalog

Open-weight models, deployed and managed for you.

18 models across 9 families for chat, reasoning, embeddings, vision, speech and reranking. Pick one, and we deploy it behind an OpenAI-compatible endpoint on shared or dedicated capacity.

Browse the catalog

Model availability depends on licensing, memory requirements and deployment configuration.

Showing 18 of 18 models

Reading the catalog

What the status and memory class mean.

Status tells you how a model can be deployed today. Memory class is an approximate guide to how much of a node a model needs.

Status

Available
Served on the shared Managed AI API and deployable on private nodes.
Private Node
Deployed on dedicated Private AI Nodes.
On Request
Deployable after a short compatibility check for your workload.
Coming Soon
Being validated. Not yet deployable.

Memory class

Approximate weight footprint, excluding the KV cache, relative to a 96 GB node.

XS < 8 GB
Small embedding, reranking and speech models
S 8–24 GB
7B–14B models at 4–8-bit, or ~30B at 4-bit
M 24–48 GB
~30B models at 8-bit or 70B-class at 4-bit
L 48–80 GB
Large models — uses most of a 96 GB node

The KV cache grows with context length and the number of concurrent requests, so the memory left after the weights determines how much context and concurrency a deployment can serve.

Bring your own model

Your weights, our infrastructure.

Fine-tuned or not in the catalog? We can deploy your own open-weight model on a dedicated node after a compatibility check.

Requirements

  • The architecture is supported by Apple Silicon runtimes such as MLX or llama.cpp.
  • The weights fit in memory together with the KV cache for your context length and concurrency — each node has 96 GB of unified memory.
  • You hold the rights to use the weights for your purpose — including fine-tunes and licensed models.

No NVIDIA CUDA: workloads that depend on CUDA kernels, TensorRT or CUDA-only libraries need a different architecture.

Model names are trademarks of their owners; their use here does not imply endorsement. License summaries are for orientation only and are not legal advice — review each model's full license for your use case. See also model documentation.

Not sure which model fits?

Tell us about your workload. An engineer will suggest a model, a deployment and a plan — and validate it on your data.