Model catalog
Open-weight models, deployed and managed for you.
18 models across 9 families for chat, reasoning, embeddings, vision, speech and reranking. Pick one, and we deploy it behind an OpenAI-compatible endpoint on shared or dedicated capacity.
Browse the catalog
Showing 18 of 18 models
Reading the catalog
What the status and memory class mean.
Status tells you how a model can be deployed today. Memory class is an approximate guide to how much of a node a model needs.
Status
- Available
- Served on the shared Managed AI API and deployable on private nodes.
- Private Node
- Deployed on dedicated Private AI Nodes.
- On Request
- Deployable after a short compatibility check for your workload.
- Coming Soon
- Being validated. Not yet deployable.
Memory class
Approximate weight footprint, excluding the KV cache, relative to a 96 GB node.
- XS < 8 GB
- Small embedding, reranking and speech models
- S 8–24 GB
- 7B–14B models at 4–8-bit, or ~30B at 4-bit
- M 24–48 GB
- ~30B models at 8-bit or 70B-class at 4-bit
- L 48–80 GB
- Large models — uses most of a 96 GB node
The KV cache grows with context length and the number of concurrent requests, so the memory left after the weights determines how much context and concurrency a deployment can serve.
Bring your own model
Your weights, our infrastructure.
Fine-tuned or not in the catalog? We can deploy your own open-weight model on a dedicated node after a compatibility check.
Requirements
- The architecture is supported by Apple Silicon runtimes such as MLX or llama.cpp.
- The weights fit in memory together with the KV cache for your context length and concurrency — each node has 96 GB of unified memory.
- You hold the rights to use the weights for your purpose — including fine-tunes and licensed models.
No NVIDIA CUDA: workloads that depend on CUDA kernels, TensorRT or CUDA-only libraries need a different architecture.
Model names are trademarks of their owners; their use here does not imply endorsement. License summaries are for orientation only and are not legal advice — review each model's full license for your use case. See also model documentation.
Not sure which model fits?
Tell us about your workload. An engineer will suggest a model, a deployment and a plan — and validate it on your data.