Reference
Models
Every request names a model by ID in the model field. This page explains the ID scheme, what each status means, and how versions and quantization are handled.
A model ID is a lowercase, stable string derived from the upstream model name — family, size and variant — for example qwen3-30b-a3b. Pass it in the model field of any request:
{ "model": "qwen3-30b-a3b", "messages": [ ... ] }- IDs do not encode quantization or runtime. Those are deployment details, documented per deployment and pinned on dedicated nodes.
GET /v1/modelsis the source of truth for what your key can call. The catalog below lists what can be deployed.- Calling a model that exists but is not enabled for your key returns
403 model_not_in_plan; an unknown ID returns404 model_not_found.
Shared means served on the multi-tenant Managed AI API. Dedicated means deployable on a single-tenant Private AI Node. Context is the model's published maximum; the served context on a deployment can be lower, because the KV cache shares memory with the weights.
| Model ID | Categories | Context | Shared | Dedicated | Status |
|---|---|---|---|---|---|
qwen3-30b-a3b30B total · 3B active (MoE) | Chat, Reasoning | 32K native (longer with RoPE scaling) | Shared: yes | Dedicated: yes | Available |
qwen3-8b8B (dense) | Chat | 32K native (longer with RoPE scaling) | Shared: yes | Dedicated: yes | Available |
bge-m3568M | Embedding | 8K | Shared: yes | Dedicated: yes | Available |
bge-reranker-v2-m3568M | Reranking | 8K | Shared: yes | Dedicated: yes | Available |
whisper-large-v3-turbo809M | Speech | Audio (30 s windows, chunked) | Shared: yes | Dedicated: yes | Available |
qwen3-32b32B (dense) | Chat, Reasoning | 32K native (longer with RoPE scaling) | Shared: no | Dedicated: yes | Private Node |
qwen3-14b14B (dense) | Chat | 32K native (longer with RoPE scaling) | Shared: no | Dedicated: yes | Private Node |
llama-3.3-70b-instruct70B (dense) | Chat | 128K | Shared: no | Dedicated: yes | Private Node |
gemma-3-27b-it27B (dense) | Chat, Vision | 128K | Shared: no | Dedicated: yes | Private Node |
mistral-small-3.2-24b-instruct24B (dense) | Chat, Vision | 128K | Shared: no | Dedicated: yes | Private Node |
whisper-large-v31.55B | Speech | Audio (30 s windows, chunked) | Shared: no | Dedicated: yes | Private Node |
llama-3.1-8b-instruct8B (dense) | Chat | 128K | Shared: no | Dedicated: yes | On Request |
gpt-oss-120b117B total · ~5B active (MoE) | Reasoning, Chat | 128K | Shared: no | Dedicated: yes | On Request |
deepseek-r1-distill-qwen-32b32B (dense) | Reasoning | 128K | Shared: no | Dedicated: yes | On Request |
qwen2.5-vl-7b-instruct7B (dense) | Vision | 32K | Shared: no | Dedicated: yes | On Request |
multilingual-e5-large-instruct560M | Embedding | 512 tokens | Shared: no | Dedicated: yes | On Request |
gpt-oss-20b21B total · ~3.6B active (MoE) | Reasoning, Chat | 128K | Shared: yes | Dedicated: yes | Coming Soon |
qwen3-embedding-8b8B | Embedding | 32K | Shared: no | Dedicated: yes | Coming Soon |
| Status | Meaning |
|---|---|
| Available | Served on the shared Managed AI API and deployable on private nodes. |
| Private Node | Deployed on dedicated Private AI Nodes. |
| On Request | Deployable after a short compatibility check for your workload. |
| Coming Soon | Being validated. Not yet deployable. |
Only models marked Available on the shared API appear in /v1/models for shared plans. Private Node and On Request models are deployed on dedicated capacity after a short fit check for your workload.
Open-weight models receive upstream revisions, and quantized builds can be regenerated. Both can change outputs, so production applications need control over when that happens.
- Dedicated nodes: the deployed build — upstream revision, quantization and runtime version — is pinned. It changes only when you agree to an update, ideally after you have re-run your evaluation set against the new build. Aliases (for example
support-assistantpointing to a pinned build) can be configured on request so you can switch builds without changing application code. - Shared API: a model ID refers to a maintained deployment of that model. Changes that can affect outputs are communicated to customers before they take effect. If you need a frozen build, use a dedicated node.
Quantization stores model weights at lower numerical precision — for example 4-bit or 8-bit instead of 16-bit. It reduces the memory a model needs, which determines what fits on a node with 96 GB of unified memory and how much room is left for context.
- 8-bit is generally close to full precision in output quality and needs roughly half the memory of 16-bit.
- 4-bit makes larger models (such as 70B-class) fit on a single node. Quality impact depends on the model and task, so evaluate on your own data.
- Native formats such as MXFP4 for gpt-oss are the format the model is distributed in, not a later compression step.
The quantization of each deployment is documented and, on dedicated nodes, chosen with you. Approximate weight footprints by memory class:
| Class | Weights | Typical models |
|---|---|---|
XS | < 8 GB | Small embedding, reranking and speech models |
S | 8–24 GB | 7B–14B models at 4–8-bit, or ~30B at 4-bit |
M | 24–48 GB | ~30B models at 8-bit or 70B-class at 4-bit |
L | 48–80 GB | Large models — uses most of a 96 GB node |
Weight footprints exclude the KV cache, which grows with context length and concurrent requests.