Skip to content

Reference

Models

Every request names a model by ID in the model field. This page explains the ID scheme, what each status means, and how versions and quantization are handled.

A model ID is a lowercase, stable string derived from the upstream model name — family, size and variant — for example qwen3-30b-a3b. Pass it in the model field of any request:

request body
{ "model": "qwen3-30b-a3b", "messages": [ ... ] }
  • IDs do not encode quantization or runtime. Those are deployment details, documented per deployment and pinned on dedicated nodes.
  • GET /v1/models is the source of truth for what your key can call. The catalog below lists what can be deployed.
  • Calling a model that exists but is not enabled for your key returns 403 model_not_in_plan; an unknown ID returns 404 model_not_found.

Shared means served on the multi-tenant Managed AI API. Dedicated means deployable on a single-tenant Private AI Node. Context is the model's published maximum; the served context on a deployment can be lower, because the KV cache shares memory with the weights.

Model catalog with categories, context window, availability and status
Model IDCategoriesContextSharedDedicatedStatus
qwen3-30b-a3b30B total · 3B active (MoE)Chat, Reasoning32K native (longer with RoPE scaling)Shared: yesDedicated: yesAvailable
qwen3-8b8B (dense)Chat32K native (longer with RoPE scaling)Shared: yesDedicated: yesAvailable
bge-m3568MEmbedding8KShared: yesDedicated: yesAvailable
bge-reranker-v2-m3568MReranking8KShared: yesDedicated: yesAvailable
whisper-large-v3-turbo809MSpeechAudio (30 s windows, chunked)Shared: yesDedicated: yesAvailable
qwen3-32b32B (dense)Chat, Reasoning32K native (longer with RoPE scaling)Shared: noDedicated: yesPrivate Node
qwen3-14b14B (dense)Chat32K native (longer with RoPE scaling)Shared: noDedicated: yesPrivate Node
llama-3.3-70b-instruct70B (dense)Chat128KShared: noDedicated: yesPrivate Node
gemma-3-27b-it27B (dense)Chat, Vision128KShared: noDedicated: yesPrivate Node
mistral-small-3.2-24b-instruct24B (dense)Chat, Vision128KShared: noDedicated: yesPrivate Node
whisper-large-v31.55BSpeechAudio (30 s windows, chunked)Shared: noDedicated: yesPrivate Node
llama-3.1-8b-instruct8B (dense)Chat128KShared: noDedicated: yesOn Request
gpt-oss-120b117B total · ~5B active (MoE)Reasoning, Chat128KShared: noDedicated: yesOn Request
deepseek-r1-distill-qwen-32b32B (dense)Reasoning128KShared: noDedicated: yesOn Request
qwen2.5-vl-7b-instruct7B (dense)Vision32KShared: noDedicated: yesOn Request
multilingual-e5-large-instruct560MEmbedding512 tokensShared: noDedicated: yesOn Request
gpt-oss-20b21B total · ~3.6B active (MoE)Reasoning, Chat128KShared: yesDedicated: yesComing Soon
qwen3-embedding-8b8BEmbedding32KShared: noDedicated: yesComing Soon
Model availability depends on licensing, memory requirements and deployment configuration. For details, licenses and use cases per model, open the model explorer.
Model status values
StatusMeaning
AvailableServed on the shared Managed AI API and deployable on private nodes.
Private NodeDeployed on dedicated Private AI Nodes.
On RequestDeployable after a short compatibility check for your workload.
Coming SoonBeing validated. Not yet deployable.

Only models marked Available on the shared API appear in /v1/models for shared plans. Private Node and On Request models are deployed on dedicated capacity after a short fit check for your workload.

Open-weight models receive upstream revisions, and quantized builds can be regenerated. Both can change outputs, so production applications need control over when that happens.

  • Dedicated nodes: the deployed build — upstream revision, quantization and runtime version — is pinned. It changes only when you agree to an update, ideally after you have re-run your evaluation set against the new build. Aliases (for example support-assistant pointing to a pinned build) can be configured on request so you can switch builds without changing application code.
  • Shared API: a model ID refers to a maintained deployment of that model. Changes that can affect outputs are communicated to customers before they take effect. If you need a frozen build, use a dedicated node.

Quantization stores model weights at lower numerical precision — for example 4-bit or 8-bit instead of 16-bit. It reduces the memory a model needs, which determines what fits on a node with 96 GB of unified memory and how much room is left for context.

  • 8-bit is generally close to full precision in output quality and needs roughly half the memory of 16-bit.
  • 4-bit makes larger models (such as 70B-class) fit on a single node. Quality impact depends on the model and task, so evaluate on your own data.
  • Native formats such as MXFP4 for gpt-oss are the format the model is distributed in, not a later compression step.

The quantization of each deployment is documented and, on dedicated nodes, chosen with you. Approximate weight footprints by memory class:

Memory classes
ClassWeightsTypical models
XS< 8 GBSmall embedding, reranking and speech models
S8–24 GB7B–14B models at 4–8-bit, or ~30B at 4-bit
M24–48 GB~30B models at 8-bit or 70B-class at 4-bit
L48–80 GBLarge models — uses most of a 96 GB node

Weight footprints exclude the KV cache, which grows with context length and concurrent requests.