Skip to content

Migrate from OpenAI

Move compatible workloads by changing three lines.

Keep the OpenAI SDK and your request code. Change the base URL, the API key and the model name — then validate behaviour on the open model before switching production traffic.

Before and after

The same SDK. A different endpoint.

For supported API patterns, this is the whole code change. Messages, streaming and response parsing stay as they are.

Before — OpenAI

client.py
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["OPENAI_API_KEY"],
)

client.chat.completions.create(
    model="gpt-4o-mini",
    messages=messages,
)

After — Lirux

client.py
from openai import OpenAI

client = OpenAI(
    base_url="https://api.lirux.ai/v1",
    api_key=os.environ["LIRUX_API_KEY"],
)

client.chat.completions.create(
    model="qwen3-30b-a3b",
    messages=messages,
)

What changes

Three values in your configuration.

01

Base URL

https://api.lirux.ai/v1

Point the SDK client at your endpoint — shared API or your dedicated node.

02

API key

LIRUX_API_KEY

Use your key from onboarding, stored as a server-side secret.

03

Model name

qwen3-30b-a3b

Choose an open-weight model ID. OpenAI model names are not aliased.

What to check

Compatible API, different models.

The request shape carries over. Model behaviour does not automatically. These are the areas where migrations need attention.

Prompt behaviour on open models

System-prompt adherence, tone, verbosity and formatting differ from GPT models. Expect to adjust prompts — usually towards more explicit instructions and examples.

Tool calling support

Tool calling is available on supported models only, and reliability varies. Test your real tool schemas, including multi-step calls and argument validation.

Context length

The served context of a deployment can be shorter than the model's published maximum. Check your longest prompts and RAG contexts against it.

Tokenizer

Token counts differ by model family. Replace tiktoken-based budgeting with the usage field, and re-check max_tokens and chunk sizes.

Evaluation

Build a set of real inputs with expected outputs before you switch. Compare quality, failure modes and latency for your workload — not on public leaderboards.

Feature coverage

Assistants, Batch, Files, fine-tuning and image generation are not currently supported. Check every OpenAI feature you use against the compatibility matrix.

Migration plan

Six steps from evaluation to production.

A pilot covers steps 3 to 5 with our engineers: model deployment on dedicated capacity, benchmarking on your workload and a migration plan.

  1. 1

    Inventory your usage

    List every place you call OpenAI: endpoints, models, parameters, tools, JSON mode, streaming and volumes.

  2. 2

    Check compatibility

    Compare that list with the compatibility matrix. Flag anything partial or unsupported and decide how to handle it.

  3. 3

    Build an evaluation set

    Collect representative prompts with expected outputs or grading criteria — including edge cases that matter to your users.

  4. 4

    Pick candidate models

    Shortlist one or two models per task using the mapping below. Run the evaluation set against each.

  5. 5

    Adapt prompts and limits

    Tune prompts, max_tokens, chunk sizes and timeouts. Re-run the evaluation until results are acceptable.

  6. 6

    Switch behind a flag

    Make base URL, key and model configuration-driven. Shift traffic gradually and keep the previous provider as a fallback during the transition.

Model mapping

Where to start evaluating.

A starting point for your evaluation, by the class of model you use today. This is not a claim of equivalence or quality parity — results depend on your prompts and data.

Model mapping guidance from current model class to open-weight models to evaluate
If you use todayStart evaluatingTypical workloads
Small, fast general model
qwen3-8bAvailable
High-volume tagging, classification, short replies and translation drafts.
General assistant model
qwen3-30b-a3bAvailable
Assistants, RAG answers and multilingual chat. A sensible default to evaluate first.
Higher-quality reasoning model
qwen3-32bPrivate Nodegpt-oss-120bOn Request
Document analysis, structured extraction and multi-step reasoning, on dedicated capacity.
Embedding model
bge-m3Available
Multilingual semantic search and RAG retrieval. Re-embed your corpus — vectors are not interchangeable between models.
Speech-to-text model
whisper-large-v3-turboAvailable
Transcription of calls, meetings and voice input via the transcription endpoint.
Model availability depends on licensing, memory requirements and deployment configuration. Status labels show where each model runs today — see the model docs or the model explorer.

Good fit, poor fit

Be clear about what moves easily.

Moves easily

  • Chat completions with plain prompts
  • Streaming responses
  • Embeddings for search and RAG
  • Speech-to-text
  • Classification and extraction

Needs more work

  • Assistants or Responses API
  • Batch and Files API
  • Heavy tool-calling agents
  • Fine-tuned OpenAI models
  • Image generation and TTS

Migrate with an engineer, not a checklist.

Start with a pilot on dedicated capacity: we deploy candidate models, benchmark them on your real workload and hand you a migration plan.