Skip to content

Benchmarks

Performance, measured — or not shown.

We publish inference performance only when we have measured it on our own hardware, together with the conditions it was measured under. Until then, this page describes exactly what we will publish and how.

Validation in progress

Benchmark data coming after production validation.

Results

Results by model.

Each row is one model, quantization and node configuration. A dash means not yet measured.

  • Qwen3 30B-A3B

    4-bit · M3 Ultra · 96 GB · Not yet measured

    Tokens / second

    —

    Time to first token

    —

    Concurrent requests

    —

    Model load time

    —

    Memory usage

    —

    Power efficiency

    —

  • Qwen3 32B

    4-bit · M3 Ultra · 96 GB · Not yet measured

    Tokens / second

    —

    Time to first token

    —

    Concurrent requests

    —

    Model load time

    —

    Memory usage

    —

    Power efficiency

    —

  • Llama 3.3 70B Instruct

    4-bit · M3 Ultra · 96 GB · Not yet measured

    Tokens / second

    —

    Time to first token

    —

    Concurrent requests

    —

    Model load time

    —

    Memory usage

    —

    Power efficiency

    —

  • Whisper Large v3 Turbo

    FP16 · M3 Ultra · 96 GB · Not yet measured

    Tokens / second

    —

    Time to first token

    —

    Concurrent requests

    —

    Model load time

    —

    Memory usage

    —

    Power efficiency

    —

Metrics

What each metric means.

Six measurements that together describe how a model behaves on a node under load.

Tokens / secondtok/s
Generation throughput for a single stream.
Time to first tokenms
Latency until the first streamed token, at a fixed prompt length.
Concurrent requestsreq
Concurrent streams sustained before latency targets are exceeded.
Model load times
Cold load from local NVMe into unified memory.
Memory usageGB
Resident memory including KV cache at the tested context length.
Power efficiencytok/s per W
Throughput per watt measured at the wall.

Methodology

What we will publish with every result.

A number without its conditions is not a benchmark. Every published result will state:

Hardware

Node class, chip and memory configuration (Apple M3 Ultra, 96 GB).

Quantization

The exact weight format and bit width of every model tested.

Context length

Prompt and output lengths used for each measurement.

Concurrency

Number of simultaneous streams and how load was generated.

Runtime version

Inference runtime and version, plus relevant configuration.

Date

When the measurement was taken, so results can be compared over time.

Principles

  • Measured on our production hardware, not taken from third-party marketing.
  • Reported with the conditions above, so they can be reproduced and challenged.
  • No comparisons against GPUs we do not operate.
  • Updated when runtimes or models change materially — with the date shown.
  • Your workload may behave differently: prompt length, output length and concurrency all matter.

Benchmark on your workload

Generic numbers only go so far. In a pilot we deploy your chosen model on dedicated capacity and measure it with your prompts, context lengths and concurrency — before you commit to a larger deployment.

See how your model performs on your traffic.

Tell us the model, prompt sizes and expected concurrency. We will measure it on dedicated capacity and share the results with the methodology.