Benchmarks
Performance, measured — or not shown.
We publish inference performance only when we have measured it on our own hardware, together with the conditions it was measured under. Until then, this page describes exactly what we will publish and how.
Validation in progress
Results
Results by model.
Each row is one model, quantization and node configuration. A dash means not yet measured.
Qwen3 30B-A3B
4-bit · M3 Ultra · 96 GB · Not yet measured
Tokens / second
—
Time to first token
—
Concurrent requests
—
Model load time
—
Memory usage
—
Power efficiency
—
Qwen3 32B
4-bit · M3 Ultra · 96 GB · Not yet measured
Tokens / second
—
Time to first token
—
Concurrent requests
—
Model load time
—
Memory usage
—
Power efficiency
—
Llama 3.3 70B Instruct
4-bit · M3 Ultra · 96 GB · Not yet measured
Tokens / second
—
Time to first token
—
Concurrent requests
—
Model load time
—
Memory usage
—
Power efficiency
—
Whisper Large v3 Turbo
FP16 · M3 Ultra · 96 GB · Not yet measured
Tokens / second
—
Time to first token
—
Concurrent requests
—
Model load time
—
Memory usage
—
Power efficiency
—
Metrics
What each metric means.
Six measurements that together describe how a model behaves on a node under load.
- Tokens / secondtok/s
- Generation throughput for a single stream.
- Time to first tokenms
- Latency until the first streamed token, at a fixed prompt length.
- Concurrent requestsreq
- Concurrent streams sustained before latency targets are exceeded.
- Model load times
- Cold load from local NVMe into unified memory.
- Memory usageGB
- Resident memory including KV cache at the tested context length.
- Power efficiencytok/s per W
- Throughput per watt measured at the wall.
Methodology
What we will publish with every result.
A number without its conditions is not a benchmark. Every published result will state:
Hardware
Node class, chip and memory configuration (Apple M3 Ultra, 96 GB).
Quantization
The exact weight format and bit width of every model tested.
Context length
Prompt and output lengths used for each measurement.
Concurrency
Number of simultaneous streams and how load was generated.
Runtime version
Inference runtime and version, plus relevant configuration.
Date
When the measurement was taken, so results can be compared over time.
Principles
- Measured on our production hardware, not taken from third-party marketing.
- Reported with the conditions above, so they can be reproduced and challenged.
- No comparisons against GPUs we do not operate.
- Updated when runtimes or models change materially — with the date shown.
- Your workload may behave differently: prompt length, output length and concurrency all matter.
Benchmark on your workload
Generic numbers only go so far. In a pilot we deploy your chosen model on dedicated capacity and measure it with your prompts, context lengths and concurrency — before you commit to a larger deployment.
See how your model performs on your traffic.
Tell us the model, prompt sizes and expected concurrency. We will measure it on dedicated capacity and share the results with the methodology.