Benchmarks

Measured performance

Per-model latency, time-to-first-token, throughput, and availability from the public SLO and host telemetry. No account required. We do not fingerprint benchmark traffic.

Not enough samples yet
10 Catalog models
0 7d latency samples
Sparse Same numbers as /v1/slo

Models

Latency, TTFT, throughput

Completion latency and 7-day success come from /v1/slo. TTFT and tokens/s are the nearest-rank percentiles of online-host EWMAs, split by trust tier. Empty cells are not enough samples yet.

No hosts or samples yet: flux-schnell, gemma-3-12b-it, gpt-oss-20b, llama-3.1-8b-instruct, llama-3.2-1b-instruct, neuronpool-tiny-chat, nomic-embed-text, qwen2.5-7b-instruct, qwen3-30b-a3b, whisper-small.

Honest numbers

What these figures are (and are not)

Methodology

Reproducible from public endpoints

Artificial Analysis lists providers at its discretion. This page is our own measurement, not theirs.

  • Success rate = completed ÷ (completed + failed) over 7 days. Unsettled jobs are excluded.
  • Latency p50/p95 = nearest-rank percentile of completed_at − created_at on completed jobs.
  • TTFT / TPS = nearest-rank percentile of per-host EWMAs (α = 0.3) among online hosts advertising that model at that trust tier.
  • Host count is point-in-time coverage (heartbeat ≤ 45 s), matching /v1/stats provider counts.