Benchmarks
Measured performance
Per-model latency, time-to-first-token, throughput, and availability from the public SLO and host telemetry. No account required. We do not fingerprint benchmark traffic.
—
Not enough samples yet
10
Catalog models
0
7d latency samples
Sparse
Same numbers as /v1/slo
Models
Latency, TTFT, throughput
Completion latency and 7-day success come from /v1/slo. TTFT and tokens/s are the nearest-rank percentiles of online-host EWMAs, split by trust tier. Empty cells are not enough samples yet.
No hosts or samples yet: flux-schnell, gemma-3-12b-it, gpt-oss-20b, llama-3.1-8b-instruct, llama-3.2-1b-instruct, neuronpool-tiny-chat, nomic-embed-text, qwen2.5-7b-instruct, qwen3-30b-a3b, whisper-small.
Honest numbers
What these figures are (and are not)
- Success rate = completed ÷ (completed + failed) over 7 days. Unsettled jobs are excluded.
- Latency p50/p95 = nearest-rank percentile of
completed_at − created_at on completed jobs.
- TTFT / TPS = nearest-rank percentile of per-host EWMAs (α = 0.3) among online hosts advertising that model at that trust tier.
- Host count is point-in-time coverage (heartbeat ≤ 45 s), matching
/v1/stats provider counts.