Benchmarks
Speed test
P50 = median response time for a standard 500-token output. Measured from EU (Amsterdam). P95 = tail latency — 95% of requests complete within this time. Three runs per model per test cycle; values are medians across cycles.
Tier S< 200 ms
Tier A< 500 ms
Tier B< 1000 ms
Tier C> 1000 ms
P50 (median)P95 (tail)
01FLUX.1 SchnellHugging Face (nscale)
Tier S0 ms
P95: 0 ms
02SDXL 1.0Hugging Face (nscale)
Tier S0 ms
P95: 0 ms
03FLUX.1 Kontext [max] — Multi-Image Fusionfal.ai
Tier S0 ms
P95: 0 ms
04FLUX.1 Kontext [pro] — Multi-Image Fusionfal.ai
Tier S0 ms
P95: 0 ms
05NVIDIA Nemotron Super 49B v1.5OpenRouter
Tier S15 ms
P95: 18 ms
06Mistral-Nemo-Instruct-2407OVH AI Endpoints (GRA)
Tier S99 ms
P95: 104 ms
07Meta-Llama-3_3-70B-InstructOVH AI Endpoints (GRA)
Tier S116 ms
P95: 120 ms
08Mistral-7B-Instruct-v0.3OVH AI Endpoints (GRA)
Tier S143 ms
P95: 154 ms
09Qwen2.5-VL-72B-InstructOVH AI Endpoints (GRA)
Tier S157 ms
P95: 581 ms
10Llama 4 MaverickOpenRouter
Tier S170 ms
P95: 171 ms
How we measure: Each model receives an identical prompt targeting a ~500-token output. We run 3 sequential calls per test cycle and compute P50/P95 across the distribution. Tests run 4× per day from a single EU endpoint. Network overhead is included.