Qwen3.5-9B
Speed analysis
Latency measured across all benchmark runs. P50 (median) and P95 (95th percentile) give a realistic picture of response speed under normal and peak load.
Quality scores
Evaluation results from judge-model scoring across diverse task categories. Scores reflect coherence, accuracy and instruction-following.
Pricing history
Direct provider rates per million tokens, plus a typical-conversation cost estimate.
Pricing over time
Input & output per 1M tokens · step-line = price changes
$0.1200
input / 1M
— stable
$0.1800
output / 1M
— stable
Tokens per second
Throughput in tokens per second, derived from measured P50 latency. Higher is better; fluctuations track provider-side load.
Estimated from P50 latency × 200 output tokens — the absolute number depends on this assumption; the trend is what matters.
Capabilities
Availability
Availability
No measurements yet
We haven't recorded enough API calls to show availability stats for this model. Data appears once the model starts receiving live traffic.
Tokonomix benchmark verdicts
Qwen3.5-9B rebounds sharply with 77.3 score, 46-point improvement
Qwen3.5-9B has demonstrated a dramatic recovery in the current benchmark window, climbing from a concerning 31.0 overall quality score to a respectable 77.3, representing a substantial 46.3-point improvement. This marks a significant turnaround from the previous period's quality decline. The model continues to excel in technical domains, maintaining exceptional coding performance at 96 and achieving strong multilingual capabilities at 91, a remarkable recovery from the previous window's zero score in that category. However, creative tasks remain a notable weakness at just 45, indicating uneven capability distribution across different task types. Latency has also improved meaningfully, with the median response time decreasing by 25 percent from 20,724ms to 15,522ms, though absolute response times remain relatively high at over 15 seconds. The current assessment is based on five test runs compared to four in the previous window. Users should expect reliable performance for coding and multilingual applications but may encounter limitations with creative writing or content generation tasks. The dramatic improvement suggests previous issues may have been addressed through model updates or infrastructure changes.
Quality
77.3
Latency p50
15,522 ms
Test runs
5
Qwen3.5-9B
by OVH AI Endpoints (GRA)
- Context window
- — tokens
- Input price
- $0.1200 / 1M
- Output price
- $0.1800 / 1M
- Tier
- —
- Modality
- Text
- API type
- REST · streaming
- Benchmark runs
- 263
More from OVH AI Endpoints (GRA)