Skip to content
Runs in:FranceMade in:China
Tokonomix Editorial Team·Reviewed by Mes Kalkan··
Section 01

Speed analysis

Latency measured across all benchmark runs. P50 (median) and P95 (95th percentile) give a realistic picture of response speed under normal and peak load.

P50 latency (median)P95 latency100 runs
401817615951237253150006-2707-22ms
Section 02

Quality scores

Evaluation results from judge-model scoring across diverse task categories. Scores reflect coherence, accuracy and instruction-following.

96
Coding
91
Multilingual
45
Creative
Section 03

Pricing history

Direct provider rates per million tokens, plus a typical-conversation cost estimate.

💰
API rates — Qwen3.5-9B
$0.1200 per 1M input tokens
$0.1800 per 1M output tokens
≈ $0.0001 per typical conversation (800 tokens)
Input vs output price (per 1M tokens)
per 1M input tokens$0.1200
per 1M output tokens$0.1800

Pricing over time

Input & output per 1M tokens · step-line = price changes

$0.1200

input / 1M

— stable

$0.1800

output / 1M

— stable

2026-06-142026-07-052026-07-19
Input
Output
Price change
⟳ synced weekly
Section 04

Tokens per second

Throughput in tokens per second, derived from measured P50 latency. Higher is better; fluctuations track provider-side load.

Throughput (tokens / s)384 / avg 386
4945

Estimated from P50 latency × 200 output tokens — the absolute number depends on this assumption; the trend is what matters.

Section 05

Capabilities

ownedBy: Qwen
Section 06

Availability

Availability

No measurements yet

We haven't recorded enough API calls to show availability stats for this model. Data appears once the model starts receiving live traffic.

Section 07

Tokonomix benchmark verdicts

⚖️
Endorsed by 1 judge
Independent LLM judges evaluated this model on our weekly intelligence tests
claude-sonnet-4-541/100 · 41 runs
14 correct2 partial25 wrong34% accuracy
2026-07-19

Qwen3.5-9B rebounds sharply with 77.3 score, 46-point improvement

Qwen3.5-9B has demonstrated a dramatic recovery in the current benchmark window, climbing from a concerning 31.0 overall quality score to a respectable 77.3, representing a substantial 46.3-point improvement. This marks a significant turnaround from the previous period's quality decline. The model continues to excel in technical domains, maintaining exceptional coding performance at 96 and achieving strong multilingual capabilities at 91, a remarkable recovery from the previous window's zero score in that category. However, creative tasks remain a notable weakness at just 45, indicating uneven capability distribution across different task types. Latency has also improved meaningfully, with the median response time decreasing by 25 percent from 20,724ms to 15,522ms, though absolute response times remain relatively high at over 15 seconds. The current assessment is based on five test runs compared to four in the previous window. Users should expect reliable performance for coding and multilingual applications but may encounter limitations with creative writing or content generation tasks. The dramatic improvement suggests previous issues may have been addressed through model updates or infrastructure changes.

Quality

77.3

Latency p50

15,522 ms

Test runs

5

Quality surged 46 points Latency improved 25% Multilingual restored to 91 Creative tasks weak at 45
Last automated test
Jul 22, 2026 · 02:00 UTC · Speed benchmark
P50 latency
521 ms
P95 latency
541 ms
Errors
0 / 6 runs
Last reviewed by Tokonomix Team·July 22, 2026