Mistral-Nemo-Instruct-2407
Speed analysis
Latency measured across all benchmark runs. P50 (median) and P95 (95th percentile) give a realistic picture of response speed under normal and peak load.
Quality scores
Evaluation results from judge-model scoring across diverse task categories. Scores reflect coherence, accuracy and instruction-following.
Pricing history
Direct provider rates per million tokens, plus a typical-conversation cost estimate.
Pricing over time
Input & output per 1M tokens · step-line = price changes
$0.1300
input / 1M
— stable
$0.1300
output / 1M
— stable
Tokens per second
Throughput in tokens per second, derived from measured P50 latency. Higher is better; fluctuations track provider-side load.
Estimated from P50 latency × 200 output tokens — the absolute number depends on this assumption; the trend is what matters.
Capabilities
Availability
Availability
No measurements yet
We haven't recorded enough API calls to show availability stats for this model. Data appears once the model starts receiving live traffic.
Tokonomix benchmark verdicts
Mistral-Nemo sees 63% latency boost, quality dips 0.9 points to 84.9
Mistral-Nemo-Instruct-2407 delivers a substantial performance improvement with latency dropping 63% from 8312ms to 3051ms at the median, making responses significantly faster for users. Overall quality remains stable at 84.9, down just 0.9 points from the previous window's 85.8. The category performance reveals a mixed picture: multilingual capabilities stay strong at 97, up slightly from 95, while coding performance dropped notably from 99 to 83. Creative writing enters the benchmark at 75, replacing reasoning which previously scored 63. The dramatic latency improvement suggests infrastructure optimizations or model serving changes that successfully addressed the previous window's performance degradation. Users can expect much faster responses while maintaining solid overall quality, though those relying heavily on coding tasks may notice reduced accuracy compared to the previous period. The model continues to excel at multilingual tasks and maintains reasonable performance across other categories.
Quality
84.9
Latency p50
3,051 ms
Test runs
5
Mistral-Nemo-Instruct-2407
by OVH AI Endpoints (GRA)
- Context window
- — tokens
- Input price
- $0.1300 / 1M
- Output price
- $0.1300 / 1M
- Tier
- —
- Modality
- Text
- API type
- REST · streaming
- Benchmark runs
- 263
More from OVH AI Endpoints (GRA)