Skip to content
Runs in:FranceMade in:France
OVH AI Endpoints (GRA)

Mistral-Nemo-Instruct-2407

Tokonomix Editorial Team·Reviewed by Mes Kalkan··
Section 01

Speed analysis

Latency measured across all benchmark runs. P50 (median) and P95 (95th percentile) give a realistic picture of response speed under normal and peak load.

P50 latency (median)P95 latency100 runs
76793215788236443150006-2707-22ms
Section 02

Quality scores

Evaluation results from judge-model scoring across diverse task categories. Scores reflect coherence, accuracy and instruction-following.

83
Coding
97
Multilingual
75
Creative
Section 03

Pricing history

Direct provider rates per million tokens, plus a typical-conversation cost estimate.

💰
API rates — Mistral-Nemo-Instruct-2407
$0.1300 per 1M input tokens
$0.1300 per 1M output tokens
≈ $0.0001 per typical conversation (800 tokens)
Input vs output price (per 1M tokens)
per 1M input tokens$0.1300
per 1M output tokens$0.1300

Pricing over time

Input & output per 1M tokens · step-line = price changes

$0.1300

input / 1M

— stable

$0.1300

output / 1M

— stable

2026-06-142026-06-282026-07-19
Input
Output
Price change
⟳ synced weekly
Section 04

Tokens per second

Throughput in tokens per second, derived from measured P50 latency. Higher is better; fluctuations track provider-side load.

Throughput (tokens / s)2151 / avg 1893
25895

Estimated from P50 latency × 200 output tokens — the absolute number depends on this assumption; the trend is what matters.

Section 05

Capabilities

ownedBy: mistralai
Section 06

Availability

Availability

No measurements yet

We haven't recorded enough API calls to show availability stats for this model. Data appears once the model starts receiving live traffic.

Section 07

Tokonomix benchmark verdicts

⚖️
Endorsed by 1 judge
Independent LLM judges evaluated this model on our weekly intelligence tests
claude-sonnet-4-582/100 · 42 runs
30 correct5 partial7 wrong71% accuracy
2026-07-19

Mistral-Nemo sees 63% latency boost, quality dips 0.9 points to 84.9

Mistral-Nemo-Instruct-2407 delivers a substantial performance improvement with latency dropping 63% from 8312ms to 3051ms at the median, making responses significantly faster for users. Overall quality remains stable at 84.9, down just 0.9 points from the previous window's 85.8. The category performance reveals a mixed picture: multilingual capabilities stay strong at 97, up slightly from 95, while coding performance dropped notably from 99 to 83. Creative writing enters the benchmark at 75, replacing reasoning which previously scored 63. The dramatic latency improvement suggests infrastructure optimizations or model serving changes that successfully addressed the previous window's performance degradation. Users can expect much faster responses while maintaining solid overall quality, though those relying heavily on coding tasks may notice reduced accuracy compared to the previous period. The model continues to excel at multilingual tasks and maintains reasonable performance across other categories.

Quality

84.9

Latency p50

3,051 ms

Test runs

5

63% faster response times Multilingual stays strong at 97 Coding drops from 99 to 83 Quality dips 0.9 points
Last automated test
Jul 22, 2026 · 02:00 UTC · Speed benchmark
P50 latency
93 ms
P95 latency
158 ms
Errors
0 / 6 runs
Last reviewed by Tokonomix Team·July 22, 2026