Skip to content
Runs in:FranceMade in:France
OVH AI Endpoints (GRA)

Mistral-7B-Instruct-v0.3

Tokonomix Editorial Team·Reviewed by Mes Kalkan··
Section 01

Speed analysis

Latency measured across all benchmark runs. P50 (median) and P95 (95th percentile) give a realistic picture of response speed under normal and peak load.

P50 latency (median)P95 latency100 runs
4845749101136271815306-2707-22ms
Section 02

Quality scores

Evaluation results from judge-model scoring across diverse task categories. Scores reflect coherence, accuracy and instruction-following.

95
Coding
99
Multilingual
57
Creative
Section 03

Pricing history

Direct provider rates per million tokens, plus a typical-conversation cost estimate.

💰
API rates — Mistral-7B-Instruct-v0.3
$0.1000 per 1M input tokens
$0.1000 per 1M output tokens
≈ <$0.0001 per typical conversation (800 tokens)
Input vs output price (per 1M tokens)
per 1M input tokens$0.1000
per 1M output tokens$0.1000

Pricing over time

Input & output per 1M tokens · step-line = price changes

$0.1000

input / 1M

— stable

$0.1000

output / 1M

— stable

2026-06-142026-07-052026-07-19
Input
Output
Price change
⟳ synced weekly
Section 04

Tokens per second

Throughput in tokens per second, derived from measured P50 latency. Higher is better; fluctuations track provider-side load.

Throughput (tokens / s)325 / avg 1412
407579

Estimated from P50 latency × 200 output tokens — the absolute number depends on this assumption; the trend is what matters.

Section 05

Capabilities

ownedBy: mistralai
Section 06

Availability

Availability

No measurements yet

We haven't recorded enough API calls to show availability stats for this model. Data appears once the model starts receiving live traffic.

Section 07

Tokonomix benchmark verdicts

⚖️
Endorsed by 1 judge
Independent LLM judges evaluated this model on our weekly intelligence tests
claude-sonnet-4-578/100 · 42 runs
25 correct8 partial9 wrong60% accuracy
2026-07-19

Quality up 4.9 points, latency cut 35%, creative scores debut at 57

Mistral-7B-Instruct-v0.3 on OVH AI Endpoints shows meaningful quality improvement, climbing from 78.7 to 83.6 overall. The model delivered exceptional performance in multilingual tasks, advancing from 88 to 99, while maintaining strong coding capabilities at 95, down slightly from the previous perfect 100. A notable shift in benchmark composition introduces creative scoring at 57, replacing the previous reasoning category that had scored 48. Latency performance improved substantially, with p50 response times dropping from 8061ms to 5262ms, representing a 35% reduction that should translate to noticeably faster interactions. The model continues to excel in technical and language tasks, though the moderate creative score suggests limitations in open-ended generative scenarios. Users seeking multilingual support or coding assistance will find this deployment particularly capable, while those prioritizing creative tasks may experience more variable results. The significant latency improvements make this endpoint more practical for real-time applications. With consistent test coverage of five runs per window, these results demonstrate reliable performance trends worth considering for production deployments requiring strong multilingual and coding capabilities.

Quality

83.6

Latency p50

5,262 ms

Test runs

5

Latency improved 35% Multilingual jumped to 99 Overall quality up 4.9 points Creative scores modest at 57
Last automated test
Jul 22, 2026 · 02:00 UTC · Speed benchmark
P50 latency
615 ms
P95 latency
1129 ms
Errors
0 / 6 runs
Last reviewed by Tokonomix Team·July 22, 2026