Mistral-7B-Instruct-v0.3
Speed analysis
Latency measured across all benchmark runs. P50 (median) and P95 (95th percentile) give a realistic picture of response speed under normal and peak load.
Quality scores
Evaluation results from judge-model scoring across diverse task categories. Scores reflect coherence, accuracy and instruction-following.
Pricing history
Direct provider rates per million tokens, plus a typical-conversation cost estimate.
Pricing over time
Input & output per 1M tokens · step-line = price changes
$0.1000
input / 1M
— stable
$0.1000
output / 1M
— stable
Tokens per second
Throughput in tokens per second, derived from measured P50 latency. Higher is better; fluctuations track provider-side load.
Estimated from P50 latency × 200 output tokens — the absolute number depends on this assumption; the trend is what matters.
Capabilities
Availability
Availability
No measurements yet
We haven't recorded enough API calls to show availability stats for this model. Data appears once the model starts receiving live traffic.
Tokonomix benchmark verdicts
Quality up 4.9 points, latency cut 35%, creative scores debut at 57
Mistral-7B-Instruct-v0.3 on OVH AI Endpoints shows meaningful quality improvement, climbing from 78.7 to 83.6 overall. The model delivered exceptional performance in multilingual tasks, advancing from 88 to 99, while maintaining strong coding capabilities at 95, down slightly from the previous perfect 100. A notable shift in benchmark composition introduces creative scoring at 57, replacing the previous reasoning category that had scored 48. Latency performance improved substantially, with p50 response times dropping from 8061ms to 5262ms, representing a 35% reduction that should translate to noticeably faster interactions. The model continues to excel in technical and language tasks, though the moderate creative score suggests limitations in open-ended generative scenarios. Users seeking multilingual support or coding assistance will find this deployment particularly capable, while those prioritizing creative tasks may experience more variable results. The significant latency improvements make this endpoint more practical for real-time applications. With consistent test coverage of five runs per window, these results demonstrate reliable performance trends worth considering for production deployments requiring strong multilingual and coding capabilities.
Quality
83.6
Latency p50
5,262 ms
Test runs
5
Mistral-7B-Instruct-v0.3
by OVH AI Endpoints (GRA)
- Context window
- — tokens
- Input price
- $0.1000 / 1M
- Output price
- $0.1000 / 1M
- Tier
- —
- Modality
- Text
- API type
- REST · streaming
- Benchmark runs
- 263
More from OVH AI Endpoints (GRA)