Mistral Small 4
Mistral AI · moe · 119.4B parameters · 1,048,576 context
Parameters
119.4B
Context Window
1024K tokens
Architecture
MoE
Best GPU
B200 SXM
Cheapest API
$0.60/M
Intelligence Brief
Mistral Small 4 is a 119.4B parameter Mixture-of-Experts (128 experts, 4 active) model from Mistral AI, featuring Multi-Head Attention (MHA) with 36 layers and 4,096 hidden dimensions. With a 1,048,576 token context window, it supports vision, multilingual. The most cost-effective API deployment is via openrouter at $0.60/M output tokens. For self-hosted inference, B200 SXM delivers optimal throughput at $4261/month.
Provider pricing
1 provider · canonical: openrouter| Provider | Input $/M | Output $/M ▲ | Notes |
|---|---|---|---|
| openroutercanonical | $0.150 | $0.600 | cheapest input · cheapest output |
Prices update via the nightly pricing cron + admin approvals at /admin/ingest-queue. The leaderboard's Input/Output cells show the canonical rate above; this table shows the full spread.
Recent changes
Loading…
Related models
5 suggestions
Mistral 7BMistral · 7.3B$0.070/M out
Mistral Medium 3Mistral · 70B$2.00/M out
Mistral Medium 3.5Mistral · 127.7B$7.50/M out
Codestral Mamba 7BCodestral · 7.3B$0.600/M out
Ministral 8BMinistral · 8B$0.100/M out
Picks: same family first, then same vendor within ±2× params, then top tag-overlap matches. Price shown is the cheapest Output $/M across providers — the row's page shows the canonical anchor.
Architecture Details
Memory Requirements
BF16 Weights
222.4 GB
FP8 Weights
111.2 GB
INT4 Weights
55.6 GB
GPU Compatibility Matrix
Mistral Small 4 is compatible with 21% of GPU configurations across 41 GPUs at 3 precision levels.
GPU Recommendations
FP8 · 1 GPU · tensorrt-llm
100/100
score
Throughput
280.0 tok/s
Latency (ITL)
3.6ms
Est. TTFT
1ms
Cost/Month
$4261
Cost/M Tokens
$5.79
FP8 · 1 GPU · tensorrt-llm
100/100
score
Throughput
280.0 tok/s
Latency (ITL)
3.6ms
Est. TTFT
1ms
Cost/Month
$4271
Cost/M Tokens
$5.80
FP8 · 1 GPU · tensorrt-llm
100/100
score
Throughput
280.0 tok/s
Latency (ITL)
3.6ms
Est. TTFT
1ms
Cost/Month
$6169
Cost/M Tokens
$8.38
Deployment Options
API Deployment
openrouter
$0.60/M
output tokens
Single GPU
B200 SXM
$4261/mo
Min VRAM: 111 GB
Multi-GPU
H100 NVL x2
280.0 tok/s
TP· $5865/mo
API Pricing Comparison
| Provider | Input $/M | Output $/M | Badges |
|---|---|---|---|
| openrouter | $0.15 | $0.60 | Cheapest |
Cost Analysis
| Provider | Input $/M | Output $/M | ~Monthly Cost |
|---|---|---|---|
| openrouterBest Value | $0.15 | $0.60 | $4 |
Cost per 1,000 Requests
Short (500 tok)
$0.20
via openrouter
Medium (2K tok)
$0.78
via openrouter
Long (8K tok)
$2.40
via openrouter
Performance Estimates
Throughput by GPU
VRAM Breakdown (B200 SXM, FP8)
Precision Impact
bf16
238.8 GB
weights/GPU
fp8
119.4 GB
weights/GPU
~280.0 tok/s
int4
59.7 GB
weights/GPU
Capabilities
Features
Supported Frameworks
Supported Precisions
Where to Deploy Mistral Small 4
Self-Hosted Infrastructure
Similar Models
Mistral Medium 3.5
127.7B params · dense
Quality: 50
from $7.50/M
Mistral Medium 3
70B params · dense
Quality: 80
from $2.00/M
Nemotron-3 Super 120B
120B params · dense
Quality: 84
from $0.45/M
Mistral Large 2411
123B params · dense
Quality: 75
from $6.00/M
Mistral Large 2
123B params · dense
Quality: 75
from $2.50/M
Frequently Asked Questions
How much VRAM does Mistral Small 4 need for inference?
Mistral Small 4 requires approximately 222.4 GB of VRAM at BF16 precision, 111.2 GB at FP8, or 55.6 GB at INT4 quantization. Additional VRAM is needed for KV-cache (589824 bytes per token) and activations (~0.00 GB).
What is the best GPU for Mistral Small 4?
The top recommended GPU for Mistral Small 4 is the B200 SXM using FP8 precision. It achieves approximately 280.0 tokens/sec at an estimated cost of $4261/month ($5.79/M tokens). Score: 100/100.
How much does Mistral Small 4 inference cost?
Mistral Small 4 API inference starts from $0.15/M input tokens and $0.60/M output tokens. Self-hosted inference costs depend on your GPU configuration — use our ROI calculator for a detailed breakdown.