Qwen 3 0.6B
Alibaba · dense · 0.6B parameters · 131,072 context
Parameters
0.6B
Context Window
128K tokens
Architecture
Dense
Best GPU
B200 SXM
Intelligence Brief
Qwen 3 0.6B is a 0.6B parameter DENSE model from Alibaba, featuring Grouped Query Attention (GQA) with 28 layers and 1,024 hidden dimensions. With a 131,072 token context window, it supports structured output, code, multilingual. For self-hosted inference, B200 SXM delivers optimal throughput at $4261/month.
Recent changes
Loading…
Related models
5 suggestions
Qwen 3 1.7BQwen 3 · 1.7B—
Qwen 3 235BQwen 3 · 22Bfree/M out
Qwen 3 30B-A3BQwen 3 · 3.3B$0.450/M out
Qwen 3 32BQwen 3 · 32.8Bfree/M out
Qwen 3 4BQwen 3 · 4B$0.100/M out
Picks: same family first, then same vendor within ±2× params, then top tag-overlap matches. Price shown is the cheapest Output $/M across providers — the row's page shows the canonical anchor.
Architecture Details
Memory Requirements
BF16 Weights
1.2 GB
FP8 Weights
0.6 GB
INT4 Weights
0.3 GB
GPU Compatibility Matrix
Qwen 3 0.6B is compatible with 100% of GPU configurations across 41 GPUs at 3 precision levels.
GPU Recommendations
FP8 · 1 GPU · tensorrt-llm
83/100
score
Throughput
3.5K tok/s
Latency (ITL)
0.3ms
Est. TTFT
0ms
Cost/Month
$4261
Cost/M Tokens
$0.46
FP8 · 1 GPU · tensorrt-llm
83/100
score
Throughput
3.5K tok/s
Latency (ITL)
0.3ms
Est. TTFT
0ms
Cost/Month
$4271
Cost/M Tokens
$0.46
FP8 · 1 GPU · tensorrt-llm
83/100
score
Throughput
3.5K tok/s
Latency (ITL)
0.3ms
Est. TTFT
0ms
Cost/Month
$6169
Cost/M Tokens
$0.67
Deployment Options
API Deployment
No API pricing available
Single GPU
B200 SXM
$4261/mo
Min VRAM: 1 GB
Multi-GPU
B200 SXM
3.5K tok/s
Best available config
API Pricing Comparison
No API pricing data available for this model.
Performance Estimates
Throughput by GPU
VRAM Breakdown (B200 SXM, FP8)
Precision Impact
bf16
1.2 GB
weights/GPU
fp8
0.6 GB
weights/GPU
~3.5K tok/s
int4
0.3 GB
weights/GPU
Capabilities
Features
Supported Frameworks
Supported Precisions
Where to Deploy Qwen 3 0.6B
Self-Hosted Infrastructure
Similar Models
Qwen 3 1.7B
1.7B params · dense
Quality: 50
Parakeet CTC 0.6B
0.6B params · dense
Quality: 50
from $0.03/M
Jina Embeddings v3
0.57B params · dense
Quality: 50
from $0.01/M
BGE M3
0.568B params · dense
Quality: 50
from $0.01/M
Frequently Asked Questions
How much VRAM does Qwen 3 0.6B need for inference?
Qwen 3 0.6B requires approximately 1.2 GB of VRAM at BF16 precision, 0.6 GB at FP8, or 0.3 GB at INT4 quantization. Additional VRAM is needed for KV-cache (57344 bytes per token) and activations (~0.20 GB).
What is the best GPU for Qwen 3 0.6B?
The top recommended GPU for Qwen 3 0.6B is the B200 SXM using FP8 precision. It achieves approximately 3.5K tokens/sec at an estimated cost of $4261/month ($0.46/M tokens). Score: 83/100.
How much does Qwen 3 0.6B inference cost?
Qwen 3 0.6B inference costs vary by provider and GPU setup. Use our calculator for detailed cost estimates across all providers.