Qwen3-ASR-0.6B-hf
Qwen · dense · 0.8B parameters · 65,536 context
Parameters
0.8B
Context Window
64K tokens
Architecture
Dense
Best GPU
RTX 4060
Intelligence Brief
Qwen3-ASR-0.6B-hf is a 0.8B parameter DENSE model from Qwen, featuring Grouped Query Attention (GQA) with 28 layers and 1,024 hidden dimensions. With a 65,536 token context window, it supports general text generation. For self-hosted inference, RTX 4060 delivers optimal throughput at $209/month.
Recent changes
Loading…
Related models
4 suggestions
Qwen-AgentWorld-35B-A3BQwen · 7B—
GLM-4.7-FlashGLM-4 · 3B—
mmE5-mllama-11b-instructintfloat · 10.6B—
siglip-so400m-14-384-flash-attn2-navitHuggingFaceM4 · 0.9B—
Picks: same family first, then same vendor within ±2× params, then top tag-overlap matches. Price shown is the cheapest Output $/M across providers — the row's page shows the canonical anchor.
Architecture Details
Memory Requirements
BF16 Weights
1.6 GB
FP8 Weights
0.8 GB
INT4 Weights
0.4 GB
GPU Compatibility Matrix
Qwen3-ASR-0.6B-hf is compatible with 100% of GPU configurations across 41 GPUs at 3 precision levels.
GPU Recommendations
BF16 · 1 GPU · vllm
90/100
score
Throughput
917.9 tok/s
Latency (ITL)
1.1ms
Est. TTFT
0ms
Cost/Month
$209
Cost/M Tokens
$0.09
BF16 · 1 GPU · vllm
90/100
score
Throughput
1.5K tok/s
Latency (ITL)
0.7ms
Est. TTFT
0ms
Cost/Month
$85
Cost/M Tokens
$0.02
FP8 · 1 GPU · tensorrt-llm
83/100
score
Throughput
3.5K tok/s
Latency (ITL)
0.3ms
Est. TTFT
0ms
Cost/Month
$4261
Cost/M Tokens
$0.46
Deployment Options
API Deployment
No API pricing available
Single GPU
RTX 4060
$209/mo
Min VRAM: 1 GB
Multi-GPU
RTX 4060
917.9 tok/s
Best available config
API Pricing Comparison
No API pricing data available for this model.
Performance Estimates
Throughput by GPU
VRAM Breakdown (RTX 4060, BF16)
Precision Impact
bf16
1.6 GB
weights/GPU
~917.9 tok/s
fp8
0.8 GB
weights/GPU
int4
0.4 GB
weights/GPU
Capabilities
Features
Supported Frameworks
Supported Precisions
Where to Deploy Qwen3-ASR-0.6B-hf
Self-Hosted Infrastructure
Similar Models
Florence 2 Large
0.77B params · dense
Quality: 50
Whisper Medium
0.769B params · dense
Quality: 50
Gemma 3 1B
1B params · dense
Quality: 35
Llama Guard 3 1B
1B params · dense
Quality: 50
Frequently Asked Questions
How much VRAM does Qwen3-ASR-0.6B-hf need for inference?
Qwen3-ASR-0.6B-hf requires approximately 1.6 GB of VRAM at BF16 precision, 0.8 GB at FP8, or 0.4 GB at INT4 quantization. Additional VRAM is needed for KV-cache (114688 bytes per token) and activations (~0.00 GB).
What is the best GPU for Qwen3-ASR-0.6B-hf?
The top recommended GPU for Qwen3-ASR-0.6B-hf is the RTX 4060 using BF16 precision. It achieves approximately 917.9 tokens/sec at an estimated cost of $209/month ($0.09/M tokens). Score: 90/100.
How much does Qwen3-ASR-0.6B-hf inference cost?
Qwen3-ASR-0.6B-hf inference costs vary by provider and GPU setup. Use our calculator for detailed cost estimates across all providers.