Qwen3-ASR-1.7B-hf
Speech / ASRQwen · Qwen · v1.7 · released
About
Qwen3-ASR 1.7B is Alibaba's open-weight automatic-speech-recognition model — a 2.04B-parameter encoder-decoder transformer that transcribes audio to text, runnable locally via HuggingFace transformers.
Architecture
- Type
- encoder-decoder
- Parameters
- 2.0B
Qwen3-ASR automatic-speech-recognition model (Qwen3ASRForConditionalGeneration), 2.04B parameters. Encoder-decoder transformer mapping audio input to text transcription. Released open-weight on HuggingFace via the transformers library.
Memory
- Weights (BF16)
- 4.08 GB
- Activation estimate
- 1.20 GB
Pricing
Free — open weights
Self-host on your own GPU. The calculator surfaces GPU-hours cost on the hardware page instead of an API price.
Provenance
- Source
- huggingface.co
- Hugging Face
- Qwen/Qwen3-ASR-1.7B-hf
- Last verified
- 2026-07-10
speech-recognitionasraudioopen-weighttransformers