Skip to content

Qwen3-ASR-1.7B-hf

Speech / ASR

Qwen · Qwen · v1.7 · released

About

Qwen3-ASR 1.7B is Alibaba's open-weight automatic-speech-recognition model — a 2.04B-parameter encoder-decoder transformer that transcribes audio to text, runnable locally via HuggingFace transformers.

Architecture

Type
encoder-decoder
Parameters
2.0B

Qwen3-ASR automatic-speech-recognition model (Qwen3ASRForConditionalGeneration), 2.04B parameters. Encoder-decoder transformer mapping audio input to text transcription. Released open-weight on HuggingFace via the transformers library.

Memory

Weights (BF16)
4.08 GB
Activation estimate
1.20 GB

Pricing

Free — open weights

Self-host on your own GPU. The calculator surfaces GPU-hours cost on the hardware page instead of an API price.

Provenance

Last verified
2026-07-10
speech-recognitionasraudioopen-weighttransformers