Universal_Audio_Tokenizer
Speech / ASRtencent · tencent · released
About
Tencent's Universal Audio Tokenizer is an open-weight audio tokenizer that bridges acoustic and semantic understanding for Audio-LLMs. It encodes speech, sound and music into discrete tokens (25 Hz, 8192 codebook, ~325 bps) for LLM integration, TTS and audio perception.
Architecture
- Type
- encoder
Compact audio tokenizer that unifies general audio perception with linguistic alignment for Audio-LLMs. Components: a WhisperVQ encoder plus an AudioDecoder with flow + HIFT modules, trained with Semantic-Acoustic Primitives supervision and a Semantic-Acoustic Equilibrium mechanism. 25 Hz frame rate, 8192-entry codebook, ~325 bits/sec. Parameter count not published.
Pricing
Free — open weights
Self-host on your own GPU. The calculator surfaces GPU-hours cost on the hardware page instead of an API price.
Provenance
- Source
- huggingface.co
- License
- other
- Hugging Face
- tencent/Universal_Audio_Tokenizer
- Last verified
- 2026-06-23