Skip to content

Universal_Audio_Tokenizer

Speech / ASR

tencent · tencent · released

About

Tencent's Universal Audio Tokenizer is an open-weight audio tokenizer that bridges acoustic and semantic understanding for Audio-LLMs. It encodes speech, sound and music into discrete tokens (25 Hz, 8192 codebook, ~325 bps) for LLM integration, TTS and audio perception.

Architecture

Type
encoder

Compact audio tokenizer that unifies general audio perception with linguistic alignment for Audio-LLMs. Components: a WhisperVQ encoder plus an AudioDecoder with flow + HIFT modules, trained with Semantic-Acoustic Primitives supervision and a Semantic-Acoustic Equilibrium mechanism. 25 Hz frame rate, 8192-entry codebook, ~325 bits/sec. Parameter count not published.

Pricing

Free — open weights

Self-host on your own GPU. The calculator surfaces GPU-hours cost on the hardware page instead of an API price.

Provenance

License
other
Last verified
2026-06-23
audiotokenizeraudio-llmspeechopen-weight