Skip to content

GPU × use-case guide · Vision-language

Is the H100 NVL 94GB (per GPU pair) a good GPU for vision-language?

The H100 NVL 94GB (per GPU pair) is a hopper NVIDIA GPU with 188 GB HBM3 (7876 GB/s), 1670 TFLOPS BF16 / 3341 TFLOPS FP8, and a 800 W TDP. Vision-language workloads care most about extra VRAM for the vision encoder and image tokens on top of the language model. Here's how the H100 NVL 94GB (per GPU pair) measures up.

VRAM
188 GB
Bandwidth
7876 GB/s
BF16
1670 TFLOPS
Cheapest
$7.49/hr

What models fit on a single H100 NVL 94GB (per GPU pair)?

Weights only, reserving ~25% of the 188 GB for KV cache, activations and fragmentation. ✓ = fits on one card.

ModelBF16FP8INT4
Llama 3.1 8B
Qwen 2.5 14B
Gemma 2 27B
Mixtral 8x7B (MoE)
Llama 3.3 70B
Qwen 2.5 72B
Llama 3.1 405B

Largest single-card fit: Llama 3.3 70B at BF16, Qwen 2.5 72B at FP8, Qwen 2.5 72B at INT4. Bigger models need tensor-parallel across 4 cards.

H100 NVL 94GB (per GPU pair) for vision-language, specifically

Vision-language is context-heavy, so the KV cache — not the weights — is what fills the 188 GB. On the H100 NVL 94GB (per GPU pair) you'll trade context length against batch size: long prompts mean fewer concurrent requests. Because it runs offline, batch aggressively to push tokens-per-dollar down. Size it precisely on the calculator.

H100 NVL 94GB (per GPU pair) pricing across providers

ProviderOn-demand $/hrReserved $/hr
coreweave$7.49$5.49
runpod$7.89

Verdict

With 188 GB, the H100 NVL 94GB (per GPU pair) is a data-center-class card that comfortably handles vision-language for models up to Llama 3.3 70B at full precision on a single card — a strong pick if your budget supports ~$7.49/hr.

See full H100 NVL 94GB (per GPU pair)specs & pricing, size your model on the calculator, or compare every GPU on the GPU list.