Gemini 2.0 Pro
Google · moe · 600B parameters · 2,000,000 context
Parameters
600B
Context Window
1953K tokens
Architecture
MoE
Best GPU
B200 NVL (pair)
Cheapest API
$4.00/M
Quality Score
88/100
Intelligence Brief
Gemini 2.0 Pro is a 600B parameter Mixture-of-Experts (16 experts, 2 active) model from Google, featuring Grouped Query Attention (GQA) with 96 layers and 12,288 hidden dimensions. With a 2,000,000 token context window, it supports tools, vision, structured output, code, math, multilingual, reasoning. On standardized benchmarks, it achieves MMLU 87, HumanEval 68, GSM8K 93. The most cost-effective API deployment is via google at $4.00/M output tokens. For self-hosted inference, B200 NVL (pair) delivers optimal throughput at $39858/month.
Provider pricing
1 provider · canonical: google| Provider | Input $/M | Output $/M ▲ | Notes |
|---|---|---|---|
| googlecanonical | $1.00 | $4.00 | cheapest input · cheapest output |
Prices update via the nightly pricing cron + admin approvals at /admin/ingest-queue. The leaderboard's Input/Output cells show the canonical rate above; this table shows the full spread.
Recent changes
Loading…
Related models
5 suggestions
Gemini 3 Pro PreviewGemini · 100B$12.00/M out
Gemini 1.5 FlashGemini · 12B$0.300/M out
Gemini 1.5 ProGemini · 40B$5.00/M out
Gemini 2.0 FlashGemini · 15B$0.400/M out
GPT-4.5 PreviewGPT · 300B$150.00/M out
Picks: same family first, then same vendor within ±2× params, then top tag-overlap matches. Price shown is the cheapest Output $/M across providers — the row's page shows the canonical anchor.
Architecture Details
Memory Requirements
BF16 Weights
1200.0 GB
FP8 Weights
600.0 GB
INT4 Weights
300.0 GB
Fits on (single GPU) — most practical first
Fits on (multi-GPU with Tensor Parallelism)
Multi-GPU configurations use Tensor Parallelism (TP) to split model layers across GPUs. Requires NVLink or NVSwitch interconnect for optimal performance.
GPU Compatibility Matrix
Gemini 2.0 Pro is compatible with 1% of GPU configurations across 41 GPUs at 3 precision levels.
GPU Recommendations
BF16 · 4 GPUs · tensorrt-llm
68/100
score
Throughput
140.0 tok/s
Latency (ITL)
7.1ms
Est. TTFT
1ms
Cost/Month
$39858
Cost/M Tokens
$108.33
BF16 · 8 GPUs · vllm
65/100
score
Throughput
140.0 tok/s
Latency (ITL)
7.1ms
Est. TTFT
1ms
Cost/Month
$18904
Cost/M Tokens
$51.38
BF16 · 8 GPUs · tensorrt-llm
63/100
score
Throughput
140.0 tok/s
Latency (ITL)
7.1ms
Est. TTFT
1ms
Cost/Month
$34088
Cost/M Tokens
$92.65
Deployment Options
API Deployment
$4.00/M
output tokens
Single GPU
Requires multi-GPU setup (600 GB VRAM needed)
Multi-GPU
B200 NVL (pair) x4
140.0 tok/s
TP· $39858/mo
API Pricing Comparison
| Provider | Input $/M | Output $/M | Badges |
|---|---|---|---|
| $1.00 | $4.00 | Cheapest |
Cost Analysis
| Provider | Input $/M | Output $/M | ~Monthly Cost |
|---|---|---|---|
| googleBest Value | $1.00 | $4.00 | $25 |
Cost per 1,000 Requests
Short (500 tok)
$1.30
via google
Medium (2K tok)
$5.20
via google
Long (8K tok)
$16.00
via google
Performance Estimates
Throughput by GPU
VRAM Breakdown (B200 NVL (pair), BF16)
Quality Benchmarks
Capabilities
Features
Supported Frameworks
Supported Precisions
Where to Deploy Gemini 2.0 Pro
Similar Models
Gemini 3 Pro Preview
600B params · moe
Quality: 50
from $12.00/M
Grok 3
600B params · moe
Quality: 90
from $15.00/M
Megatron-Turing NLG 530B
530B params · dense
Quality: 58
DeepSeek R1
671B params · moe
Quality: 88
from $2.00/M
DeepSeek V3
671B params · moe
Quality: 81
from $0.42/M
Frequently Asked Questions
How much VRAM does Gemini 2.0 Pro need for inference?
Gemini 2.0 Pro requires approximately 1200.0 GB of VRAM at BF16 precision, 600.0 GB at FP8, or 300.0 GB at INT4 quantization. Additional VRAM is needed for KV-cache (2359296 bytes per token) and activations (~10.00 GB).
What is the best GPU for Gemini 2.0 Pro?
The top recommended GPU for Gemini 2.0 Pro is the B200 NVL (pair) (x4) using BF16 precision. It achieves approximately 140.0 tokens/sec at an estimated cost of $39858/month ($108.33/M tokens). Score: 68/100.
How much does Gemini 2.0 Pro inference cost?
Gemini 2.0 Pro API inference starts from $1.00/M input tokens and $4.00/M output tokens. Self-hosted inference costs depend on your GPU configuration — use our ROI calculator for a detailed breakdown.