Skip to content

Leaderboard

Your top benchmark model rarely wins on cost

Every model ranked on quality, throughput and real dollars per million tokens, with each figure traceable to its source.

Open the leaderboard

317 models·60 GPUs·19 providers·Independently sourced, refreshed daily

Cost calculator

Compare before you deploy

Set GPUs, nodes and price, then read throughput and hourly cost before you commit a dollar. Try it right here.

GPU price (USD)
/hr
Nodes
GPUs per node
Throughput
12,480tok/s
Total cost
$160.00/hr
Utilisation
64 GPUsactive

Live rankings

Model leaderboard

Quality, throughput and real $/M tokens for every model, ranked, filterable and sourced.

Specialized models

Image, speech, embedding, video and scientific models, with their pricing basis.

ModelTypeProviderArchPricing
OtherEdify Imageimage-genNVIDIANVIDIAdiffusion$0.040 / image
NVIDIANemotron-3-Embed-1B-NVFP4image-gennvidianvidiaencoderFree · open weights
NVIDIANemotron-3-Embed-8B-BF16image-gennvidianvidiaencoderFree · open weights
OtherStyleGAN3image-genNVIDIANVIDIAganFree · open weights
OtherBioNeMo ESM-2 650MproteinNVIDIANVIDIAencoderFree · open weights
OtherRiva FastPitch (en-US)speech-ttsNVIDIANVIDIAencoder-decoderFree · open weights
OtherRiva HiFi-GAN (en-US)speech-ttsNVIDIANVIDIAganFree · open weights
OtherMaxine Eye Contactvideo-fxNVIDIANVIDIApipeline$0.000 / video min
OtherDINOv2 ViT-g/14 (NVIDIA-optimized)vision-embeddingNVIDIANVIDIAencoderFree · open weights
OtherNV-CLIPvision-embeddingNVIDIANVIDIAencoderPer embedding

Top picks

The standout model on each axis: best value, best quality, cheapest and fastest.

Market map

Every model plotted by quality against price, so the frontier and the outliers are obvious.

As of 02 AUG 2026
$2.19/Mbuys quality 88
vs
$150/Mfor the same tier
68× spread
708090100$0.03$0.30$3$30o1 · $60/M · Q 93GPT-4.5 Preview · $150/M · Q 93Grok 3 · $15/M · Q 90Claude Opus 4 · $75/M · Q 90Gemini 2.0 Pro · $4.00/M · Q 88o3-mini · $4.40/M · Q 86Nemotron Ultra 253B · $6.00/M · Q 86Claude Sonnet 4 · $15/M · Q 86Nemotron 340B · $4.20/M · Q 85GPT-4o · $10/M · Q 85Llama 4 Maverick · $1.80/M · Q 84Nemotron-3 Super 120B · $2.40/M · Q 84Llama 3.1 Nemotron 70B Instruct · $1.00/M · Q 83Qwen 3 235B · $3.00/M · Q 83o1-mini · $12/M · Q 83MiniMax M2.7 · $2.80/M · Q 82Llama 3.1 405B · $3.50/M · Q 81Command A · $10/M · Q 81Llama 3.1 Nemotron 70B Reward · $0.500/M · Q 80Qwen 2.5 Coder 32B · $0.800/M · Q 80Llama 3 70B · $0.880/M · Q 80Gemini 1.5 Pro · $5.00/M · Q 801. Gemini 2.0 Flash · $0.400/M · Q 8012. DeepSeek V3 · $0.420/M · Q 8123. HelpSteer2 Llama 3.1 70B · $0.500/M · Q 8234. Nemotron 70B · $0.880/M · Q 8345. Llama 3.2 90B Vision Instruct · $1.20/M · Q 8456. DeepSeek R1 · $2.19/M · Q 8867. Grok-3 · $15/M · Q 9178. Llama 4 Behemoth · $16/M · Q 938
30 plotted8 on frontierPareto$/1M out · log
On the Frontier
Model$/MQ
1Gemini 2.0 FlashPick$0.40080
2DeepSeek V3$0.42081
3HelpSteer2 Llama 3.1 70B$0.50082
4Nemotron 70B$0.88083
5Llama 3.2 90B Vision Instruct$1.2084
6DeepSeek R1$2.1988
7Grok-3$1591
8Llama 4 Behemoth$1693
Cheapest 85+ quality$2.19/M
Median price plotted$4.00/M
Frontier vs median10.0× cheaper

Market moves

The biggest provider price drops and the newest models, with the dates they landed.

Best in class, by workload

Top model for Code, Math, Reasoning and Vision, plus each size class — no single model wins everything.

Workloads

Chatbot, code generation, document analysis and more, each with a defensible starter model.

Provider spotlight

The top inference providers ranked by reputation, with model coverage and price ranges.

Market insight

Output prices span 4 orders of magnitude

From $0.005 to $150 per million output tokens across 189 priced models. The model you pick, not the task, drives the bill.

Compare on the leaderboard
Pricing · live catalogue

$0.005 to $150 per million tokens

Cheapest output price for each of 189 priced models, on a log scale. The model you choose can swing your bill by orders of magnitude.

median $0.50$0.005$150$/M output tokens (log)
Cheapest
$0.005
Median
$0.50
Priced models
189
Compare

How InferenceBench works

How the numbers are sourced

(1)Collect
  • Provider pricing
  • Model and GPU specs
  • Benchmark results
  • A source for each
(2)Standardise
  • Normalise units
  • Convert formats
  • Validate on import
(3)Recommend
  • Rank by quality and cost
  • Match to your workload
  • Score real value
Verified benchmarks

Signed, reproducible benchmarks

Run the open-source InferenceBench CLI to generate a signed record — the hardware, software stack, dataset hash and metrics, cryptographically sealed so anyone can verify the run.

  • Hardware + software stack fingerprint
  • Dataset hash + cryptographically signed record
  • Vendor-neutral, independently verifiable

Infrastructure Intelligence

Where the GPUs actually live: facilities, connectivity and carbon intensity across the global compute map.

2,254
Data centres tracked
53
Countries covered
19
GPU providers tracked
75
Internet exchanges
Explore Infrastructure Intelligence

Data centres by country

2,254 total
United States536 · 24%
China333 · 15%
Japan135 · 6%
India111 · 5%
Germany91 · 4%
South Korea88 · 4%

Playground

Test and compare models live

Send one prompt to several models and compare the answers, latency and cost side by side.

Worked examples

Build versus buy, the throughput bottleneck and a price forecast, each computed end to end.

Build vs Buy

Llama 3 70B 1M Context chatbot, 1B tokens / month

Two cost paths for the same workload: rent the API per-token from the cheapest provider, or rent 2× H100 PCIe 24×7 and serve it yourself.

API (cheapest provider)$740/mo
Self-host (2× H100 PCIe)$3.3k/mo
$2.6kspent extra/mo · -352% pricier to self-host
Bottleneck X-Ray

Llama 3 70B 1M Context FP8 on H100 PCIe, batch 1

Decode is memory-bandwidth dominated at batch 1: every token reloads the full weight matrix. Compute sits idle. Splitting across 2 GPUs (TP=2) doubles the BW ceiling.

99%BW-bound
BW ceiling28 tok/s
Compute ceiling11k tok/s
+28 tok/sat TP=2 (2× BW)
Where prices live

Output-token prices · top 20 by quality

The spread between budget and premium $/M, today. Sparkline shows the sorted distribution on log scale.

63×spread · p10 → p90
  • Budget · p10Llama 3.2 90B Vision Instruct$1.20
  • MedianNemotron Ultra 253B$6.00
  • Premium · p90Claude Opus 4$75.00
20 models with verified pricing & quality

Tools and calculators

Free, no signup — pick the part of the decision you care about.

For developers

Build on the data

Everything behind the leaderboard as a free public API — no key required.

pricing.ts
curl-friendly · zero auth
// Get the cheapest provider for Llama 3.1 70B
const res = await fetch(
  'https://inferencebench.io/api/v1/models/meta-llama/llama-3.1-70b/pricing'
);
const { providers } = await res.json();

providers
  .sort((a, b) => a.output_per_m - b.output_per_m)
  .slice(0, 3)
  .forEach((p) =>
    console.log(`${p.provider} · $${p.output_per_m}/M`)
  );

// DeepInfra   · $0.40/M
// Groq        · $0.79/M
// Openrouter  · $0.79/M
Public REST API

Build whatever you want on the data

Same data the leaderboard renders from — exposed as plain JSON. No API key. Rate-limit generous.

  • GET/api/v1/models
    307
  • GET/api/v1/gpus
    60
  • GET/api/v1/providers
    19
  • GET/api/v1/pricing
    live snapshots
  • GET/api/v1/leaderboard
    9 categories
Read the API docs

FAQ

Frequently asked questions

What is an AI inference benchmark?

An AI inference benchmark measures how fast a GPU or cloud provider can generate tokens from a large language model (LLM). Key metrics include tokens per second (throughput), time to first token (TTFT), inter-token latency (ITL), and cost per million tokens.

How does InferenceBench measure GPU performance?

InferenceBench uses a roofline performance model combined with CUDA kernel-level modeling (FlashAttention, PagedAttention, fused kernels) to predict real-world inference throughput. Results are validated against actual benchmarks from the HuggingFace LLM Perf Leaderboard and provider-reported data.

Which GPU is fastest for LLM inference?

Performance depends on model size. For large models (70B+), the NVIDIA B200 and H200 lead in throughput. For mid-size models (7B–30B), the H100 SXM offers the best price-performance. For budget deployments, the RTX 4090 and L40S are strong contenders.

How often is benchmark data updated?

Pricing data is refreshed every 6 hours via automated API calls to providers. Benchmark results are updated when new GPU hardware or model architectures are released. Community-submitted data is verified before inclusion.