Skip to content
Live pricing · updated daily

The definitive ranking of 345 AI models — by quality, cost, and value.

Independent. Open-source. Every number traced to its source. Compare 345 models across 19 providers on 60 GPUs — the data desk for everyone deciding which model to ship.

345Models
60GPUs
19Providers
12 moPrice history
FreePublic API
Open methodology
MIT-licensed
No signup
Refreshed daily
THE WIREEDITORS
0 / 24h
Pricing · last 12 months

Output prices fell 41% YoY.

Median $/1M output across the top 20 models, weighted by request volume. Milestone markers mark the major model releases.

GPT-4o-miniLlama 3.1 405Bo1-previewDeepSeek R1Llama 3.3Today$2$3$4
Today
$2.14
12 mo ago
$3.62
Analysis

Specialized models

Non-LLM — image-gen, speech, embeddings, video, protein

Specialized modelsimage-gen · speech · embeddings · video · protein

View all specialized
ModelTypeProviderArchPricing
Edify Imageimage-genNVIDIAdiffusion$0.040 / image
Nemotron-3-Embed-1B-NVFP4image-gennvidiaencoderFree · open weights
Nemotron-3-Embed-8B-BF16image-gennvidiaencoderFree · open weights
StyleGAN3image-genNVIDIAganFree · open weights
BioNeMo ESM-2 650MproteinNVIDIAencoderFree · open weights
Riva FastPitch (en-US)speech-ttsNVIDIAencoder-decoderFree · open weights
Riva HiFi-GAN (en-US)speech-ttsNVIDIAganFree · open weights
Maxine Eye Contactvideo-fxNVIDIApipeline$0.000 / video min
DINOv2 ViT-g/14 (NVIDIA-optimized)vision-embeddingNVIDIAencoderFree · open weights
NV-CLIPvision-embeddingNVIDIAencoderPer embedding

Top picks

One winner per axis — best value, best quality, cheapest, fastest

Market Map

Where every model sits on the price-vs-quality curve, and which few set the floor for everyone else.

Open comparator
As of 17 JUL 2026
$2.19/Mbuys quality 88
vs
$150/Mfor the same tier
68× spread
708090100$0.03$0.30$3$30o1 · $60/M · Q 93GPT-4.5 Preview · $150/M · Q 93Grok 3 · $15/M · Q 90Claude Opus 4 · $75/M · Q 90Gemini 2.0 Pro · $4.00/M · Q 88o3-mini · $4.40/M · Q 86Nemotron Ultra 253B · $6.00/M · Q 86Claude Sonnet 4 · $15/M · Q 86Nemotron 340B · $4.20/M · Q 85GPT-4o · $10/M · Q 85Llama 4 Maverick · $1.80/M · Q 84Nemotron-3 Super 120B · $2.40/M · Q 84Llama 3.1 Nemotron 70B Instruct · $1.00/M · Q 83Qwen 3 235B · $3.00/M · Q 83o1-mini · $12/M · Q 83MiniMax M2.7 · $2.80/M · Q 82Llama 3.1 405B · $3.50/M · Q 81Command A · $10/M · Q 81Llama 3.1 Nemotron 70B Reward · $0.500/M · Q 80Qwen 2.5 Coder 32B · $0.800/M · Q 80Llama 3 70B · $0.880/M · Q 80Gemini 1.5 Pro · $5.00/M · Q 801. Gemini 2.0 Flash · $0.400/M · Q 8012. DeepSeek V3 · $0.420/M · Q 8123. HelpSteer2 Llama 3.1 70B · $0.500/M · Q 8234. Nemotron 70B · $0.880/M · Q 8345. Llama 3.2 90B Vision Instruct · $1.20/M · Q 8456. DeepSeek R1 · $2.19/M · Q 8867. Grok-3 · $15/M · Q 9178. Llama 4 Behemoth · $16/M · Q 938
30 plotted8 on frontierPareto$/1M out · log
On the Frontier
Model$/MQ
1Gemini 2.0 FlashPick$0.40080
2DeepSeek V3$0.42081
3HelpSteer2 Llama 3.1 70B$0.50082
4Nemotron 70B$0.88083
5Llama 3.2 90B Vision Instruct$1.2084
6DeepSeek R1$2.1988
7Grok-3$1591
8Llama 4 Behemoth$1693
Cheapest 85+ quality$2.19/M
Median price plotted$4.00/M
Frontier vs median10.0× cheaper

Market Moves

Biggest price drops and the freshest releases — past 30 days

Class Leaders

Best for Code · Math · Reasoning · Vision · Small (<15B) · Large (70B+)

All categories

Workloads

Curated model + GPU shortlist for the workload you're building

Provider spotlight

The top inference providers ranked by reputation

Sandbox

Chat with any model · compare two at once · vote in the arena

Open arena

News & research

Latest deep-dives from the benchmark team

All posts

Worked examples

Pre-computed scenarios from the engine — Build vs Buy · Bottleneck · Forecast

See all analyses
Build vs Buy

Llama 3 70B 1M Context chatbot, 1B tokens / month

Two cost paths for the same workload: rent the API per-token from the cheapest provider, or rent 2× H100 PCIe 24×7 and serve it yourself.

API (cheapest provider)$740/mo
Self-host (2× H100 PCIe)$3.3k/mo
$2.6kspent extra/mo · -352% pricier to self-host
Bottleneck X-Ray

Llama 3 70B 1M Context FP8 on H100 PCIe, batch 1

Decode is memory-bandwidth dominated at batch 1: every token reloads the full weight matrix. Compute sits idle. Splitting across 2 GPUs (TP=2) doubles the BW ceiling.

99%BW-bound
BW ceiling28 tok/s
Compute ceiling11k tok/s
+28 tok/sat TP=2 (2× BW)
Where prices live

Output-token prices · top 20 by quality

The spread between budget and premium $/M, today. Sparkline shows the sorted distribution on log scale.

63×spread · p10 → p90
  • Budget · p10Llama 3.2 90B Vision Instruct$1.20
  • MedianNemotron Ultra 253B$6.00
  • Premium · p90Claude Opus 4$75.00
20 models with verified pricing & quality

Build on it

The data that powers the leaderboard, exposed as a free public API

pricing.ts
curl-friendly · zero auth
// Get the cheapest provider for Llama 3.1 70B
const res = await fetch(
  'https://inferencebench.io/api/v1/models/meta-llama/llama-3.1-70b/pricing'
);
const { providers } = await res.json();

providers
  .sort((a, b) => a.output_per_m - b.output_per_m)
  .slice(0, 3)
  .forEach((p) =>
    console.log(`${p.provider} · $${p.output_per_m}/M`)
  );

// DeepInfra   · $0.40/M
// Groq        · $0.79/M
// Openrouter  · $0.79/M
Public REST API

Build whatever you want on the data

Same data the leaderboard renders from — exposed as plain JSON. No API key. Rate-limit generous.

  • GET/api/v1/models
    307
  • GET/api/v1/gpus
    60
  • GET/api/v1/providers
    19
  • GET/api/v1/pricing
    live snapshots
  • GET/api/v1/leaderboard
    9 categories
Read the API docs

Tools & calculators

Go deeper into the parts you care about — every page is free, no signup

Methodology & trust

How we measure, where the data comes from, how to build on it

Frequently asked questions

Common questions about the benchmark, methodology, and how the data is sourced

What is an AI inference benchmark?

An AI inference benchmark measures how fast a GPU or cloud provider can generate tokens from a large language model (LLM). Key metrics include tokens per second (throughput), time to first token (TTFT), inter-token latency (ITL), and cost per million tokens.

How does InferenceBench measure GPU performance?

InferenceBench uses a roofline performance model combined with CUDA kernel-level modeling (FlashAttention, PagedAttention, fused kernels) to predict real-world inference throughput. Results are validated against actual benchmarks from the HuggingFace LLM Perf Leaderboard and provider-reported data.

Which GPU is fastest for LLM inference?

Performance depends on model size. For large models (70B+), the NVIDIA B200 and H200 lead in throughput. For mid-size models (7B–30B), the H100 SXM offers the best price-performance. For budget deployments, the RTX 4090 and L40S are strong contenders.

How often is benchmark data updated?

Pricing data is refreshed every 6 hours via automated API calls to providers. Benchmark results are updated when new GPU hardware or model architectures are released. Community-submitted data is verified before inclusion.

Ready to pick a model?

The full live ranking is one click away — sort, filter, and compare every model by quality, cost, and value.

Built with care · Open source · MIT-licensed data · No signup