Kimi K3
Moonshot AI · moe · 2779.9B parameters · 1,048,576 context
Parameters
2.8T
Context Window
1024K tokens
Architecture
MoE
Best GPU
B200 SXM
Cheapest API
$15.00/M
Intelligence Brief
Kimi K3 is a 2779.9B parameter Mixture-of-Experts (896 experts, 16 active) model from Moonshot AI, featuring Multi-Head Attention (MHA) with 93 layers and 7,168 hidden dimensions. With a 1,048,576 token context window, it supports tools, vision, structured output, code, math, multilingual, reasoning. The most cost-effective API deployment is via openrouter at $15.00/M output tokens. For self-hosted inference, B200 SXM delivers optimal throughput at $136352/month.
Provider pricing
1 provider · canonical: openrouter| Provider | Input $/M | Output $/M ▲ | Notes |
|---|---|---|---|
| openroutercanonical | $3.00 | $15.00 | cheapest input · cheapest output |
Prices update via the nightly pricing cron + admin approvals at /admin/ingest-queue. The leaderboard's Input/Output cells show the canonical rate above; this table shows the full spread.
Recent changes
Loading…
Related models
5 suggestions
Kimi K2.5Kimi · 32Bfree/M out
Kimi K2.7 CodeKimi · 32.9B$3.50/M out
DeepSeek R1DeepSeek R1 · 37B$2.00/M out
DeepSeek V3-0324DeepSeek V3 · 37Bfree/M out
Qwen 3 235BQwen 3 · 22Bfree/M out
Picks: same family first, then same vendor within ±2× params, then top tag-overlap matches. Price shown is the cheapest Output $/M across providers — the row's page shows the canonical anchor.
Architecture Details
Memory Requirements
BF16 Weights
5178.0 GB
FP8 Weights
2589.0 GB
INT4 Weights
1294.5 GB
Fits on (multi-GPU with Tensor Parallelism)
Multi-GPU configurations use Tensor Parallelism (TP) to split model layers across GPUs. Requires NVLink or NVSwitch interconnect for optimal performance.
This model requires multi-GPU deployment. Minimum: 6x B300 (288GB each) with Tensor Parallelism.
GPU Compatibility Matrix
Kimi K3 is compatible with 0% of GPU configurations across 41 GPUs at 3 precision levels.
GPU Recommendations
FP8 · 32 GPUs · tensorrt-llm
83/100
score
Throughput
140.0 tok/s
Latency (ITL)
7.1ms
Est. TTFT
1ms
Cost/Month
$136352
Cost/M Tokens
$370.60
FP8 · 16 GPUs · tensorrt-llm
83/100
score
Throughput
140.0 tok/s
Latency (ITL)
7.1ms
Est. TTFT
1ms
Cost/Month
$159432
Cost/M Tokens
$433.33
FP8 · 32 GPUs · tensorrt-llm
80/100
score
Throughput
140.0 tok/s
Latency (ITL)
7.1ms
Est. TTFT
1ms
Cost/Month
$81690
Cost/M Tokens
$222.03
Deployment Options
API Deployment
openrouter
$15.00/M
output tokens
Single GPU
Requires multi-GPU setup (2589 GB VRAM needed)
Multi-GPU
B200 SXM x32
140.0 tok/s
TP· $136352/mo
API Pricing Comparison
| Provider | Input $/M | Output $/M | Badges |
|---|---|---|---|
| openrouter | $3.00 | $15.00 | Cheapest |
Cost Analysis
| Provider | Input $/M | Output $/M | ~Monthly Cost |
|---|---|---|---|
| openrouterBest Value | $3.00 | $15.00 | $90 |
Cost per 1,000 Requests
Short (500 tok)
$4.50
via openrouter
Medium (2K tok)
$18.00
via openrouter
Long (8K tok)
$54.00
via openrouter
Performance Estimates
Throughput by GPU
VRAM Breakdown (B200 SXM, FP8)
Precision Impact
bf16
173.7 GB
weights/GPU
fp8
86.9 GB
weights/GPU
~140.0 tok/s
int4
43.4 GB
weights/GPU
Capabilities
Features
Supported Frameworks
Supported Precisions
Where to Deploy Kimi K3
Self-Hosted Infrastructure
Similar Models
Kimi K2.7 Code
1026.9B params · moe
Quality: 50
from $3.50/M
Kimi K2.5
1000B params · moe
Quality: 54
from $0.00/M
Llama 4 Behemoth
2000B params · moe
Quality: 93
from $16.00/M
GPT-4.5 Preview
1500B params · moe
Quality: 93
from $150.00/M
Frequently Asked Questions
How much VRAM does Kimi K3 need for inference?
Kimi K3 requires approximately 5178.0 GB of VRAM at BF16 precision, 2589.0 GB at FP8, or 1294.5 GB at INT4 quantization. Additional VRAM is needed for KV-cache (2678400 bytes per token) and activations (~0.00 GB).
What is the best GPU for Kimi K3?
The top recommended GPU for Kimi K3 is the B200 SXM (x32) using FP8 precision. It achieves approximately 140.0 tokens/sec at an estimated cost of $136352/month ($370.60/M tokens). Score: 83/100.
How much does Kimi K3 inference cost?
Kimi K3 API inference starts from $3.00/M input tokens and $15.00/M output tokens. Self-hosted inference costs depend on your GPU configuration — use our ROI calculator for a detailed breakdown.