Kimi-Linear-48B-A3B-Base
moonshotai · moe · 48B parameters · 1,048,576 context
Parameters
48B
Context Window
1024K tokens
Architecture
MoE
Best GPU
B200 SXM
Intelligence Brief
Kimi-Linear-48B-A3B-Base is a 48B parameter Mixture-of-Experts (256 experts, 8 active) model from moonshotai, featuring Multi-Head Attention (MHA) with 27 layers and 2,304 hidden dimensions. With a 1,048,576 token context window, it supports code, math, multilingual. For self-hosted inference, B200 SXM delivers optimal throughput at $4261/month.
Recent changes
Loading…
Related models
3 suggestions
Jamba 1.5 LargeJamba · 52B$8.00/M out
Jamba 1.5 MiniJamba · 12B$0.400/M out
Llama 4 ScoutLlama 4 · 17B$0.300/M out
Picks: same family first, then same vendor within ±2× params, then top tag-overlap matches. Price shown is the cheapest Output $/M across providers — the row's page shows the canonical anchor.
Architecture Details
Memory Requirements
BF16 Weights
96.0 GB
FP8 Weights
48.0 GB
INT4 Weights
24.0 GB
GPU Compatibility Matrix
Kimi-Linear-48B-A3B-Base is compatible with 40% of GPU configurations across 41 GPUs at 3 precision levels.
GPU Recommendations
FP8 · 1 GPU · tensorrt-llm
100/100
score
Throughput
1.1K tok/s
Latency (ITL)
1.0ms
Est. TTFT
0ms
Cost/Month
$4261
Cost/M Tokens
$1.54
FP8 · 1 GPU · tensorrt-llm
100/100
score
Throughput
1.1K tok/s
Latency (ITL)
1.0ms
Est. TTFT
0ms
Cost/Month
$4271
Cost/M Tokens
$1.55
FP8 · 1 GPU · tensorrt-llm
100/100
score
Throughput
1.1K tok/s
Latency (ITL)
1.0ms
Est. TTFT
0ms
Cost/Month
$6169
Cost/M Tokens
$2.24
Deployment Options
API Deployment
No API pricing available
Single GPU
B200 SXM
$4261/mo
Min VRAM: 48 GB
Multi-GPU
A100 80GB SXM x2
1.1K tok/s
TP· $2259/mo
API Pricing Comparison
No API pricing data available for this model.
Performance Estimates
Throughput by GPU
VRAM Breakdown (B200 SXM, FP8)
Precision Impact
bf16
96.0 GB
weights/GPU
fp8
48.0 GB
weights/GPU
~1.1K tok/s
int4
24.0 GB
weights/GPU
Capabilities
Features
Supported Frameworks
Supported Precisions
Where to Deploy Kimi-Linear-48B-A3B-Base
Self-Hosted Infrastructure
Similar Models
Mixtral 8x7B Instruct
46.7B params · moe
Quality: 69
from $0.24/M
Mixtral 8x7B
46.7B params · moe
Quality: 67
from $0.50/M
Amazon Nova Pro
50B params · dense
Quality: 36
from $3.20/M
Gemini 1.5 Flash
50B params · moe
Quality: 75
from $0.30/M
Gemini 2.0 Flash
50B params · moe
Quality: 80
from $0.40/M
Frequently Asked Questions
How much VRAM does Kimi-Linear-48B-A3B-Base need for inference?
Kimi-Linear-48B-A3B-Base requires approximately 96.0 GB of VRAM at BF16 precision, 48.0 GB at FP8, or 24.0 GB at INT4 quantization. Additional VRAM is needed for KV-cache (64512 bytes per token) and activations (~6.00 GB).
What is the best GPU for Kimi-Linear-48B-A3B-Base?
The top recommended GPU for Kimi-Linear-48B-A3B-Base is the B200 SXM using FP8 precision. It achieves approximately 1.1K tokens/sec at an estimated cost of $4261/month ($1.54/M tokens). Score: 100/100.
How much does Kimi-Linear-48B-A3B-Base inference cost?
Kimi-Linear-48B-A3B-Base inference costs vary by provider and GPU setup. Use our calculator for detailed cost estimates across all providers.