siglip-so400m-14-384-flash-attn2-navit
HuggingFaceM4 · dense · 0.9B parameters · 8,192 context
Parameters
0.9B
Context Window
8K tokens
Architecture
Dense
Best GPU
RTX 4060
Intelligence Brief
siglip-so400m-14-384-flash-attn2-navit is a 0.9B parameter DENSE model from HuggingFaceM4, featuring Multi-Head Attention (MHA) with 27 layers and 1,152 hidden dimensions. With a 8,192 token context window, it supports vision. For self-hosted inference, RTX 4060 delivers optimal throughput at $209/month.
Recent changes
Loading…
Related models
3 suggestions
MiniMax-M2.1MiniMax-M2.1 · 7B—
MiniMax-M2.5MiniMax-M2 · 7B$1.20/M out
MiniMax-M2MiniMax-M2 · 7B$1.10/M out
Picks: same family first, then same vendor within ±2× params, then top tag-overlap matches. Price shown is the cheapest Output $/M across providers — the row's page shows the canonical anchor.
Architecture Details
Memory Requirements
BF16 Weights
1.8 GB
FP8 Weights
0.9 GB
INT4 Weights
0.5 GB
GPU Compatibility Matrix
siglip-so400m-14-384-flash-attn2-navit is compatible with 100% of GPU configurations across 41 GPUs at 3 precision levels.
GPU Recommendations
BF16 · 1 GPU · vllm
90/100
score
Throughput
734.4 tok/s
Latency (ITL)
1.4ms
Est. TTFT
0ms
Cost/Month
$209
Cost/M Tokens
$0.11
BF16 · 1 GPU · vllm
90/100
score
Throughput
1.2K tok/s
Latency (ITL)
0.8ms
Est. TTFT
0ms
Cost/Month
$85
Cost/M Tokens
$0.03
BF16 · 1 GPU · tensorrt-llm
78/100
score
Throughput
3.5K tok/s
Latency (ITL)
0.3ms
Est. TTFT
0ms
Cost/Month
$4261
Cost/M Tokens
$0.46
Deployment Options
API Deployment
No API pricing available
Single GPU
RTX 4060
$209/mo
Min VRAM: 1 GB
Multi-GPU
RTX 4060
734.4 tok/s
Best available config
API Pricing Comparison
No API pricing data available for this model.
Performance Estimates
Throughput by GPU
VRAM Breakdown (RTX 4060, BF16)
Capabilities
Features
Supported Frameworks
Supported Precisions
Where to Deploy siglip-so400m-14-384-flash-attn2-navit
Self-Hosted Infrastructure
Similar Models
Gemma 3 1B
1B params · dense
Quality: 35
Llama Guard 3 1B
1B params · dense
Quality: 50
Canary 1B
1B params · dense
Quality: 50
from $0.04/M
Qwen3-ASR-0.6B-hf
0.8B params · dense
Quality: 50
CSM-1B
1B params · dense
Quality: 50
Frequently Asked Questions
How much VRAM does siglip-so400m-14-384-flash-attn2-navit need for inference?
siglip-so400m-14-384-flash-attn2-navit requires approximately 1.8 GB of VRAM at BF16 precision, 0.9 GB at FP8, or 0.5 GB at INT4 quantization. Additional VRAM is needed for KV-cache (124416 bytes per token) and activations (~0.00 GB).
What is the best GPU for siglip-so400m-14-384-flash-attn2-navit?
The top recommended GPU for siglip-so400m-14-384-flash-attn2-navit is the RTX 4060 using BF16 precision. It achieves approximately 734.4 tokens/sec at an estimated cost of $209/month ($0.11/M tokens). Score: 90/100.
How much does siglip-so400m-14-384-flash-attn2-navit inference cost?
siglip-so400m-14-384-flash-attn2-navit inference costs vary by provider and GPU setup. Use our calculator for detailed cost estimates across all providers.