Nemotron-3-Embed-1B-BF16
nvidia · dense · 1.1B parameters · 262,144 context
Parameters
1.1B
Context Window
256K tokens
Architecture
Dense
Best GPU
RTX 3080
Intelligence Brief
Nemotron-3-Embed-1B-BF16 is a 1.1B parameter DENSE model from nvidia, featuring Grouped Query Attention (GQA) with 16 layers and 2,048 hidden dimensions. With a 262,144 token context window, it supports general text generation. For self-hosted inference, RTX 3080 delivers optimal throughput at $133/month.
Recent changes
Loading…
Related models
3 suggestions
multilingual-e5-large-instructintfloat · 0.6B—
mmE5-mllama-11b-instructintfloat · 10.6B—
NV Retriever v1NV Retriever · 0.33B$0.0060/M out
Picks: same family first, then same vendor within ±2× params, then top tag-overlap matches. Price shown is the cheapest Output $/M across providers — the row's page shows the canonical anchor.
Architecture Details
Memory Requirements
BF16 Weights
2.2 GB
FP8 Weights
1.1 GB
INT4 Weights
0.6 GB
GPU Compatibility Matrix
Nemotron-3-Embed-1B-BF16 is compatible with 100% of GPU configurations across 41 GPUs at 3 precision levels.
GPU Recommendations
BF16 · 1 GPU · vllm
90/100
score
Throughput
1.9K tok/s
Latency (ITL)
0.5ms
Est. TTFT
0ms
Cost/Month
$133
Cost/M Tokens
$0.03
BF16 · 1 GPU · vllm
90/100
score
Throughput
667.6 tok/s
Latency (ITL)
1.5ms
Est. TTFT
0ms
Cost/Month
$209
Cost/M Tokens
$0.12
BF16 · 1 GPU · vllm
90/100
score
Throughput
1.1K tok/s
Latency (ITL)
0.9ms
Est. TTFT
0ms
Cost/Month
$85
Cost/M Tokens
$0.03
Deployment Options
API Deployment
No API pricing available
Single GPU
RTX 3080
$133/mo
Min VRAM: 1 GB
Multi-GPU
RTX 3080
1.9K tok/s
Best available config
API Pricing Comparison
No API pricing data available for this model.
Performance Estimates
Throughput by GPU
VRAM Breakdown (RTX 3080, BF16)
Precision Impact
bf16
2.2 GB
weights/GPU
~1.9K tok/s
fp8
1.1 GB
weights/GPU
int4
0.6 GB
weights/GPU
Capabilities
Features
Supported Frameworks
Supported Precisions
Where to Deploy Nemotron-3-Embed-1B-BF16
Self-Hosted Infrastructure
Similar Models
SantaCoder 1.1B
1.1B params · dense
Quality: 50
Parakeet TDT 1.1B
1.1B params · dense
Quality: 50
from $0.04/M
TinyLlama 1.1B Chat
1.1B params · dense
Quality: 50
TinyLlama 1.1B
1.1B params · dense
Quality: 50
Gemma 3 1B
1B params · dense
Quality: 35
Frequently Asked Questions
How much VRAM does Nemotron-3-Embed-1B-BF16 need for inference?
Nemotron-3-Embed-1B-BF16 requires approximately 2.2 GB of VRAM at BF16 precision, 1.1 GB at FP8, or 0.6 GB at INT4 quantization. Additional VRAM is needed for KV-cache (65536 bytes per token) and activations (~0.00 GB).
What is the best GPU for Nemotron-3-Embed-1B-BF16?
The top recommended GPU for Nemotron-3-Embed-1B-BF16 is the RTX 3080 using BF16 precision. It achieves approximately 1.9K tokens/sec at an estimated cost of $133/month ($0.03/M tokens). Score: 90/100.
How much does Nemotron-3-Embed-1B-BF16 inference cost?
Nemotron-3-Embed-1B-BF16 inference costs vary by provider and GPU setup. Use our calculator for detailed cost estimates across all providers.