Qwen-AgentWorld-35B-A3B
Qwen · moe · 34.7B parameters · 262,144 context
Parameters
34.7B
Context Window
256K tokens
Architecture
MoE
Best GPU
B200 SXM
Intelligence Brief
Qwen-AgentWorld-35B-A3B is a 34.7B parameter Mixture-of-Experts (256 experts, 8 active) model from Qwen, featuring Grouped Query Attention (GQA) with 40 layers and 2,048 hidden dimensions. With a 262,144 token context window, it supports tools, vision. For self-hosted inference, B200 SXM delivers optimal throughput at $4261/month.
Recent changes
Loading…
Related models
4 suggestions
Qwen3-ASR-0.6B-hfQwen · 0.8B—
mmE5-mllama-11b-instructintfloat · 10.6B—
siglip-so400m-14-384-flash-attn2-navitHuggingFaceM4 · 0.9B—
multilingual-e5-large-instructintfloat · 0.6B—
Picks: same family first, then same vendor within ±2× params, then top tag-overlap matches. Price shown is the cheapest Output $/M across providers — the row's page shows the canonical anchor.
Architecture Details
Memory Requirements
BF16 Weights
69.4 GB
FP8 Weights
34.7 GB
INT4 Weights
17.4 GB
GPU Compatibility Matrix
Qwen-AgentWorld-35B-A3B is compatible with 57% of GPU configurations across 41 GPUs at 3 precision levels.
GPU Recommendations
FP8 · 1 GPU · tensorrt-llm
100/100
score
Throughput
1.1K tok/s
Latency (ITL)
1.0ms
Est. TTFT
0ms
Cost/Month
$4261
Cost/M Tokens
$1.54
FP8 · 1 GPU · tensorrt-llm
100/100
score
Throughput
1.1K tok/s
Latency (ITL)
1.0ms
Est. TTFT
0ms
Cost/Month
$2553
Cost/M Tokens
$0.93
FP8 · 1 GPU · tensorrt-llm
100/100
score
Throughput
1.1K tok/s
Latency (ITL)
1.0ms
Est. TTFT
0ms
Cost/Month
$1794
Cost/M Tokens
$0.65
Deployment Options
API Deployment
No API pricing available
Single GPU
B200 SXM
$4261/mo
Min VRAM: 35 GB
Multi-GPU
RTX A6000 x2
411.6 tok/s
TP· $930/mo
API Pricing Comparison
No API pricing data available for this model.
Performance Estimates
Throughput by GPU
VRAM Breakdown (B200 SXM, FP8)
Precision Impact
bf16
69.4 GB
weights/GPU
fp8
34.7 GB
weights/GPU
~1.1K tok/s
int4
17.4 GB
weights/GPU
Capabilities
Features
Supported Frameworks
Supported Precisions
Where to Deploy Qwen-AgentWorld-35B-A3B
Self-Hosted Infrastructure
Similar Models
Aya 23 35B
35B params · dense
Quality: 50
from $1.50/M
Command R (August 2024)
35B params · dense
Quality: 68
from $0.60/M
Command R
35B params · dense
Quality: 68
from $0.50/M
Yi 1.5 34B
34.4B params · dense
Quality: 72
from $0.80/M
Code Llama 34B
34B params · dense
Quality: 55
from $0.78/M
Frequently Asked Questions
How much VRAM does Qwen-AgentWorld-35B-A3B need for inference?
Qwen-AgentWorld-35B-A3B requires approximately 69.4 GB of VRAM at BF16 precision, 34.7 GB at FP8, or 17.4 GB at INT4 quantization. Additional VRAM is needed for KV-cache (81920 bytes per token) and activations (~0.00 GB).
What is the best GPU for Qwen-AgentWorld-35B-A3B?
The top recommended GPU for Qwen-AgentWorld-35B-A3B is the B200 SXM using FP8 precision. It achieves approximately 1.1K tokens/sec at an estimated cost of $4261/month ($1.54/M tokens). Score: 100/100.
How much does Qwen-AgentWorld-35B-A3B inference cost?
Qwen-AgentWorld-35B-A3B inference costs vary by provider and GPU setup. Use our calculator for detailed cost estimates across all providers.