DeepSeek
236 Billion Parameters (21B Active MoE)
DeepSeek V2.5 / V3 MoE VRAM Requirements Guide
High-throughput Mixture-of-Experts architecture using Multi-head Latent Attention (MLA). Demands substantial RAM/VRAM to hold all expert weights.
4-Bit (Q4_K_M) - Recommended
143.50 GB
>98% perplexity retention. Standard Ollama/llama.cpp quantization.
8-Bit (Q8_0) - Near Lossless
260.00 GB
Virtually identical to FP16 (~99.9% fidelity). Demands high-capacity VRAM.
16-Bit (FP16 / BF16) - Full
490.00 GB
Uncompressed weights. Requires multi-GPU clusters or datacenter cards.
Recommended Hardware Sweet Spot
Quad RTX 3090 / Cloud 8x H100 / Mac Ultra 192GB
Provides sufficient headroom for weights, 8k+ context KV-cache, and CUDA runtime buffers without Out-Of-Memory (OOM) crashes.
Run DeepSeek V2.5 / V3 MoE Locally
Ollama:
ollama run deepseek-v2.5:236b-instruct-q4_K_M
vLLM:
vllm serve deepseek-ai/DeepSeek-V2.5 --tensor-parallel-size 4 --gpu-memory-utilization 0.95
Test Custom Hardware & Context Windows
Simulate your exact GPU, calculate token throughput speeds, or compare cloud rental break-even costs.
Open in GPU Garden Matchmaker →