Kann meine GPU Open-Weight KI ausführen?
Präzise, clientseitige VRAM-Abschätzung für LLaMA 3.1, Qwen 2.5, DeepSeek und FLUX. Berücksichtigt quantisierte Gewichte, Grouped Query Attention (GQA) KV-Cache und CUDA-Treiber-Puffer.
Mathematische Formel & Aufschlüsselung prüfen Rechenweg zeigen
Params * (BPW / 8) * overhead
2 * Layers * KV_Heads * Head_Dim * Context_Tokens * Precision_Bytes
Transient Activation Buffer & Driver Baseline (1.50 GB)
Total Required VRAM vs Available Capacity & Headroom
Beliebte GPU-VRAM-Kompatibilität & Modell-Sweet-Spots
Finde den exakten Speicherbedarf für Open-Weight-KI auf deiner Hardware. Klicke auf eine Karte, um Parameter und KV-Cache sofort im Matchmaker zu testen.
NVIDIA GeForce RTX 4090
Der unangefochtene Consumer-Champion. Führt Qwen 2.5 Coder 32B bei Q4 mit 32k Kontext, FLUX.1 [dev] in unkomprimiertem FP16 und LLaMA 3.1 8B in FP16 bei über 120 Tokens/Sek. aus.
Dual NVIDIA RTX 3090 / 4090 Rig
Der Homelab-Goldstandard von r/LocalLLaMA. 48GB kombinierter VRAM bewältigen Meta LLaMA 3.3 70B und Qwen 2.5 72B bei Q4_K_M mit 32k Kontext komplett im VRAM ohne CPU-Auslagerung.
NVIDIA GeForce RTX 4070 Ti Super
Der 16GB-Preis-Leistungs-Sieger. Verfügt über einen 256-Bit-Speicherbus, ideal für Qwen 2.5 14B, LLaMA 3.2 11B Vision, Mistral Small 24B (Q4) und FLUX Schnell.
NVIDIA GeForce RTX 3060 12GB
Der unbestrittene Budget-König unter 300 €. Der großzügige 12GB-Puffer stemmt LLaMA 3.1 8B mit riesigen 64k-Kontextfenstern sowie 14B-Modelle bei Q4_K_M absolut ohne OOM-Fehler.
Apple M3 / M4 Max (64GB Unified)
Die lautlose Entwickler-Workstation. macOS Metal allokiert rund 48GB Unified RAM direkt für llama.cpp und MLX und betreibt 70B-Modelle bei Q4_K_M mit vollem 128k-Kontext.
NVIDIA GeForce RTX 4060 Ti 16GB
Der günstigste Einstieg in moderne 16GB VRAM (~449 €). Läuft mit kühlen 165W TDP und betreibt FLUX.1 Schnell NF4 sowie Qwen 2.5 14B mühelos.
KI-Modell VRAM-Anforderungsleitfaden & Sweet-Spot-Matrix
Exakte VRAM-Anforderungen über 4-Bit (Q4_K_M), 8-Bit (Q8_0) und 16-Bit (FP16) Präzisionen mit empfohlenen Hardware-Sweet-Spots.
| Modell & Architektur | Parameter | 4-Bit (Q4_K_M) VRAM | 8-Bit (Q8_0) VRAM | 16-Bit (FP16) VRAM | Mindest-VRAM | Sweet-Spot GPU | Aktion |
|---|---|---|---|---|---|---|---|
|
LLaMA 3.3 70B
Meta AI • Flagship LLM
|
70.6B Dense GQA
128k context
|
~41.7 GB
46.1 GB w/ 8k KV
|
~77.3 GB
81.7 GB w/ 8k KV
|
~144.0 GB
148.4 GB w/ 8k KV
|
48 GB |
Dual RTX 3090 / 4090
48GB VRAM (Dual GPU)
|
|
|
Qwen 2.5 Coder 32B
Alibaba Cloud • SOTA Coding
|
32.5B Dense GQA
128k context
|
~19.2 GB
23.1 GB w/ 8k KV
|
~35.6 GB
39.5 GB w/ 8k KV
|
~66.3 GB
70.2 GB w/ 8k KV
|
24 GB |
RTX 4090 24GB
1,008 GB/s (Consumer King)
|
|
|
DeepSeek V2.5 / V3 MoE
DeepSeek • MLA MoE
|
236B MoE MLA
21B active params
|
~139.4 GB
142.5 GB w/ 8k KV
|
~258.3 GB
261.4 GB w/ 8k KV
|
~481.4 GB
484.6 GB w/ 8k KV
|
64 GB+ (Offload) |
Apple M3 Max 64GB
Unified RAM / Dual 3090
|
|
|
Mistral Small 24B
Mistral AI • Math & Reasoning
|
23.6B Dense
128k context
|
~14.2 GB
17.8 GB w/ 8k KV
|
~26.3 GB
29.9 GB w/ 8k KV
|
~49.0 GB
52.6 GB w/ 8k KV
|
16 GB |
RTX 4070 Ti Super 16GB
672 GB/s (256-bit bus)
|
|
|
LLaMA 3.2 11B Vision
Meta AI • Multimodal Vision
|
13.8B Multimodal
128k context
|
~8.2 GB
11.1 GB w/ 8k KV
|
~15.1 GB
18.0 GB w/ 8k KV
|
~28.2 GB
31.1 GB w/ 8k KV
|
12 GB - 16 GB |
RTX 4070 Ti Super 16GB
Full 16GB Headroom
|
|
|
LLaMA 3.1 8B
Meta AI • Open Benchmark
|
8.03B Dense
128k context
|
~4.7 GB
7.6 GB w/ 8k KV
|
~8.8 GB
11.7 GB w/ 8k KV
|
~16.4 GB
19.3 GB w/ 8k KV
|
8 GB |
RTX 3060 12GB
Budget King (<$300)
|
|
|
FLUX.1 [dev]
Black Forest Labs • SOTA DiT
|
12.0B DiT
+4.9B T5-XXL / CLIP
|
~10.2 GB
15.5 GB w/ Act
|
~18.7 GB
24.0 GB w/ Act
|
~34.6 GB
39.9 GB w/ Act
|
16 GB (Q4) / 24 GB |
RTX 4090 24GB
Native FP16 / FP8
|
|
|
SDXL 1.0
Stability AI • UNet 1024px
|
2.6B UNet
+0.8B CLIP-L/G
|
~3.3 GB
7.0 GB w/ Act
|
~4.7 GB
8.4 GB w/ Act
|
~7.1 GB
10.8 GB w/ Act
|
8 GB - 12 GB |
RTX 3060 12GB
Full LoRA + ControlNet Headroom
|