🌱⚡
GPU Garden
Baue deine eigene Compute-Power
Link in die Zwischenablage kopiert! Bereit zum Teilen.

Kann meine GPU Open-Weight KI ausführen?

Präzise, clientseitige VRAM-Abschätzung für LLaMA 3.1, Qwen 2.5, DeepSeek und FLUX. Berücksichtigt quantisierte Gewichte, Grouped Query Attention (GQA) KV-Cache und CUDA-Treiber-Puffer.

Schnellszenarien:

GGUF / Ollama / vLLM
8k tokens
KV-Cache-Präzision:
Hardware-Urteil
Läuft komfortabel

Gesamter VRAM erforderlich
7.64 GB
GPU-VRAM-Kapazität
8.00 GB
Verfügbarer Spielraum
+0.36 GB
Geschätzter Durchsatz
~68 tok/s
VRAM-Zuweisungsaufschlüsselung 100% Kapazität
Gewichte: 4.74 GB
KV-Cache: 1.00 GB
CUDA/Puffer: 1.90 GB
Spielraum: 0.36 GB
1. Modellgewichte-Formel 4.74 GB
Params * (BPW / 8) * overhead
8.03B params × (4.50 BPW / 8) × 1.05 overhead = 4.74 GB
2. KV-Cache-Formel 1.00 GB
2 * Layers * KV_Heads * Head_Dim * Context_Tokens * Precision_Bytes
2 × 32 layers × 8 KV heads × 128 dim × 8,192 ctx × 2B (FP16) = 1.00 GB
3. Transienter Aktivierungspuffer & Treiber-Basis (1.50 GB) 1.90 GB
Transient Activation Buffer & Driver Baseline (1.50 GB)
Activations (0.40 GB) + Driver Baseline (1.50 GB) = 1.90 GB
4. Gesamter VRAM-Bedarf vs. Verfügbare Kapazität & Spielraum 7.64 GB
Total Required VRAM vs Available Capacity & Headroom
Erforderlich: 4.74 GB + 1.00 GB + 1.90 GB = 7.64 GB
Verfügbar: 8.00 GB Spielraum: +0.36 GB
Referenz-Spickzettel

Beliebte GPU-VRAM-Kompatibilität & Modell-Sweet-Spots

Finde den exakten Speicherbedarf für Open-Weight-KI auf deiner Hardware. Klicke auf eine Karte, um Parameter und KV-Cache sofort im Matchmaker zu testen.

Aktualisiert für LLaMA 3.2, Qwen 2.5 & FLUX.1
24GB GDDR6X 1,008 GB/s

NVIDIA GeForce RTX 4090

Der unangefochtene Consumer-Champion. Führt Qwen 2.5 Coder 32B bei Q4 mit 32k Kontext, FLUX.1 [dev] in unkomprimiertem FP16 und LLaMA 3.1 8B in FP16 bei über 120 Tokens/Sek. aus.

Amazon →
48GB VRAM (Dual) 1,872+ GB/s

Dual NVIDIA RTX 3090 / 4090 Rig

Der Homelab-Goldstandard von r/LocalLLaMA. 48GB kombinierter VRAM bewältigen Meta LLaMA 3.3 70B und Qwen 2.5 72B bei Q4_K_M mit 32k Kontext komplett im VRAM ohne CPU-Auslagerung.

Amazon →
16GB GDDR6X 672 GB/s

NVIDIA GeForce RTX 4070 Ti Super

Der 16GB-Preis-Leistungs-Sieger. Verfügt über einen 256-Bit-Speicherbus, ideal für Qwen 2.5 14B, LLaMA 3.2 11B Vision, Mistral Small 24B (Q4) und FLUX Schnell.

Amazon →
12GB GDDR6 360 GB/s

NVIDIA GeForce RTX 3060 12GB

Der unbestrittene Budget-König unter 300 €. Der großzügige 12GB-Puffer stemmt LLaMA 3.1 8B mit riesigen 64k-Kontextfenstern sowie 14B-Modelle bei Q4_K_M absolut ohne OOM-Fehler.

Amazon →
~48GB Allocatable 400 GB/s

Apple M3 / M4 Max (64GB Unified)

Die lautlose Entwickler-Workstation. macOS Metal allokiert rund 48GB Unified RAM direkt für llama.cpp und MLX und betreibt 70B-Modelle bei Q4_K_M mit vollem 128k-Kontext.

Amazon →
16GB GDDR6 288 GB/s

NVIDIA GeForce RTX 4060 Ti 16GB

Der günstigste Einstieg in moderne 16GB VRAM (~449 €). Läuft mit kühlen 165W TDP und betreibt FLUX.1 Schnell NF4 sowie Qwen 2.5 14B mühelos.

Amazon →
Referenzmatrix

KI-Modell VRAM-Anforderungsleitfaden & Sweet-Spot-Matrix

Exakte VRAM-Anforderungen über 4-Bit (Q4_K_M), 8-Bit (Q8_0) und 16-Bit (FP16) Präzisionen mit empfohlenen Hardware-Sweet-Spots.

Quantized GGUF • AWQ • FP16 / bfloat16
Modell & Architektur Parameter 4-Bit (Q4_K_M) VRAM 8-Bit (Q8_0) VRAM 16-Bit (FP16) VRAM Mindest-VRAM Sweet-Spot GPU Aktion
LLaMA 3.3 70B
Meta AI • Flagship LLM
70.6B Dense GQA
128k context
~41.7 GB
46.1 GB w/ 8k KV
~77.3 GB
81.7 GB w/ 8k KV
~144.0 GB
148.4 GB w/ 8k KV
48 GB
Dual RTX 3090 / 4090
48GB VRAM (Dual GPU)
Qwen 2.5 Coder 32B
Alibaba Cloud • SOTA Coding
32.5B Dense GQA
128k context
~19.2 GB
23.1 GB w/ 8k KV
~35.6 GB
39.5 GB w/ 8k KV
~66.3 GB
70.2 GB w/ 8k KV
24 GB
RTX 4090 24GB
1,008 GB/s (Consumer King)
DeepSeek V2.5 / V3 MoE
DeepSeek • MLA MoE
236B MoE MLA
21B active params
~139.4 GB
142.5 GB w/ 8k KV
~258.3 GB
261.4 GB w/ 8k KV
~481.4 GB
484.6 GB w/ 8k KV
64 GB+ (Offload)
Apple M3 Max 64GB
Unified RAM / Dual 3090
Mistral Small 24B
Mistral AI • Math & Reasoning
23.6B Dense
128k context
~14.2 GB
17.8 GB w/ 8k KV
~26.3 GB
29.9 GB w/ 8k KV
~49.0 GB
52.6 GB w/ 8k KV
16 GB
RTX 4070 Ti Super 16GB
672 GB/s (256-bit bus)
LLaMA 3.2 11B Vision
Meta AI • Multimodal Vision
13.8B Multimodal
128k context
~8.2 GB
11.1 GB w/ 8k KV
~15.1 GB
18.0 GB w/ 8k KV
~28.2 GB
31.1 GB w/ 8k KV
12 GB - 16 GB
RTX 4070 Ti Super 16GB
Full 16GB Headroom
LLaMA 3.1 8B
Meta AI • Open Benchmark
8.03B Dense
128k context
~4.7 GB
7.6 GB w/ 8k KV
~8.8 GB
11.7 GB w/ 8k KV
~16.4 GB
19.3 GB w/ 8k KV
8 GB
RTX 3060 12GB
Budget King (<$300)
FLUX.1 [dev]
Black Forest Labs • SOTA DiT
12.0B DiT
+4.9B T5-XXL / CLIP
~10.2 GB
15.5 GB w/ Act
~18.7 GB
24.0 GB w/ Act
~34.6 GB
39.9 GB w/ Act
16 GB (Q4) / 24 GB
RTX 4090 24GB
Native FP16 / FP8
SDXL 1.0
Stability AI • UNet 1024px
2.6B UNet
+0.8B CLIP-L/G
~3.3 GB
7.0 GB w/ Act
~4.7 GB
8.4 GB w/ Act
~7.1 GB
10.8 GB w/ Act
8 GB - 12 GB
RTX 3060 12GB
Full LoRA + ControlNet Headroom