Black Forest Labs
12.0 Billion Parameters (Diffusion Transformer / Flow)
Black Forest Labs FLUX.1 [dev] VRAM Requirements Guide
State-of-the-art open diffusion transformer. Combines a 12B MMDiT core with T5-XXL and CLIP-L text encoders for photorealistic generation.
4-Bit (Q4_K_M) - Recommended
12.00 GB
>98% perplexity retention. Standard Ollama/llama.cpp quantization.
8-Bit (Q8_0) - Near Lossless
18.50 GB
Virtually identical to FP16 (~99.9% fidelity). Demands high-capacity VRAM.
16-Bit (FP16 / BF16) - Full
26.50 GB
Uncompressed weights. Requires multi-GPU clusters or datacenter cards.
Recommended Hardware Sweet Spot
NVIDIA RTX 3060 12GB (NF4) or RTX 4090 24GB (FP16/Q8)
Provides sufficient headroom for weights, 8k+ context KV-cache, and CUDA runtime buffers without Out-Of-Memory (OOM) crashes.
Run Black Forest Labs FLUX.1 [dev] Locally
Ollama:
# Run via ComfyUI: git clone https://github.com/comfyanonymous/ComfyUI
vLLM:
# Use Diffusers pipeline: pipe = FluxPipeline.from_pretrained('black-forest-labs/FLUX.1-dev')
Test Custom Hardware & Context Windows
Simulate your exact GPU, calculate token throughput speeds, or compare cloud rental break-even costs.
Open in GPU Garden Matchmaker →