WeeLLM — Running 20B Diffusion Models on a 4GB GPU
Posted on Sat 22 August 2026 in Inference • Tagged with diffusion, low-vram, layer-streaming, flux, local-inference
WeeLLM is a layer-streaming inference engine for large diffusion models. Its one claim is unusual: run a ~12B FLUX model, or even a ~20B Qwen-Image model, in under 4GB of VRAM — with no quantization. Full bfloat16 weights, degraded only in speed, never in precision.
The Core Idea
Layer Streaming The …
Continue reading
