Ground truth

The Model Matrix

The most popular open-source models with their real VRAM requirements, mapped to machines you can actually buy. LLM sizes are Q4_K_M quantization (the standard for local inference); image & video models are FP16 unless noted. Pick your VRAM to see what lights up.

MY VRAM:

How to read this: "Q4" = 4-bit quantized weights (GGUF Q4_K_M — near-lossless, the default for local LLM serving via llama.cpp, Ollama or LM Studio). "Q8" = 8-bit. "FP16" = full precision, needed for training/fine-tuning and the default for ComfyUI image models. VRAM figures include weights + KV-cache headroom for typical use. Multi-GPU machines split large models across cards automatically (tensor / pipeline / layer parallel) via vLLM, SGLang, ExLlamaV2 or ComfyUI.