If you run local AI models for coding, writing, or research, the question is always the same: which model actually fits in your GPU without crashing mid-task? The Stack published a tier-by-tier breakdown covering hardware from 4GB all the way to 512GB.
The core principle: do not load the largest model your GPU can technically hold. Leave room for runtime operations, context handling, and repository files. Quantized builds, specifically Q4 and Q5, compress model weights and reduce memory demand without significant performance loss.
Model recommendations by GPU tier
- 4 to 6 GB: Quen 3.52B or 3.54B (Q4 builds). Single-function tasks only: text generation or basic image recognition.
- 8 to 16 GB: Ornith 1.59B at Q4 to Q8 quantization. At 16GB, OpenAI’s GPTO OSS 20B handles multi-functional tasks with moderate complexity.
- 20 to 24 GB: Quen 3.827B (Q4KM build) for advanced text or image processing. Ornith 1.535B stays reliable for lighter workloads.
- 32 to 48 GB: Quen 3.827B again, now at Q5 to Q8 precision as memory allows. Well-suited for multi-modal workflows.
- 64 to 96 GB: Quen 3 Coder Next (4-bit) or GPTOSS 120B. Apple Silicon users in this range should consider hybrid setups that combine GPU memory with SSD streaming.
- 128 to 192 GB: Quen 3.8 Flash Next (4-bit), GPTOSS 120B, or DeepSk 4.1 Flash depending on system architecture. May require custom software configurations.
- 256 to 512 GB: GLM 5.3 Flash (4-bit) or Ornith 397B (6-bit). Mac users at this tier can use DeepSk V4.1 Flash with SSD streaming.
A note on unified memory systems
Apple Silicon and similar unified memory architectures share VRAM with the rest of the system. Deduct your system memory usage from the total before picking a model. SSD streaming is the practical workaround for running larger models on these machines without overloading the GPU.
What to do if your GPU falls between tiers
Use the nearest lower tier as your starting point. A 12GB card follows the 8 to 16GB recommendations. A 40GB card follows the 32 to 48GB guidance. Adjust quantization upward as your actual free memory allows. Test with real tasks before committing to a setup.
The takeaway is straightforward: pick the smallest model that handles your actual workload. Headroom is not wasted, it is what keeps your setup stable.
