VRAM Requirements for Running LLMs Locally
The primary consideration when deploying a local LLM is straightforward: will it fit within your GPU's capacity? The answer hinges on the model size, the degree of quantization, and the length of the context. This guide offers a practical foundation for determining the appropriate amount of VRAM.
The Impact of Quantization on VRAM
Quantization lowers the precision used to store model weights. Applying lower-bit quantization shrinks the model size and reduces VRAM consumption, albeit with a slight compromise in output quality.
| Quantisation | Bits per weight | Typical use |
|---|---|---|
| Q8_0 | 8 | Very high quality |
| Q6_K | ~6.6 | Very good quality |
| Q5_K_M | ~5.5 | Good quality and size |
| Q4_K_M | ~4.5 | Good balance of size and quality |
| Q3_K_M | ~3.5 | Lower VRAM, more quality loss |
Q4_K_M is a frequent choice when VRAM is constrained. If you have greater VRAM available, opting for Q5 or Q6 allows you to run the same model with less aggressive quantization.
Estimated VRAM by Model Size
The following figures are rough estimates for the model weights alone. The actual VRAM demand is higher, as the runtime, KV cache, and context also consume memory.
| Model size | Q8_0 | Q6_K | Q4_K_M | Q3_K_M |
|---|---|---|---|---|
| 4B | ~5 GB | ~4 GB | ~3 GB | ~2.5 GB |
| 8B | ~9 GB | ~7 GB | ~5.5 GB | ~4.5 GB |
| 12B | ~13 GB | ~10 GB | ~8 GB | ~6.5 GB |
| 14B | ~16 GB | ~12 GB | ~9 GB | ~7.5 GB |
| 27B | ~30 GB | ~22 GB | ~17 GB | ~13 GB |
| 32B | ~36 GB | ~27 GB | ~20 GB | ~16 GB |
| 70B | ~80 GB | ~60 GB | ~42 GB | ~34 GB |
These figures are estimates rather than strict limits. Variances in model architectures and quantization formats can influence the actual footprint.
Model Compatibility with Different VRAM Capacities
| VRAM | Practical range | Current examples |
|---|---|---|
| 8 GB | Small models around 4B to 9B | Gemma 4 E4B, Qwen3.5 9B |
| 12 GB | Small to mid-sized models around 9B to 14B | Gemma 4 12B, Qwen3.5 9B |
| 16 GB | 12B to 27B with lower quantization | Gemma 4 26B-A4B, Qwen3.6 27B at Q4 |
| 24 GB | 27B to 35B at Q4 to Q6 | Qwen3.8 27B, Gemma 4 31B |
| 32 GB | 27B to 35B at higher quantization | Qwen3.8 27B, Gemma 4 31B |
| 48 GB | Large dense models at lower quantization | 70B-class models at Q3 to Q4 |
| 80 GB | Large dense models at higher quantization | 70B-class models at Q4 to Q6 |
These ranges apply to models whose weights can reside on the GPU. Large MoE models behave differently: while only a subset of parameters is active per token, the model must still store its full set of weights. Consequently, a model with 100B or more total parameters cannot fit into a 100B-sized VRAM budget simply because it utilizes fewer active parameters.
MoE Models
Mixture-of-Experts models consist of multiple parameter groups known as experts. Only specific experts are engaged for each token, potentially making inference more efficient than a dense model with an equivalent total parameter count.
However, inactive experts remain part of the model architecture. As a result, large MoE models can demand significantly more memory than their active parameter count implies. Very large models may necessitate multiple GPUs or offloading to system RAM.
Context Length Also Consumes VRAM
Model weights represent only a portion of the total memory requirement. The KV cache expands as context length increases, meaning that running the same model at a 64K context can demand substantially more VRAM than at 4K.
- Longer contexts require more VRAM.
- KV-cache precision impacts memory usage.
- Batch size and concurrent user load also increase memory consumption.
- Reserve some VRAM for the runtime rather than filling the GPU entirely with model weights.
Practical Advice
- Verify the actual size of the specific quantized model you intend to run.
- Do not rely solely on the model file size for VRAM requirements. Allow space for the KV cache and runtime.
- If a model does not fit entirely in VRAM, portions can be offloaded to system RAM, though this typically results in slower inference.
- For long-context or agentic workloads, allocate more VRAM than the model weights alone require.
- Multiple GPUs can be used to split a model if a single GPU lacks sufficient VRAM.
Run on DaDesktop
You do not need to purchase a GPU to run a local LLM. DaDesktop provides a cloud desktop equipped with the VRAM required, allowing you to run models directly without owning the hardware.
Select the VRAM tier that suits your model, load it, and begin usage. There is no setup, no hardware purchase, and no driver issues. View the available GPUs to see the options.