Local LLM VRAM & KV-Cache Scaling Dataset (2026 Lab Matrix)
Empirical memory profiling on our test bench across NVIDIA RTX 4090 (24GB), RTX 4070 Ti Super (16GB), and RTX 4060 (8GB). Measured under CUDA 12.6, Ollama v0.5.8, and llama.cpp b4120.
How to Cite This Benchmark
BibTeX / APA availablePraveenTechWorld Research Lab. (2026). Local LLM VRAM & KV-Cache Benchmark Matrix (v1.2).
https://www.praveentechworld.com/benchmarks/vram-llm1. Model Weight VRAM Footprint by Quantization
| Model | Parameters | FP16 Weight | Q8_0 Weight | Q4_K_M Weight | Min GPU Recommended |
|---|---|---|---|---|---|
| DeepSeek-R1-Distill-Qwen-7B | 7.6B | 15.2 GB | 8.1 GB | 4.7 GB | RTX 3060 / 4060 (8GB) |
| DeepSeek-R1-Distill-Qwen-14B | 14.8B | 29.6 GB | 15.6 GB | 9.0 GB | RTX 4060 Ti (16GB) |
| Phi-4 (Microsoft) | 14.7B | 29.4 GB | 15.4 GB | 8.9 GB | RTX 4060 Ti / 4070 (16GB) |
| Llama-3.3-70B-Instruct | 70.6B | 141.2 GB | 74.2 GB | 42.8 GB | Dual RTX 3090 / 4090 (48GB) |
2. KV-Cache Memory Scaling Overhead by Context Depth
VRAM is consumed not only by model weights, but by active attention key-value caches. Here is the measured overhead in gigabytes added to base model weights:
| Model Architecture | 4K Context | 8K Context | 16K Context | 32K Context |
|---|---|---|---|---|
| Qwen 2.5 / DeepSeek-R1 (7B) | 0.32 GB | 0.64 GB | 1.28 GB | 2.56 GB |
| Phi-4 (14B) | 0.45 GB | 0.90 GB | 1.80 GB | 3.60 GB |
| Llama 3.3 (70B with GQA) | 0.62 GB | 1.25 GB | 2.50 GB | 5.00 GB |