
ai-automation
How to Enable FP8 KV Cache in Ollama & vLLM (128K Context)
Break the VRAM wall in local LLMs. Learn the exact memory math, vLLM FP8 KV-cache flags, and Ollama Flash Attention settings to run 128K context on 24GB GPUs.
7m read
2 articles

Break the VRAM wall in local LLMs. Learn the exact memory math, vLLM FP8 KV-cache flags, and Ollama Flash Attention settings to run 128K context on 24GB GPUs.

Why does Ollama slow down or shift to CPU offload during long chats? Fix num_ctx memory degradation and force full GPU offload.