
ai-automation
Fix Ollama 2048 Context Truncation: num_ctx Guide
Why Ollama silently truncates prompts at 2048 tokens: the hidden num_ctx bug in Open WebUI, API, and Modelfiles, and how to fix RAG drops.
10m read
3 articles

Why Ollama silently truncates prompts at 2048 tokens: the hidden num_ctx bug in Open WebUI, API, and Modelfiles, and how to fix RAG drops.

We ran a full RAG pipeline on a single server using Ollama, Open-WebUI, and ChromaDB — no cloud, no API costs. Here is the exact setup and the mistakes we made.

How our team set up Open WebUI and Ollama with nomic-embed-text to index internal IT runbooks, network topologies, and incident notes 100% offline.