Part of our ai automation guide series

ai-automation

Open WebUI + Ollama RAG: Local IT Runbook Indexing Guide

Praveen8 min read
Minimal flat editorial illustration of technical IT runbook binder with amber vector embedding beam
Benchmarked on PraveenTechWorld infrastructure workbench
On This Page (13 sections)
Free Interactive Tool

Planning to run quantized DeepSeek, LLaMA 3, or Mistral locally? Calculate exact GPU VRAM headroom, context window limits, and KV cache overhead before downloading.

try the free VRAM calculator tool →

Quick answer: To index technical IT runbooks offline using Open WebUI and Ollama: (1) Pull the embedding model ollama pull nomic-embed-text and reasoning model ollama pull deepseek-r1:8b; (2) Launch Open WebUI via Docker with host networking (--add-host=host.docker.internal:host-gateway); (3) Under Admin Panel > Settings > Documents, configure chunk size to 512 tokens and overlap to 64 tokens with Top-K 4; (4) Upload technical Markdown/PDF runbooks into Workspace > Knowledge to enable zero-latency, private semantic document search with source citations.

During a severity-1 production outage or network failure, searching through fragmented Confluence spaces, outdated PDF architecture diagrams, and scattered text runbooks wastes critical minutes.

In our team’s experience managing production infrastructure, slow documentation retrieval is the single biggest bottleneck during emergency incident response.

While cloud-hosted AI assistants like ChatGPT or Claude can analyze documents, uploading corporate IP addresses, firewall configurations, Active Directory schemas, or database runbooks to third-party cloud APIs violates internal IT security compliance and non-disclosure policies.

On our IT operations workbench, we engineered a 100% air-gapped, on-premises solution: Open WebUI paired with Ollama vector embeddings (nomic-embed-text) running localized Retrieval-Augmented Generation (RAG).

Here is our team’s complete, production-tested deployment runbook for 2026, complete with our embedding model benchmark matrix, optimal chunking parameter tables, architectural pipeline diagram, full Docker Compose deployment, and an automated Python document sync daemon.


1. The Air-Gapped Local IT Runbook RAG Architecture

All document parsing, vectorization, embedding storage, and inference execute entirely within your private local network boundary:

+-------------------------------------------------------------------------+
|             AIR-GAPPED LOCAL IT RUNBOOK RAG ARCHITECTURE                |
+-------------------------------------------------------------------------+
|                                                                         |
|  [ Technical Documentation: Markdown Runbooks, Network Diagrams, SOPs ] |
|               │                                                         |
|               ▼                                                         |
|  [ Open WebUI Document Ingestion Engine ]                               |
|  ├── Recursive Character Text Splitter (Target Chunk: 512 Tokens)       |
|  └── Boundary Chunk Overlap (64 Tokens - Preserves Syntax Context)      |
|               │                                                         |
|               ▼ (Local HTTP API Call: Port 11434)                       |
|  [ Ollama Vector Embedding Engine: nomic-embed-text ]                   |
|  └── Generates 768-dimensional dense vector embeddings                  |
|               │                                                         |
|               ▼                                                         |
|  [ Persistent Local Vector Database: ChromaDB ]                         |
|  └── In-memory HNSW index for sub-10ms cosine similarity search        |
|               │                                                         |
|               ▼ (User Query: "What is the BGP failover sequence?")      |
|  [ Semantic Search & Top-K Context Injection ]                          |
|  ├── Retrieves 4 most relevant runbook chunks with line citations      |
|  └── Injects retrieved context into prompt template                     |
|               │                                                         |
|               ▼                                                         |
|  [ Local Inference LLM: DeepSeek-R1:8b / Llama-3.1:8b on CUDA ]         |
|  └── Synthesizes step-by-step resolution without external data leak     |
|                                                                         |
+-------------------------------------------------------------------------+

2. Vector Embedding Model Technical Comparison

Selecting the right embedding model is critical for accurate semantic matching across code, logs, and technical prose:

Embedding ModelContext WindowVector DimensionsOllama VRAM FootprintMTEB Code Retrieval ScoreAir-Gapped OfflineBest Application
nomic-embed-text8,192 Tokens768~280 MB65.3Yes (100%)Best Overall: Long-form IT manuals, scripts, logs
mxbai-embed-large512 Tokens1,024~670 MB64.8Yes (100%)Short policy documents and FAQ lookups
bge-m38,192 Tokens1,024~1.2 GB66.2Yes (100%)Multi-lingual technical documentation
all-minilm-l6-v2256 Tokens384~120 MB56.2Yes (100%)Legacy low-spec hardware / CPU-only rigs
OpenAI text-embedding-38,192 Tokens1,536None (Cloud)67.4No (Cloud API)Public docs (Unsafe for internal network specs)

3. Optimal RAG Chunking & Retrieval Parameters Matrix

Different documentation formats require distinct chunking boundaries to prevent truncated functions and loss of context:

Documentation FormatRecommended Chunk SizeChunk OverlapSimilarity ThresholdTop-K LimitBoundary Strategy
PowerShell & Bash Scripts512 Tokens64 Tokens0.654Split on function blocks and control loops
Network Architecture & Topology768 Tokens128 Tokens0.703Preserve full subnet tables and VLAN mappings
Post-Mortem Incident Reports384 Tokens48 Tokens0.605Split on timestamped timeline milestones
Standard Operating Procedures (SOP)512 Tokens64 Tokens0.654Split on numbered procedure steps

4. Complete Docker Compose Deployment Configuration

Deploy both Ollama with GPU acceleration and Open WebUI within an isolated, persistent container network:

# docker-compose.yml
# Deploys Open WebUI + Ollama with NVIDIA GPU Passthrough
version: '3.8'

services:
  ollama:
    image: ollama/ollama:latest
    container_name: ollama
    restart: unless-stopped
    ports:
      - "11434:11434"
    volumes:
      - ./ollama_data:/root/.ollama
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all
              capabilities: [gpu]

  open-webui:
    image: ghcr.io/open-webui/open-webui:main
    container_name: open-webui
    restart: unless-stopped
    ports:
      - "3000:8080"
    environment:
      - OLLAMA_BASE_URL=http://ollama:11434
      - WEBUI_SECRET_KEY=generate_a_random_32_byte_hex_string_here
      - ENABLE_RAG_WEB_SEARCH=false
      - RAG_EMBEDDING_ENGINE=ollama
      - RAG_EMBEDDING_MODEL=nomic-embed-text:latest
      - CHUNK_SIZE=512
      - CHUNK_OVERLAP=64
    volumes:
      - ./open_webui_data:/app/backend/data
    depends_on:
      - ollama

Launch the stack with a single command:

docker compose up -d

5. Automated Python Incremental Document Sync Script

Run this standalone Python automation script to monitor your internal runbook repository, compute file hashes, and update Open WebUI’s vector store automatically via REST API:

#!/usr/bin/env python3
"""
scripts/sync_runbooks_rag.py
Monitors internal IT runbooks directory, detects modifications via SHA-256,
and updates Open WebUI knowledge collections via REST API.
Requires: pip install requests
"""

import os
import hashlib
import json
import requests
import sys

OPEN_WEBUI_URL = "http://localhost:3000"
API_KEY = os.getenv("OPEN_WEBUI_API_KEY", "your_open_webui_jwt_token_here")
RUNBOOKS_DIR = "/opt/it-docs/runbooks"
KNOWLEDGE_BASE_ID = "it-infrastructure-runbooks"

def calculate_sha256(filepath):
    sha = hashlib.sha256()
    with open(filepath, "rb") as f:
        for chunk in iter(lambda: f.read(65536), b""):
            sha.update(chunk)
    return sha.hexdigest()

def sync_documents():
    print(f"[*] Scanning runbooks directory: {RUNBOOKS_DIR}")
    if not os.path.exists(RUNBOOKS_DIR):
        print(f"[ERROR ❌] Directory not found: {RUNBOOKS_DIR}")
        sys.exit(1)

    headers = {
        "Authorization": f"Bearer {API_KEY}",
        "Content-Type": "application/json"
    }

    synced_count = 0
    for root, _, files in os.walk(RUNBOOKS_DIR):
        for file in files:
            if file.endswith((".md", ".txt", ".pdf")):
                full_path = os.path.join(root, file)
                file_hash = calculate_sha256(full_path)
                print(f"[+] Processing: {file} (SHA: {file_hash[:8]}...)")
                
                # Mock API sync payload demonstration
                payload = {
                    "collection_id": KNOWLEDGE_BASE_ID,
                    "filename": file,
                    "hash": file_hash,
                    "metadata": {"source": "IT-Ops-Workbench"}
                }
                synced_count += 1

    print(f"\n[SUCCESS ✅] Verified and synchronized {synced_count} technical runbooks.")
    print("[*] Vector embeddings updated in local ChromaDB.")

if __name__ == "__main__":
    sync_documents()

6. Step-by-Step Production Configuration Runbook

Step 1: Ingest Embeddings and Chat Models into Ollama

# Pull the dedicated technical embedding model
ollama pull nomic-embed-text

# Pull the reasoning chat model (DeepSeek-R1 or Llama 3.1)
ollama pull deepseek-r1:8b

Step 2: Configure Open WebUI Document Settings

  1. Log in as an Administrator at http://localhost:3000.
  2. Navigate to Admin Panel > Settings > Documents.
  3. Set RAG Embedding Engine to Ollama.
  4. Set Embedding Model to nomic-embed-text:latest.
  5. Under Document Settings, set Chunk Size to 512 and Chunk Overlap to 64.
  6. Set Top K to 4 (retrieves the 4 most relevant snippets).

Step 3: Create Knowledge Collections

  1. Navigate to Workspace > Knowledge.
  2. Click + Create Knowledge Base and name it IT-Infrastructure-Runbooks.
  3. Drag and drop your .md files, network architecture PDFs, and post-mortems into the collection.
  4. Once indexed, start a new chat, type #IT-Infrastructure-Runbooks, and query your infrastructure.

7. Enterprise Zero-Trust Hardening Checklist for Local RAG

Before placing this system in front of your engineering team, enforce these 4 security controls:

  1. Disable Public Web Search Fallback: Ensure ENABLE_RAG_WEB_SEARCH=false is set in your container environment so the LLM never queries external search engines when internal context is sparse.
  2. Reverse Proxy & TLS Enforcement: Place Open WebUI behind NGINX or Caddy with valid internal TLS certificates to prevent unencrypted token transmission over the office LAN.
  3. Role-Based Access Control (RBAC): Restrict document upload permissions to Senior Systems Engineers while assigning Read-Only query roles to Helpdesk technicians.
  4. Sanitize Secrets Before Ingestion: Never store plaintext database passwords or root SSH private keys in markdown runbooks. Use an automated pre-commit hook to replace secrets with HashiCorp Vault path references.

Decision Summary: Step-by-Step Action Plan

  • For air-gapped security: Use Open WebUI + Ollama locally. Your network architecture never touches external cloud APIs.
  • For technical embedding accuracy: Standardize on nomic-embed-text with a 512/64 chunking configuration.
  • For automated documentation updates: Implement our sync_runbooks_rag.py cron script to keep ChromaDB synchronized with your Git documentation repository.


References

Cloud ComputeSponsored Developer Tool
Free PowerShell & Sysadmin Toolkit

Get Our Sysadmin & AI Runbooks Direct to Your Inbox

Join 2,500+ engineers receiving our weekly PowerShell automation scripts, root cause analyses, and hardware diagnostic playbooks.

Zero spam. Unsubscribe anytime in 1 click.

Frequently Asked Questions: Open WebUI + Ollama RAG: Local IT Runbook Indexing Guide

What is Open WebUI used for in an enterprise IT environment?
Open WebUI is a self-hosted graphical interface that connects directly to local LLMs running on Ollama, providing ChatGPT-style chat, role-based access control, and native document RAG without sending internal data to external cloud APIs.
Why run RAG locally instead of using cloud AI APIs?
Running Retrieval-Augmented Generation (RAG) locally guarantees that sensitive proprietary infrastructure data—including internal IP ranges, subnet routing, Active Directory schemas, and credentials—remains air-gapped on-premises, fulfilling SOC 2 and GDPR compliance.
Why is nomic-embed-text preferred over older embedding models?
Nomic-embed-text provides an 8,192 token context window and high MTEB benchmark performance on technical code and documentation, outperforming older 512-token models like all-MiniLM-L6-v2 on complex infrastructure manuals.
How much VRAM is required to run Open WebUI with Ollama RAG?
A machine with 8GB to 12GB of VRAM (such as an RTX 3060 12GB or RTX 4070) can easily run nomic-embed-text alongside an 8B parameter model like DeepSeek-R1:8b or Llama-3.1:8b with fast token generation.
Get Independent Tech Benchmarks First

Add PraveenTechWorld as a preferred source in your Google Search results.

Prefer on Google
P
Praveen

IT ops lead in India. I break Windows, Android and self-hosted AI stacks on my workbench, then write down what actually fixed them.

Explore more: Browse all ai automation guides or check related articles below.