#!/usr/bin/env bash # # LLM Host Setup — Ollama, docs/project-plan.md Phase 3 # Target: Debian 12 (Bookworm), a SEPARATE physical machine from the container host. # # Stands up the one thing this host exists for: an Ollama server on the LAN that # everything else in this project calls for inference — # - Home Assistant's Ollama conversation integration (Assist, AI Task) [Phase 3/4] # - digest-engine's quarter-daily synthesis + counter-run [Phase 12] # - pantry-vision's grocery-item recognition (vision model) [Phase 17] # - chores' bin/dishes/litter checks + reminder phrasing [Phase 20] # # THE GUARDRAIL THIS HOST IS BUILT AROUND: **everything above must degrade to # "unavailable", never to "broken", when this machine is off.** docs/project-plan.md's # testing checklist states it directly — "Does the reactive path (presence → light on) # work with the LLM host powered off? (It must.)" — and every consumer above already # falls back to a deterministic template, a plain lookup table, or simply skipping the # run. Nothing here should ever become load-bearing for the reactive smart-home # baseline. That is why this is a separate host at all: so it CAN be off. # # Two hardware tiers, auto-detected (override with TIER below): # gpu — NVIDIA card present. Runs the 14B-class model per Phase 3. # cpu — no GPU. Runs a 7B/3B model. Slower (single-digit tok/s), enough to validate # the whole pipeline before spending money on a card, per docs/components.md's # "Fallback: skip GPU" line. # # Run as: sudo ./setup-llm-host.sh # # EDIT THE VARIABLES BELOW BEFORE RUNNING. set -euo pipefail # --------------------------------------------------------------------------- # CONFIGURATION # # Every value here is `${VAR:-default}`, so anything already in the environment wins — # which is how the ISO from tools/build-llm-host-iso.sh configures this script without # editing it (its first-boot unit sources /opt/llm-host/llm-host.env, generated from # CoreSystemConfig.json). Edit these directly ONLY for a standalone run. # --------------------------------------------------------------------------- BASE_DIR="${BASE_DIR:-/opt/llm-host}" # Config + model storage. Models are BIG (a 14B # Q4 model is ~9GB, a vision model another 5-8GB) # — make sure this lives on a disk with room. TIER="${TIER:-auto}" # auto | gpu | cpu OLLAMA_PORT="${OLLAMA_PORT:-11434}" # Ollama's own default. # Models to pull, per tier. Phase 3 specifies Qwen2.5-14B-Instruct (GPU) or 7B/3B # (CPU). Tags are Ollama library names — `ollama list` on a real host to confirm what # you actually ended up with, since library tags do get renamed upstream. GPU_TEXT_MODEL="${GPU_TEXT_MODEL:-qwen2.5:14b-instruct}" CPU_TEXT_MODEL="${CPU_TEXT_MODEL:-qwen2.5:7b-instruct}" # The vision model, for pantry-vision (grocery items) and chores (bin/dishes/litter). # NOT A CONSIDERED CHOICE — `llava` is the default those services already ship with, # and docs/project-plan.md open decision #18 flags the pick as unmade and completely # unbenchmarked. Treat this as "something to measure", not "the answer": if grocery # recognition is too slow or too wrong to be usable, this is the first knob to turn # (qwen2.5vl and moondream are the obvious alternatives to try). VISION_MODEL="${VISION_MODEL:-llava}" PULL_VISION_MODEL="${PULL_VISION_MODEL:-true}" # false to skip — saves several GB if you're not # running pantry-vision/chores camera checks yet. # --- Contention between interactive and batch callers ------------------------------ # The real scheduling problem on one GPU (docs/project-plan.md open decision #4): # Assist is INTERACTIVE (a person is standing there waiting), while digest-engine is # BATCH (every 6h, nobody watching) and the vision callers are occasional but want a # DIFFERENT model resident. Defaults below optimise for the interactive case, because # that's the one where latency is felt: # # KEEP_ALIVE — how long a model stays resident after its last request. Ollama's own # default is 5m, which means a household that talks to Assist a few times an hour # pays the model-load cost almost every time. 30m keeps it warm through normal use. OLLAMA_KEEP_ALIVE="${OLLAMA_KEEP_ALIVE:-30m}" # MAX_LOADED_MODELS — how many distinct models may be resident at once. **1 is # deliberate on a single consumer GPU**: a 14B text model and a vision model do not # fit together in 8-12GB, and letting Ollama try produces VRAM thrash or an OOM # mid-request rather than an honest swap. 1 means "swap predictably, pay the reload # cost when the vision model is actually needed." Raise it only if you have the VRAM # to hold both and have checked that you do. OLLAMA_MAX_LOADED_MODELS="${OLLAMA_MAX_LOADED_MODELS:-1}" # NUM_PARALLEL — concurrent requests served per loaded model. 1 keeps latency # predictable for whoever is speaking to Assist; higher trades that for throughput # nothing in this project currently needs. OLLAMA_NUM_PARALLEL="${OLLAMA_NUM_PARALLEL:-1}" # --------------------------------------------------------------------------- # End of configuration # --------------------------------------------------------------------------- log() { echo -e "\n\033[1;34m==>\033[0m $*"; } warn() { echo -e "\033[1;33m[warn]\033[0m $*" >&2; } die() { echo -e "\033[1;31m[error]\033[0m $*" >&2; exit 1; } [[ $EUID -eq 0 ]] || die "Run this with sudo." # --------------------------------------------------------------------------- # Tier detection # --------------------------------------------------------------------------- detect_tier() { if [[ "$TIER" != "auto" ]]; then echo "$TIER" return fi # nvidia-smi existing AND succeeding are different things — a leftover driver # package on a machine whose card has been pulled would satisfy `command -v` alone # and send us down the GPU path to fail later at container start. if command -v nvidia-smi >/dev/null 2>&1 && nvidia-smi >/dev/null 2>&1; then echo "gpu" else echo "cpu" fi } # --------------------------------------------------------------------------- # Docker — same install path as tools/setup-container-host.sh. # Ollama is run as a container rather than natively installed for the same reason # everything else in this project is: no `curl | sh` into a root shell, a pinned # image, and an uninstall that's `docker rm`. The native installer is a legitimate # alternative (see README.md); it is not the default here. # --------------------------------------------------------------------------- install_docker() { if command -v docker >/dev/null 2>&1; then log "Docker already installed — skipping" return fi log "Installing Docker" apt-get update apt-get install -y ca-certificates curl gnupg install -m 0755 -d /etc/apt/keyrings curl -fsSL https://download.docker.com/linux/debian/gpg \ | gpg --dearmor -o /etc/apt/keyrings/docker.gpg chmod a+r /etc/apt/keyrings/docker.gpg echo "deb [arch=$(dpkg --print-architecture) signed-by=/etc/apt/keyrings/docker.gpg] \ https://download.docker.com/linux/debian $(. /etc/os-release && echo "$VERSION_CODENAME") stable" \ > /etc/apt/sources.list.d/docker.list apt-get update apt-get install -y docker-ce docker-ce-cli containerd.io docker-buildx-plugin docker-compose-plugin } # --------------------------------------------------------------------------- # NVIDIA Container Toolkit — only what lets Docker see the GPU. It does NOT install # the driver: a working `nvidia-smi` is a prerequisite this script checks for rather # than tries to satisfy, because driver installation is the single most # hardware/kernel-specific step on this host and silently picking a driver version # for someone is a good way to produce an unbootable machine. # --------------------------------------------------------------------------- install_nvidia_toolkit() { if ! nvidia-smi >/dev/null 2>&1; then die "TIER=gpu but nvidia-smi doesn't work. Install the NVIDIA driver first (Debian: enable non-free-firmware, then 'apt install nvidia-driver firmware-misc-nonfree', reboot, confirm 'nvidia-smi' prints your card), or set TIER=cpu to run without a GPU." fi if command -v nvidia-ctk >/dev/null 2>&1; then log "NVIDIA Container Toolkit already installed — skipping" else log "Installing NVIDIA Container Toolkit" curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey \ | gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg curl -fsSL https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list \ | sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' \ > /etc/apt/sources.list.d/nvidia-container-toolkit.list apt-get update apt-get install -y nvidia-container-toolkit fi nvidia-ctk runtime configure --runtime=docker systemctl restart docker } # --------------------------------------------------------------------------- # Compose file # --------------------------------------------------------------------------- write_compose() { local tier="$1" local gpu_block="" if [[ "$tier" == "gpu" ]]; then gpu_block=" deploy: resources: reservations: devices: - driver: nvidia count: all capabilities: [gpu]" fi mkdir -p "$BASE_DIR/models" cat > "$BASE_DIR/docker-compose.yml" </dev/null 2>&1; then return 0 fi sleep 2 done die "Ollama didn't come up within 2 minutes. Check: docker logs ollama" } # --------------------------------------------------------------------------- # Main # --------------------------------------------------------------------------- TIER_RESOLVED="$(detect_tier)" log "Tier: ${TIER_RESOLVED}$([[ "$TIER" == "auto" ]] && echo " (auto-detected)")" if [[ "$TIER_RESOLVED" == "cpu" ]]; then warn "Running CPU-only. Expect single-digit tokens/sec — usable for validating the whole pipeline end to end, slow enough to be annoying for live Assist. This is docs/components.md's deliberate 'skip the GPU for now' fallback, not a failure." fi install_docker [[ "$TIER_RESOLVED" == "gpu" ]] && install_nvidia_toolkit write_compose "$TIER_RESOLVED" log "Starting Ollama" docker compose -f "$BASE_DIR/docker-compose.yml" up -d wait_for_ollama if [[ "$TIER_RESOLVED" == "gpu" ]]; then TEXT_MODEL="$GPU_TEXT_MODEL" else TEXT_MODEL="$CPU_TEXT_MODEL" fi # A failed pull is deliberately not fatal — the server is up and useful, and the # likeliest cause is a renamed upstream tag, which a human fixes in seconds and a # script cannot guess at. pull_model "$TEXT_MODEL" || true if [[ "$PULL_VISION_MODEL" == "true" ]]; then pull_model "$VISION_MODEL" || true fi HOST_IP="$(hostname -I | awk '{print $1}')" log "Installed models:" docker exec ollama ollama list || warn "Could not list models" cat < Devices & Services -> Add Integration -> Ollama, URL http://${HOST_IP}:${OLLAMA_PORT}, model ${TEXT_MODEL}. Then Settings -> Voice assistants -> your Assist pipeline -> Conversation agent. 2. On the CONTAINER host, in each service's env file: digest-engine.env : OLLAMA_HOST=http://${HOST_IP}:${OLLAMA_PORT} OLLAMA_MODEL=${TEXT_MODEL} pantry-vision.env : OLLAMA_HOST=http://${HOST_IP}:${OLLAMA_PORT} OLLAMA_VISION_MODEL=${VISION_MODEL} chores.env : OLLAMA_HOST=http://${HOST_IP}:${OLLAMA_PORT} OLLAMA_VISION_MODEL=${VISION_MODEL} OLLAMA_TEXT_MODEL=${TEXT_MODEL} (optional, reminder phrasing) NOTE the full scheme+port form: those services build URLs by string concatenation, so a bare IP will not work. 3. THEN DO THE THING THIS HOST EXISTS TO SURVIVE — power it off and confirm the reactive baseline still works: - presence -> light on/off still fires (Phase 2 automations, no LLM in the loop) - the door panel still shows weather/who's-home - chores still nudges (plain template instead of LLM-phrased wording) - digest-engine skips its run rather than erroring the timer docs/project-plan.md's testing checklist calls this out explicitly. If anything above BREAKS rather than degrading, that's a bug in the consumer, not here. 4. SECURITY: Ollama has no authentication of any kind, and its API can pull and DELETE models, not just generate. Anyone who can reach :${OLLAMA_PORT} can do all of that. Keep this host on the smart-home VLAN, never port-forwarded — see docs/network-integration.md. EOF