Commit Graph

2 Commits (c10d803a1264971b7432515162c62cda535fbbcd)

Author SHA1 Message Date
Amir Alexander Abdelbaki ea82ee70ad Add tools/ and CoreSystemConfig.json — one source of truth for every build
Installation was six scripts each carrying its own copy of the container host's
IP, three that had to agree on IDENTITY_TOKEN, and every service URL typed by
hand with a port in it. Any one could be wrong, and the symptom was always the
same and always late: an image that boots fine and then can't reach something,
found after a 40-minute build and a reboot.

Two properties fix that class of bug:

- Nothing is written twice. No script in tools/ contains an IP, port or token.
- Anything derivable is derived. You give the subnet prefix once and one last
  octet per host; every address and service URL is computed from those.

THE TWINNED PAIR. container_host.ip_last_octet 12 and llm_host 13 mean the
container host's OLLAMA_HOST *is* http://<prefix>.13:11434 — computed in the
same build, not typed into two files and kept in sync. Move the LLM host to .21
and the container host's Ollama URL follows; change the subnet and both halves
move along with every kiosk's URLs. Neither image can be built pointing at an
address the other isn't using. Both carry the same SMARTHOME_PAIR_ID (a hash of
the config's meaning, not its bytes) so two USB sticks can be checked against
each other later.

validate-config.py runs before every build and refuses to start on an error, so
a mistake costs seconds not an hour. It catches duplicate ports (including the
music_assistant/pantry_vision 8095 clash that Compose can't see because MA runs
network_mode:host — open decision #31), both hosts on one address, duplicate
hostnames across kiosks and audio endpoints, placeholder tokens (checked before
the length check, so padding "changeme" to 32 chars doesn't pass), a private key
pasted where the public one goes, and a kiosk pointed at a disabled service.

build-all.sh is the normal entry point — the images are a set that has to agree
with itself, so building one is the exception. It builds the core pair, every
kiosk, and every audio endpoint including both architectures (amd64 live-build
ISO and arm64 rpi-image-gen img are different toolchains, not one image).

The two new host ISOs install unattended with everything burnt in, including
service env files generated from derived values — which permanently removes the
class of bug that had chores.env shipping IDENTITY_URL=http://127.0.0.1:8097.
setup-container-host.sh and setup-llm-host.sh now read every config value as
${VAR:-default} so the images configure them without editing.

That also makes every ISO a credential: Wi-Fi PSK, tokens, MQTT and HA
credentials are readable by anyone holding the stick. .gitignore covers the
filled-in CoreSystemConfig.json and build-output/.

Tested: 43 config validation/derivation checks and 44 builder checks against the
real code paths with only `lb` stubbed — every generated env file, preseed,
network config, first-boot unit and build stamp is verified, including that a
port collision refuses the build before writing anything. No ISO has been built;
`lb build` needs live-build, root and a long fetch. tools/README.md says so.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 13:22:25 +02:00
Amir Alexander Abdelbaki 564c4a801d Add hosts/llm-host (Ollama) and CalDAV integration notes
Two unchecked items from the README status list that were buildable in-repo
rather than blocked on hardware.

hosts/llm-host/ — Phase 3's LLM machine:
- Auto-detects gpu vs cpu tier (nvidia-smi must both exist AND succeed; a
  leftover driver package on a machine whose card was pulled satisfies only
  the first and would fail later at container start).
- Runs Ollama as a pinned container rather than curl|sh into a root shell,
  matching how everything else here is deployed. Deliberately does NOT install
  the GPU driver — the most hardware/kernel-specific step on that box.
- Sets OLLAMA_HOST=0.0.0.0 inside the container. Ollama binds loopback by
  default, which in Docker means the published port forwards to nothing and
  every caller sees a connection refused indistinguishable from "the host is
  off" — and since every consumer here is built to tolerate exactly that, it
  degrades silently. Same class of bug as chores' 127.0.0.1 env values.
- Takes a position on Ollama contention (open decision #4's resource half):
  MAX_LOADED_MODELS=1 so a 14B text model and a vision model swap predictably
  instead of thrashing VRAM or OOM-ing mid-request, NUM_PARALLEL=1 for
  predictable Assist latency, KEEP_ALIVE=30m so a household that talks to
  Assist a few times an hour isn't paying model-load cost every time.
- Documents that Ollama has NO authentication and its API can delete models,
  not just generate — added to network-integration.md's port table, since the
  network is the entire boundary.

docs/caldav-integration.md — Phase 8's notes:
- The four independent clients and their directions (digest-engine read-only,
  chores' busy-check read-only, trash-calendar create-only under a UID-prefix
  ownership invariant, HA's own bridge).
- Why they share one Nextcloud app password, and the two costs: rotation
  touches three env files plus HA and fails quietly, and the read-only
  invariant is a CODE property, not a permission boundary — an app password
  can't be scoped read-only or per-calendar, so the server would not catch a
  regression that started writing.
- The two traps worth knowing before debugging them: unexpanded recurrence
  reporting a meeting on the day it was created, and CALDAV_VERIFY_TLS=false.

Neither has been run — no Debian machine, no GPU, no live Nextcloud. The script
is syntax-checked and its generated compose validated as YAML for both tiers;
that is the whole of the testing, and both READMEs say so.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 12:50:46 +02:00