# FAILURES.md Append-only record of every failure encountered building the Mechanical Compiler environment. | | | |---|---| | Scope | All instances. Staging entries are marked `srv-b`. | | Rule | Append only. Never edit an entry except to add a `Resolution` line. | | Numbering | Sequential, never reused. See §0 on the renumbering. | --- ## 0. Why this file exists separately This is the primary input to provisioning automation. Every entry below is something a script written from `ENVIRONMENT.md` alone would have got wrong. `F-003` and `F-004` in particular would have produced a container that booted, reported success, and was quietly broken. The specification cannot anticipate them; only contact with a host reveals them. When the production automation is written, it should be read **before** `ENVIRONMENT.md`, not after. ### Renumbering note, 2026-08-15 The staging checkpoint and the reconciliation checkpoint were written independently and both allocated `F-008` and `F-009`. The reconciliation entries have been renumbered to `F-015` and `F-016`. Original phase-1 numbering is preserved because it is cited elsewhere. ### Entry format ``` ### F-nnn — one-line summary Host / container. Phase. **Observed:** what was seen, verbatim where possible. **Cause:** proven cause, or explicitly "unproven". **Correction:** the smallest change that fixed it. **Consequence:** what the specification or automation must do differently. ``` A failure with no `Consequence` line is either not yet understood or not worth recording. --- ## Phase 1 — initial container build ### F-001 — `arping` absent on the Proxmox host `srv-b`. **Observed:** `command -v arping` → not installed. **Cause:** Not part of the base PVE install. **Correction:** `apt install iputils-arping`. **Consequence:** Automation must install `iputils-arping` before duplicate address detection, or the DAD check silently does not happen. --- ### F-002 — invalid Proxmox inspection command `srv-b`. **Observed:** `pvesm config local-lvm` returned CLI usage text. **Cause:** Operator command error. Not a host fault. **Correction:** Read `/etc/pve/storage.cfg` directly. **Consequence:** Read storage configuration from the file, not from a subcommand that does not exist. --- ### F-003 — CT 101 booted with degraded systemd CT 101. **Observed:** ``` status=226/NAMESPACE Failed to set up mount namespacing: Permission denied systemd-logind.service failed systemd-networkd.service failed systemd-timedated.service failed systemd-networkd.socket failed ``` **Cause:** Debian 12 systemd requires mount namespacing that an unprivileged LXC denies without `nesting`. **Correction:** `features: nesting=1`. Deliberately tested alone rather than copying `nesting=1,keyctl=1` from CT 100 — `keyctl=1` proved unnecessary. **Consequence:** **Every Debian 12 unprivileged container needs `nesting=1`**, including ones running nothing but nginx. `keyctl=1` is needed only where Docker runs. Promoted into the specification. This is the clearest example of why the manual build was correct: the container started, appeared healthy at a glance, and was broken in four services. --- ### F-004 — locale configured but never generated CT 100, CT 101. **Observed:** `LANG=en_US.UTF-8` set, but `locale -a` listed only `C`, `C.utf8`, `POSIX`. `locale` emitted `Cannot set LC_*` warnings. **Cause:** Setting `LANG` does not generate a locale. The container template ships none. **Correction:** install `locales`, enable `en_US.UTF-8 UTF-8`, `locale-gen`, `update-locale`. **Consequence:** Specification must state the generation step, not just the value. Silent until something depends on collation or encoding. --- ### F-005 — locale repair exposed a pending upgrade backlog CT 100, CT 101. **Observed:** Installing `locales` pulled glibc-related packages and revealed 52 further pending upgrades. **Cause:** Template was not current. **Correction:** Full upgrade of each container, one at a time. Both reached Debian 12.15, zero pending, no reboot required. **Consequence:** Automation must `apt full-upgrade` immediately after first boot, before anything else is installed. --- ### F-006 — transient `curl` / `git` command-not-found CT 100. **Observed:** An egress test returned `curl: command not found` and `git: command not found`. Immediately afterward both packages were installed and current, `PATH` normal, both executables present and working. **Cause:** **Unproven.** Hypothesis only: the test ran during the F-005 upgrade, while `dpkg` had the binaries briefly unlinked mid-transaction. Adjacent in time, consistent with the symptom, not demonstrated. **Correction:** None required. Repeat tests with absolute paths passed. **Consequence:** Do not close. If it recurs, capture `dpkg` lock state at the moment of failure. --- ### F-007 — recursive `chown` failed on ext4 `lost+found` CT 100. **Observed:** `chown: cannot read directory '/var/lib/mechcomp/lost+found': Permission denied`. That directory is `nobody:nogroup`, mode `0700`. **Cause:** `/var/lib/mechcomp` is a real filesystem root, not a directory. `lost+found` is created by `mkfs` and is not ours to manage. **Correction:** Own explicit application paths only. Leave `lost+found` alone. **Consequence:** **Never `chown -R` a mount point root.** Enumerate the directories the application actually uses. Promoted into the specification. --- ### F-008 — Git "dubious ownership" CT 100. **Observed:** `fatal: detected dubious ownership in repository at '/var/www/mechcomp'` when verification ran as root. **Cause:** Repository cloned as `mechcomp`, inspected as `root`. **Correction:** Re-ran verification as the owning user. **No root `safe.directory` exception was added** — the correct fix, since the exception would have masked every future instance of the same mistake. **Consequence:** All repository operations run as the service user. --- ### F-009 — dependency manifests absent from the application repository CT 100. **Observed:** `requirements-base.txt` and `requirements-cad.txt` do not exist; the repository is `LICENSE` and `README.md` at commit `e85c4f4e`. **Cause:** Application implementation has not started. Not a provisioning failure. **Correction:** None. Do not invent pins. **Consequence:** Infrastructure and application acceptance are separately gated. Infrastructure may complete without the application existing. --- ## Phase 2 — architect corrections ### F-010 — placeholder PID file rewrite failed to substitute CT 100. **Observed:** An automated rewrite left the literal `$!` in the PID file. **Cause:** Quoting error in the rewriting command. **Correction:** Identified the real PID by inspection. **Consequence:** Superseded by F-019 — the placeholder should be a systemd unit and have no PID file at all. --- ### F-011 — root could not overwrite a `mechcomp`-owned file in `/tmp` CT 100. **Observed:** `Permission denied` writing a file owned by `mechcomp:mechcomp` in `/tmp`, as root. **Cause:** `fs.protected_regular = 2` with `/tmp` mode `1777`. Expected kernel behaviour, not a fault. **Correction:** Rewrote the file as the owning user. **Consequence:** Do not use `/tmp` for state shared across users. Moot once services run under systemd with `PrivateTmp=true`. --- ### F-012 — first request after nginx reload returned the Debian welcome page CT 101. **Observed:** `nginx -t` passed; the immediately following request returned the 615-byte Debian default page. All subsequent requests proxied correctly. **Cause:** **Unproven.** The obvious candidate — Debian's `default` site still enabled and acting as `default_server` — was tested and ruled out: `sites-enabled` contains only the project vhost. Remaining hypothesis is a graceful worker transition. Not demonstrated. **Correction:** None required. **Consequence:** Do not close. Independently, declare `default_server` explicitly on the project vhost so that adding a second server block cannot make catch-all behaviour depend on file ordering. --- ### F-013 — TLS test validated against an IP literal CT 101. **Observed:** `SSL: no alternative certificate subject name matches target host name '10.0.0.21'`. **Cause:** Test used an address; the certificate SAN contains a DNS name. Correct behaviour by both curl and OpenSSL. **Correction:** Test by hostname. **Consequence:** Acceptance checks must use the service FQDN. An IP-literal TLS test is always wrong unless the certificate carries an IP SAN. --- ### F-014 — CT 101 had no `mechcomp` group CT 101. **Observed:** Adding `sandor` to `mechcomp` failed; the group did not exist. **Cause:** Specification said the admin joins the `mechcomp` group in *both* containers, but the group is created as a side effect of creating the service user, which exists only in CT 100. **Correction:** Created the system group and added `sandor`. **Consequence:** Specification defect. Either create the group explicitly where it is required, or scope the group membership to CT 100. GID 996 now matches across both containers, which is harmless and mildly useful. --- ### F-015 — CT 101 verification lacked `curl` CT 101. *(Renumbered from a duplicate `F-008`.)* **Observed:** `curl: command not found`. **Cause:** Not in the base template; CT 101's package list did not include it. **Correction:** Installed `curl` and `ca-certificates`. **Consequence:** The proxy container needs its own minimal toolset. Do not assume packages installed in CT 100 exist in CT 101. --- ### F-016 — placeholder PID capture stored a literal `$!` CT 100. *(Renumbered from a duplicate `F-009`.)* **Observed:** PID file contained `$!` rather than a number. The listener itself was healthy. **Cause:** Quoting error. **Correction:** Identified PID `12102` by inspection. **Consequence:** Superseded by F-019. --- ## Phase 3 — network isolation ### F-017 — containers were provisioned on the home LAN `srv-b`, CT 100, CT 101. **Observed:** Both containers had `net0` on `vmbr0` at `10.0.0.20/24` and `10.0.0.21/24`, gateway `10.0.0.1` — the ISP router's network. Anything on the home LAN could reach them. **Cause:** **Specification defect in `ENVIRONMENT.md` revision 4.** The document placed containers on both bridges, giving each a management address on the LAN, on the unstated assumption that a LAN workstation would browse staging. That assumption was never confirmed and was wrong. **Correction:** Both containers moved to `vmbr1` only, gateway `10.20.0.1`. `net1` deleted. Host NAT extended for `10.20.0.0/24`. **Consequence:** Containers are on the portless service bridge only. `srv-b` is router and bastion; access arrives via WireGuard. Promoted into the specification. --- ### F-018 — removing the LAN interface did not remove LAN reachability CT 100, CT 101. **Observed:** After F-017, both containers still reached `10.0.0.1` successfully. **Cause:** Removing an interface removes an address, not a route. `ip_forward=1` plus the new `MASQUERADE -s 10.20.0.0/24 -o vmbr0` — added to give the containers internet access — also gave them the LAN, translated to the host's address. **Correction:** A `RETURN` in `nat POSTROUTING` ahead of both masquerades, and a `DROP` in `FORWARD`, both scoped `-s 10.20.0.0/24 -d 10.0.0.0/24`, plus an explicit `ACCEPT` for container-to-container traffic. **Consequence:** **Interface removal is not isolation.** Automation must assert the negative — that the LAN is unreachable — not merely that the interface is gone. Note that the containers still reach `srv-b` itself at `10.0.0.12`, because a container addressing its gateway takes the INPUT path and never enters FORWARD. That is required, not a leak. --- ### F-019 — placeholder backend did not survive reboot CT 100. **Observed:** After the F-017 reboot, nothing listened on `10.20.0.10:8770`. **Cause:** The placeholder was a bare foreground process with a PID file. No supervision. **Correction:** Pending — to be recreated as `mechcomp-placeholder.service`. **Consequence:** Anything a proof depends on must be supervised, or the proof expires silently at the next reboot and the following session diagnoses a proxy fault that does not exist. Applies to temporary scaffolding as much as to real services. --- ### F-020 — service FQDN mapping changed twice `srv-b`, CT 100, CT 101. **Observed:** `mechanical-compiler.dev.infra` was mapped to `10.20.0.10`, then corrected to `10.0.0.21`, then corrected again to `10.20.0.11`. **Cause:** Three different states, each correct for its moment. `10.20.0.10` was wrong — it pointed at the application rather than the proxy. `10.0.0.21` was right while CT 101 had a LAN address and workstation access was assumed. `10.20.0.11` became right once F-017 removed the LAN interface and access moved to WireGuard through `srv-b`. **Correction:** `10.20.0.11` on all three hosts. **Consequence:** Not an error in either direction — a value that tracked a topology decision. Recorded because the reasoning matters more than the value: **the service FQDN names the proxy, and the proxy has exactly one address.** --- ## Open, not closed | # | Status | |---|---| | F-006 | Cause unproven. Recurrence should capture `dpkg` lock state. | | F-012 | Cause unproven. Leading candidate ruled out by inspection. | | F-019 | Correction pending. |