Files
mechanical-compiler/docs/FAILURES.md
T
2026-08-17 08:25:14 -04:00

545 lines
23 KiB
Markdown

# FAILURES.md
Append-only record of every failure encountered building the Mechanical
Compiler environment.
| | |
|---|---|
| Scope | All instances. Staging entries are marked `srv-b`. |
| Updated | 2026-08-17, after work order 002 |
| Rule | Append only. Never edit an entry except to add a `Resolution` line. |
| Numbering | Sequential, never reused. See §0 on the renumbering. |
---
## 0. Why this file exists separately
This is the primary input to provisioning automation.
Every entry below is something a script written from `ENVIRONMENT.md` alone
would have got wrong. `F-003` and `F-004` in particular would have produced a
container that booted, reported success, and was quietly broken. The
specification cannot anticipate them; only contact with a host reveals them.
When the production automation is written, it should be read **before**
`ENVIRONMENT.md`, not after.
### Renumbering note, 2026-08-15
The staging checkpoint and the reconciliation checkpoint were written
independently and both allocated `F-008` and `F-009`. The reconciliation
entries have been renumbered to `F-015` and `F-016`. Original phase-1 numbering
is preserved because it is cited elsewhere.
### Entry format
```
### F-nnn — one-line summary
Host / container. Phase.
**Observed:** what was seen, verbatim where possible.
**Cause:** proven cause, or explicitly "unproven".
**Correction:** the smallest change that fixed it.
**Consequence:** what the specification or automation must do differently.
```
A failure with no `Consequence` line is either not yet understood or not worth
recording.
---
## Phase 1 — initial container build
### F-001 — `arping` absent on the Proxmox host
`srv-b`.
**Observed:** `command -v arping` → not installed.
**Cause:** Not part of the base PVE install.
**Correction:** `apt install iputils-arping`.
**Consequence:** Automation must install `iputils-arping` before duplicate
address detection, or the DAD check silently does not happen.
---
### F-002 — invalid Proxmox inspection command
`srv-b`.
**Observed:** `pvesm config local-lvm` returned CLI usage text.
**Cause:** Operator command error. Not a host fault.
**Correction:** Read `/etc/pve/storage.cfg` directly.
**Consequence:** Read storage configuration from the file, not from a
subcommand that does not exist.
---
### F-003 — CT 101 booted with degraded systemd
CT 101.
**Observed:**
```
status=226/NAMESPACE
Failed to set up mount namespacing: Permission denied
systemd-logind.service failed
systemd-networkd.service failed
systemd-timedated.service failed
systemd-networkd.socket failed
```
**Cause:** Debian 12 systemd requires mount namespacing that an unprivileged
LXC denies without `nesting`.
**Correction:** `features: nesting=1`. Deliberately tested alone rather than
copying `nesting=1,keyctl=1` from CT 100 — `keyctl=1` proved unnecessary.
**Consequence:** **Every Debian 12 unprivileged container needs
`nesting=1`**, including ones running nothing but nginx. `keyctl=1` is needed
only where Docker runs. Promoted into the specification.
This is the clearest example of why the manual build was correct: the container
started, appeared healthy at a glance, and was broken in four services.
---
### F-004 — locale configured but never generated
CT 100, CT 101.
**Observed:** `LANG=en_US.UTF-8` set, but `locale -a` listed only `C`,
`C.utf8`, `POSIX`. `locale` emitted `Cannot set LC_*` warnings.
**Cause:** Setting `LANG` does not generate a locale. The container template
ships none.
**Correction:** install `locales`, enable `en_US.UTF-8 UTF-8`, `locale-gen`,
`update-locale`.
**Consequence:** Specification must state the generation step, not just the
value. Silent until something depends on collation or encoding.
---
### F-005 — locale repair exposed a pending upgrade backlog
CT 100, CT 101.
**Observed:** Installing `locales` pulled glibc-related packages and revealed
52 further pending upgrades.
**Cause:** Template was not current.
**Correction:** Full upgrade of each container, one at a time. Both reached
Debian 12.15, zero pending, no reboot required.
**Consequence:** Automation must `apt full-upgrade` immediately after first
boot, before anything else is installed.
---
### F-006 — transient `curl` / `git` command-not-found
CT 100.
**Observed:** An egress test returned `curl: command not found` and
`git: command not found`. Immediately afterward both packages were installed
and current, `PATH` normal, both executables present and working.
**Cause:** **Unproven.** Hypothesis only: the test ran during the F-005
upgrade, while `dpkg` had the binaries briefly unlinked mid-transaction.
Adjacent in time, consistent with the symptom, not demonstrated.
**Correction:** None required. Repeat tests with absolute paths passed.
**Consequence:** Do not close. If it recurs, capture `dpkg` lock state at the
moment of failure.
---
### F-007 — recursive `chown` failed on ext4 `lost+found`
CT 100.
**Observed:** `chown: cannot read directory '/var/lib/mechcomp/lost+found':
Permission denied`. That directory is `nobody:nogroup`, mode `0700`.
**Cause:** `/var/lib/mechcomp` is a real filesystem root, not a directory.
`lost+found` is created by `mkfs` and is not ours to manage.
**Correction:** Own explicit application paths only. Leave `lost+found` alone.
**Consequence:** **Never `chown -R` a mount point root.** Enumerate the
directories the application actually uses. Promoted into the specification.
---
### F-008 — Git "dubious ownership"
CT 100.
**Observed:** `fatal: detected dubious ownership in repository at
'/var/www/mechcomp'` when verification ran as root.
**Cause:** Repository cloned as `mechcomp`, inspected as `root`.
**Correction:** Re-ran verification as the owning user. **No root
`safe.directory` exception was added** — the correct fix, since the exception
would have masked every future instance of the same mistake.
**Consequence:** All repository operations run as the service user.
---
### F-009 — dependency manifests absent from the application repository
CT 100.
**Observed:** `requirements-base.txt` and `requirements-cad.txt` do not exist;
the repository is `LICENSE` and `README.md` at commit `e85c4f4e`.
**Cause:** Application implementation has not started. Not a provisioning
failure.
**Correction:** None. Do not invent pins.
**Consequence:** Infrastructure and application acceptance are separately
gated. Infrastructure may complete without the application existing.
---
## Phase 2 — architect corrections
### F-010 — placeholder PID file rewrite failed to substitute
CT 100.
**Observed:** An automated rewrite left the literal `$!` in the PID file.
**Cause:** Quoting error in the rewriting command.
**Correction:** Identified the real PID by inspection.
**Consequence:** Superseded by F-019 — the placeholder should be a systemd
unit and have no PID file at all.
---
### F-011 — root could not overwrite a `mechcomp`-owned file in `/tmp`
CT 100.
**Observed:** `Permission denied` writing a file owned by `mechcomp:mechcomp`
in `/tmp`, as root.
**Cause:** `fs.protected_regular = 2` with `/tmp` mode `1777`. Expected kernel
behaviour, not a fault.
**Correction:** Rewrote the file as the owning user.
**Consequence:** Do not use `/tmp` for state shared across users. Moot once
services run under systemd with `PrivateTmp=true`.
---
### F-012 — first request after nginx reload returned the Debian welcome page
CT 101.
**Observed:** `nginx -t` passed; the immediately following request returned the
615-byte Debian default page. All subsequent requests proxied correctly.
**Cause:** **Unproven.** The obvious candidate — Debian's `default` site still
enabled and acting as `default_server` — was tested and ruled out:
`sites-enabled` contains only the project vhost. Remaining hypothesis is a
graceful worker transition. Not demonstrated.
**Correction:** None required.
**Consequence:** Do not close. Independently, declare `default_server`
explicitly on the project vhost so that adding a second server block cannot
make catch-all behaviour depend on file ordering.
---
### F-013 — TLS test validated against an IP literal
CT 101.
**Observed:** `SSL: no alternative certificate subject name matches target
host name '10.0.0.21'`.
**Cause:** Test used an address; the certificate SAN contains a DNS name.
Correct behaviour by both curl and OpenSSL.
**Correction:** Test by hostname.
**Consequence:** Acceptance checks must use the service FQDN. An IP-literal
TLS test is always wrong unless the certificate carries an IP SAN.
---
### F-014 — CT 101 had no `mechcomp` group
CT 101.
**Observed:** Adding `sandor` to `mechcomp` failed; the group did not exist.
**Cause:** Specification said the admin joins the `mechcomp` group in *both*
containers, but the group is created as a side effect of creating the service
user, which exists only in CT 100.
**Correction:** Created the system group and added `sandor`.
**Consequence:** Specification defect. Either create the group explicitly where
it is required, or scope the group membership to CT 100. GID 996 now matches
across both containers, which is harmless and mildly useful.
---
### F-015 — CT 101 verification lacked `curl`
CT 101. *(Renumbered from a duplicate `F-008`.)*
**Observed:** `curl: command not found`.
**Cause:** Not in the base template; CT 101's package list did not include it.
**Correction:** Installed `curl` and `ca-certificates`.
**Consequence:** The proxy container needs its own minimal toolset. Do not
assume packages installed in CT 100 exist in CT 101.
---
### F-016 — placeholder PID capture stored a literal `$!`
CT 100. *(Renumbered from a duplicate `F-009`.)*
**Observed:** PID file contained `$!` rather than a number. The listener itself
was healthy.
**Cause:** Quoting error.
**Correction:** Identified PID `12102` by inspection.
**Consequence:** Superseded by F-019.
---
## Phase 3 — network isolation
### F-017 — containers were provisioned on the home LAN
`srv-b`, CT 100, CT 101.
**Observed:** Both containers had `net0` on `vmbr0` at `10.0.0.20/24` and
`10.0.0.21/24`, gateway `10.0.0.1` — the ISP router's network. Anything on the
home LAN could reach them.
**Cause:** **Specification defect in `ENVIRONMENT.md` revision 4.** The
document placed containers on both bridges, giving each a management address on
the LAN, on the unstated assumption that a LAN workstation would browse
staging. That assumption was never confirmed and was wrong.
**Correction:** Both containers moved to `vmbr1` only, gateway `10.20.0.1`.
`net1` deleted. Host NAT extended for `10.20.0.0/24`.
**Consequence:** Containers are on the portless service bridge only. `srv-b` is
router and bastion; access arrives via WireGuard. Promoted into the
specification.
---
### F-018 — removing the LAN interface did not remove LAN reachability
CT 100, CT 101.
**Observed:** After F-017, both containers still reached `10.0.0.1`
successfully.
**Cause:** Removing an interface removes an address, not a route.
`ip_forward=1` plus the new `MASQUERADE -s 10.20.0.0/24 -o vmbr0` — added to
give the containers internet access — also gave them the LAN, translated to the
host's address.
**Correction:** A `RETURN` in `nat POSTROUTING` ahead of both masquerades, and
a `DROP` in `FORWARD`, both scoped `-s 10.20.0.0/24 -d 10.0.0.0/24`, plus an
explicit `ACCEPT` for container-to-container traffic.
**Consequence:** **Interface removal is not isolation.** Automation must assert
the negative — that the LAN is unreachable — not merely that the interface is
gone. Note that the containers still reach `srv-b` itself at `10.0.0.12`,
because a container addressing its gateway takes the INPUT path and never
enters FORWARD. That is required, not a leak.
---
### F-019 — placeholder backend did not survive reboot
CT 100.
**Observed:** After the F-017 reboot, nothing listened on `10.20.0.10:8770`.
**Cause:** The placeholder was a bare foreground process with a PID file. No
supervision.
**Correction:** Recreated as `mechcomp-placeholder.service`.
**Consequence:** Anything a proof depends on must be supervised, or the proof
expires silently at the next reboot and the following session diagnoses a
proxy fault that does not exist. Applies to temporary scaffolding as much as to
real services.
**Resolution (2026-08-16):** Installed `/usr/local/libexec/mechcomp-placeholder.py`
and `/etc/systemd/system/mechcomp-placeholder.service`, enabled at
`multi-user.target`, running as `mechcomp:mechcomp`, reading
`/etc/mechcomp/mechcomp.env`, binding `10.20.0.10:8770`, with `Restart=on-failure`
and the standard hardening set. CT 100 was rebooted: the unit restarted
automatically, the listener returned on the service address only, nginx proxied
successfully, `X-Forwarded-Proto: https` was re-observed at the backend, and the
container settled to `running` with zero failed units.
---
### F-020 — service FQDN mapping changed twice
`srv-b`, CT 100, CT 101.
**Observed:** `mechanical-compiler.dev.infra` was mapped to `10.20.0.10`, then
corrected to `10.0.0.21`, then corrected again to `10.20.0.11`.
**Cause:** Three different states, each correct for its moment. `10.20.0.10`
was wrong — it pointed at the application rather than the proxy. `10.0.0.21`
was right while CT 101 had a LAN address and workstation access was assumed.
`10.20.0.11` became right once F-017 removed the LAN interface and access moved
to WireGuard through `srv-b`.
**Correction:** `10.20.0.11` on all three hosts.
**Consequence:** Not an error in either direction — a value that tracked a
topology decision. Recorded because the reasoning matters more than the value:
**the service FQDN names the proxy, and the proxy has exactly one address.**
---
### F-021 — deleting the Proxmox interface left a stale guest `eth1`
CT 100, CT 101. Network isolation.
**Observed:** After the F-017 topology change and reboot, both containers
reported `systemctl is-system-running` → `degraded`, with
`networking.service` and `ifupdown-wait-online.service` failed. Both guests'
`/etc/network/interfaces` still carried an `auto eth1` / static `iface eth1`
stanza although neither had an `eth1` link. The boot journal on both:
```
Cannot find device "eth1"
ifup: failed to bring up eth1
```
**Cause:** **Proven.** `pct set --delete net1` removes the LXC interface but
does not remove the stanza already written into the guest. `ifup -a` therefore
exited non-zero at boot even though `eth0` came up correctly.
`ifupdown-wait-online` failed as a consequence, not independently.
`systemd-networkd` was investigated and explicitly ruled out — its units were
disabled on both containers.
**Correction:** Preserved a pre-correction copy, removed only the stale `eth1`
stanza on each guest, restarted the affected units. Both containers then
rebooted: the stanza did not return, both units succeeded, both settled to
`running` with zero failed units.
**Consequence:** **Removing a container interface is not proof that the guest
converged.** After any topology mutation, automation must inspect the guest
interface file and assert `systemctl is-system-running = running` with zero
failed units *after a reboot*. Connectivity alone is insufficient — the
surviving interface works fine while boot remains degraded, which is precisely
how this went unnoticed through an entire verification pass.
---
### F-022 — transient DNS resolution timeout in CT 101
CT 101. Formal acceptance.
**Observed:** During the first acceptance pass, two unrelated public HTTPS
tests both failed at name resolution:
```
deb.debian.org: curl: (28) Resolving timed out after 5000 ms
gitea.barternetwork.us: curl: (28) Resolving timed out after 5000 ms
```
IP routing to `10.110.0.1`, `10.0.0.12` and CT 100 remained working throughout.
**Cause:** **Unproven.** Diagnostics showed resolver configuration identical to
CT 100 and the host (`nameserver 75.75.75.75`, `hosts: files dns`),
`systemd-resolved` absent, `75.75.75.75` reachable, and `getent ahostsv4`
resolving both names immediately afterward. Repeat HTTPS tests returned 200.
Worth noting as context, not as cause: since F-017, container DNS traverses the
host's masquerade to an external resolver. That dependency is new.
**Correction:** None. No configuration was changed.
**Consequence:** Do not convert a one-shot resolver timeout into a
configuration change without evidence. On recurrence, capture resolver state
and DNS traffic at the moment of failure before touching anything. A candidate
mitigation — a second `nameserver` line, so a single hiccup retries rather than
fails — is recorded but deliberately not applied on one unexplained event.
---
### F-023 — relay accepted the mail; final delivery failed
`srv-b`. Alerting.
**Observed:** The relay was discovered at `10.110.0.1:25` over `wg0`, banner
`wg-pk.diagnostics.kane-il.us`, offering STARTTLS with a self-signed
`CN = wg-pk`. It accepted unauthenticated SMTP from `10.110.0.12` through
`RCPT TO`, before and after STARTTLS. Ports 465 and 587 were unavailable, and
the public address `198.58.111.109` did not expose SMTP on this path.
With `relayhost = [10.110.0.1]:25` and `root: sandor@kane-il.us`, `srv-b`
recorded successful handoff:
```
relay=10.110.0.1[10.110.0.1]:25 dsn=2.0.0
status=sent (250 2.0.0 Ok: queued as 87E446243A)
```
The local queue emptied. The operator then received a **delivery-failure**
message at `sandor@kane-il.us`.
**Cause:** **Unproven**, downstream of the demonstrated handoff. The bounce
notice itself arriving at `sandor@kane-il.us` establishes that the relay can
deliver to that address, which narrows the problem to the failing message
rather than the destination. Leading hypothesis, untested: the envelope sender
is `root@srv-b.dev.infra`, and `dev.infra` does not resolve publicly, so a
downstream MTA rejects on sender-domain verification. Candidate remedies are
`myorigin` or `smtp_generic_maps`.
**Correction:** None. Mail alerting was deferred by operator decision.
**Consequence:** **SMTP 250 from the relay and an empty local queue prove
handoff, not delivery.** Alerting acceptance requires demonstrated end-to-end
receipt. Until then mail, `smartd` alerting, and any mail-dependent backup
alerting are unaccepted. Note separately that `postfix check` reports
divergence between `/var/spool/postfix` copies and their host originals,
including `/etc/hosts` and NSS libraries — a known cause of resolution failure
inside the chroot, and adjacent enough to this failure to be checked first.
**Resolution (2026-08-17):** Cause proven, three hops downstream of the origin.
`mx1` runs `permit_mynetworks, permit_auth_destination, reject` on both relay
and recipient restrictions. `wg-pk` connected from public addresses absent from
`mx1`'s `mynetworks`, and `kane-il.us` is not an authorised destination there,
so `RCPT TO <sandor@kane-il.us>` was rejected `554 5.7.1 Access denied` over
both IPv4 and IPv6. Added only `198.58.111.109/32` and
`[2600:3c00::f03c:92ff:fe42:43d7]/128` to `mx1`'s `mynetworks` and reloaded.
Four hypotheses were tested and ruled out with evidence before any change: the
`srv-b` alias (`postalias -q root` resolved correctly and the journal showed
`orig_to=<root>` forwarded), sender-domain rejection (`wg-pk` rewrites the
sender to `postmaster@diagnostics.kane-il.us` and `MAIL FROM` was accepted),
address-family asymmetry (both families failed identically), and routing on
`wg-pk` (it connected and received the rejection).
Method worth reusing: **RCPT-only probes before and after the change, over both
address families, with no message body.** `554` before, `250` after. That
proves the change caused the fix rather than coinciding with it — the standard
F-012 was written to enforce. Two messages then delivered end to end with full
headers captured, empty queue, no deferred entries. **Closed.**
---
### F-024 — work order specified a log file that does not exist
`srv-b`. Work order 002.
**Observed:** WORK ORDER 002 instructed `grep /var/log/mail.log`. The file does
not exist on `srv-b`.
**Cause:** **Proven.** Proxmox VE ships without `rsyslog`. Postfix logs to
journald only.
**Correction:** The operator used `journalctl -u postfix@-` and completed the
diagnosis. No configuration was changed; installing `rsyslog` to satisfy a
document would have been the wrong direction.
**Consequence:** **Specification defect.** All log inspection on a Proxmox host
uses `journalctl`, not files under `/var/log`. Corrected in `ENVIRONMENT.md`
revision 5. Automation that greps a logfile path will silently find nothing,
which is worse than failing.
---
### F-025 — the F-023 correction granted wider relay than required
`mx1`, `wg-pk`. Estate scope.
**Observed:** Adding `wg-pk`'s two public addresses to `mx1`'s `mynetworks`
authorises them under `permit_mynetworks`, which appears in **both**
`smtpd_relay_restrictions` and `smtpd_recipient_restrictions`. That grants relay
to **any** destination, not only to `kane-il.us`.
`wg-pk` in turn carries `mynetworks = 10.110.0.0/22` with
`permit_mynetworks permit_sasl_authenticated defer_unauth_destination`. The
resulting chain:
```
any peer on 10.110.0.0/22 (including CT 100 and CT 101, which arrive
as 10.110.0.12 through the srv-b masquerade)
-> wg-pk permit_mynetworks
-> mx1 permit_mynetworks
-> any destination on the internet, as kane-il.us infrastructure
```
Before the correction `mx1` rejected at the final hop, so the path was closed by
accident rather than by policy.
**Cause:** **Proven.** `permit_mynetworks` is destination-agnostic by design.
**Correction:** **Pending — operator decision.** Two independent parts:
- *Ours:* block SMTP egress from the container network at `srv-b`. The
containers have no reason to originate mail; if they ever should, that is a
deliberate decision rather than something inherited from a masquerade. Local,
precise, touches no estate policy.
- *Estate:* whether `wg-pk`'s `mynetworks` should be explicit `/32` entries for
the hosts that legitimately originate mail rather than the whole tunnel range.
**Constraint:** `wg-pk` must continue to relay to arbitrary external
destinations — that is its purpose, since Hubzilla registration and notification
mail depends on it and the ISP blocks port 25. Restricting it by *recipient*
would break that. The question is which clients may ask, not where it may send.
**Consequence:** A correction that fixes the observed failure may widen an
adjacent boundary. Record what a trust change grants, not only what it repairs.
---
## Open, not closed
| # | Status |
|---|---|
| F-006 | **Open** — cause unproven. Recurrence should capture `dpkg` lock state. |
| F-012 | **Open** — cause unproven. Leading candidate ruled out by inspection. `default_server` was added as independent hardening and does **not** close this. |
| F-019 | **Corrected** 2026-08-16. Reboot persistence proven. |
| F-021 | **Corrected** 2026-08-16. Reboot persistence proven. |
| F-022 | **Open** — cause unproven, no correction applied. |
| F-023 | **Closed** 2026-08-17. Cause proven at `mx1`; delivery proven twice. |
| F-024 | **Closed** 2026-08-17. Specification corrected. |
| F-025 | **Open** — correction pending operator decision. |
Everything else is closed with a proven cause and a proven correction.