462 lines
19 KiB
Markdown
462 lines
19 KiB
Markdown
# FAILURES.md
|
|
|
|
Append-only record of every failure encountered building the Mechanical
|
|
Compiler environment.
|
|
|
|
| | |
|
|
|---|---|
|
|
| Scope | All instances. Staging entries are marked `srv-b`. |
|
|
| Updated | 2026-08-16, after staging acceptance |
|
|
| Rule | Append only. Never edit an entry except to add a `Resolution` line. |
|
|
| Numbering | Sequential, never reused. See §0 on the renumbering. |
|
|
|
|
---
|
|
|
|
## 0. Why this file exists separately
|
|
|
|
This is the primary input to provisioning automation.
|
|
|
|
Every entry below is something a script written from `ENVIRONMENT.md` alone
|
|
would have got wrong. `F-003` and `F-004` in particular would have produced a
|
|
container that booted, reported success, and was quietly broken. The
|
|
specification cannot anticipate them; only contact with a host reveals them.
|
|
|
|
When the production automation is written, it should be read **before**
|
|
`ENVIRONMENT.md`, not after.
|
|
|
|
### Renumbering note, 2026-08-15
|
|
|
|
The staging checkpoint and the reconciliation checkpoint were written
|
|
independently and both allocated `F-008` and `F-009`. The reconciliation
|
|
entries have been renumbered to `F-015` and `F-016`. Original phase-1 numbering
|
|
is preserved because it is cited elsewhere.
|
|
|
|
### Entry format
|
|
|
|
```
|
|
### F-nnn — one-line summary
|
|
Host / container. Phase.
|
|
**Observed:** what was seen, verbatim where possible.
|
|
**Cause:** proven cause, or explicitly "unproven".
|
|
**Correction:** the smallest change that fixed it.
|
|
**Consequence:** what the specification or automation must do differently.
|
|
```
|
|
|
|
A failure with no `Consequence` line is either not yet understood or not worth
|
|
recording.
|
|
|
|
---
|
|
|
|
## Phase 1 — initial container build
|
|
|
|
### F-001 — `arping` absent on the Proxmox host
|
|
`srv-b`.
|
|
|
|
**Observed:** `command -v arping` → not installed.
|
|
**Cause:** Not part of the base PVE install.
|
|
**Correction:** `apt install iputils-arping`.
|
|
**Consequence:** Automation must install `iputils-arping` before duplicate
|
|
address detection, or the DAD check silently does not happen.
|
|
|
|
---
|
|
|
|
### F-002 — invalid Proxmox inspection command
|
|
`srv-b`.
|
|
|
|
**Observed:** `pvesm config local-lvm` returned CLI usage text.
|
|
**Cause:** Operator command error. Not a host fault.
|
|
**Correction:** Read `/etc/pve/storage.cfg` directly.
|
|
**Consequence:** Read storage configuration from the file, not from a
|
|
subcommand that does not exist.
|
|
|
|
---
|
|
|
|
### F-003 — CT 101 booted with degraded systemd
|
|
CT 101.
|
|
|
|
**Observed:**
|
|
|
|
```
|
|
status=226/NAMESPACE
|
|
Failed to set up mount namespacing: Permission denied
|
|
systemd-logind.service failed
|
|
systemd-networkd.service failed
|
|
systemd-timedated.service failed
|
|
systemd-networkd.socket failed
|
|
```
|
|
|
|
**Cause:** Debian 12 systemd requires mount namespacing that an unprivileged
|
|
LXC denies without `nesting`.
|
|
**Correction:** `features: nesting=1`. Deliberately tested alone rather than
|
|
copying `nesting=1,keyctl=1` from CT 100 — `keyctl=1` proved unnecessary.
|
|
**Consequence:** **Every Debian 12 unprivileged container needs
|
|
`nesting=1`**, including ones running nothing but nginx. `keyctl=1` is needed
|
|
only where Docker runs. Promoted into the specification.
|
|
|
|
This is the clearest example of why the manual build was correct: the container
|
|
started, appeared healthy at a glance, and was broken in four services.
|
|
|
|
---
|
|
|
|
### F-004 — locale configured but never generated
|
|
CT 100, CT 101.
|
|
|
|
**Observed:** `LANG=en_US.UTF-8` set, but `locale -a` listed only `C`,
|
|
`C.utf8`, `POSIX`. `locale` emitted `Cannot set LC_*` warnings.
|
|
**Cause:** Setting `LANG` does not generate a locale. The container template
|
|
ships none.
|
|
**Correction:** install `locales`, enable `en_US.UTF-8 UTF-8`, `locale-gen`,
|
|
`update-locale`.
|
|
**Consequence:** Specification must state the generation step, not just the
|
|
value. Silent until something depends on collation or encoding.
|
|
|
|
---
|
|
|
|
### F-005 — locale repair exposed a pending upgrade backlog
|
|
CT 100, CT 101.
|
|
|
|
**Observed:** Installing `locales` pulled glibc-related packages and revealed
|
|
52 further pending upgrades.
|
|
**Cause:** Template was not current.
|
|
**Correction:** Full upgrade of each container, one at a time. Both reached
|
|
Debian 12.15, zero pending, no reboot required.
|
|
**Consequence:** Automation must `apt full-upgrade` immediately after first
|
|
boot, before anything else is installed.
|
|
|
|
---
|
|
|
|
### F-006 — transient `curl` / `git` command-not-found
|
|
CT 100.
|
|
|
|
**Observed:** An egress test returned `curl: command not found` and
|
|
`git: command not found`. Immediately afterward both packages were installed
|
|
and current, `PATH` normal, both executables present and working.
|
|
**Cause:** **Unproven.** Hypothesis only: the test ran during the F-005
|
|
upgrade, while `dpkg` had the binaries briefly unlinked mid-transaction.
|
|
Adjacent in time, consistent with the symptom, not demonstrated.
|
|
**Correction:** None required. Repeat tests with absolute paths passed.
|
|
**Consequence:** Do not close. If it recurs, capture `dpkg` lock state at the
|
|
moment of failure.
|
|
|
|
---
|
|
|
|
### F-007 — recursive `chown` failed on ext4 `lost+found`
|
|
CT 100.
|
|
|
|
**Observed:** `chown: cannot read directory '/var/lib/mechcomp/lost+found':
|
|
Permission denied`. That directory is `nobody:nogroup`, mode `0700`.
|
|
**Cause:** `/var/lib/mechcomp` is a real filesystem root, not a directory.
|
|
`lost+found` is created by `mkfs` and is not ours to manage.
|
|
**Correction:** Own explicit application paths only. Leave `lost+found` alone.
|
|
**Consequence:** **Never `chown -R` a mount point root.** Enumerate the
|
|
directories the application actually uses. Promoted into the specification.
|
|
|
|
---
|
|
|
|
### F-008 — Git "dubious ownership"
|
|
CT 100.
|
|
|
|
**Observed:** `fatal: detected dubious ownership in repository at
|
|
'/var/www/mechcomp'` when verification ran as root.
|
|
**Cause:** Repository cloned as `mechcomp`, inspected as `root`.
|
|
**Correction:** Re-ran verification as the owning user. **No root
|
|
`safe.directory` exception was added** — the correct fix, since the exception
|
|
would have masked every future instance of the same mistake.
|
|
**Consequence:** All repository operations run as the service user.
|
|
|
|
---
|
|
|
|
### F-009 — dependency manifests absent from the application repository
|
|
CT 100.
|
|
|
|
**Observed:** `requirements-base.txt` and `requirements-cad.txt` do not exist;
|
|
the repository is `LICENSE` and `README.md` at commit `e85c4f4e`.
|
|
**Cause:** Application implementation has not started. Not a provisioning
|
|
failure.
|
|
**Correction:** None. Do not invent pins.
|
|
**Consequence:** Infrastructure and application acceptance are separately
|
|
gated. Infrastructure may complete without the application existing.
|
|
|
|
---
|
|
|
|
## Phase 2 — architect corrections
|
|
|
|
### F-010 — placeholder PID file rewrite failed to substitute
|
|
CT 100.
|
|
|
|
**Observed:** An automated rewrite left the literal `$!` in the PID file.
|
|
**Cause:** Quoting error in the rewriting command.
|
|
**Correction:** Identified the real PID by inspection.
|
|
**Consequence:** Superseded by F-019 — the placeholder should be a systemd
|
|
unit and have no PID file at all.
|
|
|
|
---
|
|
|
|
### F-011 — root could not overwrite a `mechcomp`-owned file in `/tmp`
|
|
CT 100.
|
|
|
|
**Observed:** `Permission denied` writing a file owned by `mechcomp:mechcomp`
|
|
in `/tmp`, as root.
|
|
**Cause:** `fs.protected_regular = 2` with `/tmp` mode `1777`. Expected kernel
|
|
behaviour, not a fault.
|
|
**Correction:** Rewrote the file as the owning user.
|
|
**Consequence:** Do not use `/tmp` for state shared across users. Moot once
|
|
services run under systemd with `PrivateTmp=true`.
|
|
|
|
---
|
|
|
|
### F-012 — first request after nginx reload returned the Debian welcome page
|
|
CT 101.
|
|
|
|
**Observed:** `nginx -t` passed; the immediately following request returned the
|
|
615-byte Debian default page. All subsequent requests proxied correctly.
|
|
**Cause:** **Unproven.** The obvious candidate — Debian's `default` site still
|
|
enabled and acting as `default_server` — was tested and ruled out:
|
|
`sites-enabled` contains only the project vhost. Remaining hypothesis is a
|
|
graceful worker transition. Not demonstrated.
|
|
**Correction:** None required.
|
|
**Consequence:** Do not close. Independently, declare `default_server`
|
|
explicitly on the project vhost so that adding a second server block cannot
|
|
make catch-all behaviour depend on file ordering.
|
|
|
|
---
|
|
|
|
### F-013 — TLS test validated against an IP literal
|
|
CT 101.
|
|
|
|
**Observed:** `SSL: no alternative certificate subject name matches target
|
|
host name '10.0.0.21'`.
|
|
**Cause:** Test used an address; the certificate SAN contains a DNS name.
|
|
Correct behaviour by both curl and OpenSSL.
|
|
**Correction:** Test by hostname.
|
|
**Consequence:** Acceptance checks must use the service FQDN. An IP-literal
|
|
TLS test is always wrong unless the certificate carries an IP SAN.
|
|
|
|
---
|
|
|
|
### F-014 — CT 101 had no `mechcomp` group
|
|
CT 101.
|
|
|
|
**Observed:** Adding `sandor` to `mechcomp` failed; the group did not exist.
|
|
**Cause:** Specification said the admin joins the `mechcomp` group in *both*
|
|
containers, but the group is created as a side effect of creating the service
|
|
user, which exists only in CT 100.
|
|
**Correction:** Created the system group and added `sandor`.
|
|
**Consequence:** Specification defect. Either create the group explicitly where
|
|
it is required, or scope the group membership to CT 100. GID 996 now matches
|
|
across both containers, which is harmless and mildly useful.
|
|
|
|
---
|
|
|
|
### F-015 — CT 101 verification lacked `curl`
|
|
CT 101. *(Renumbered from a duplicate `F-008`.)*
|
|
|
|
**Observed:** `curl: command not found`.
|
|
**Cause:** Not in the base template; CT 101's package list did not include it.
|
|
**Correction:** Installed `curl` and `ca-certificates`.
|
|
**Consequence:** The proxy container needs its own minimal toolset. Do not
|
|
assume packages installed in CT 100 exist in CT 101.
|
|
|
|
---
|
|
|
|
### F-016 — placeholder PID capture stored a literal `$!`
|
|
CT 100. *(Renumbered from a duplicate `F-009`.)*
|
|
|
|
**Observed:** PID file contained `$!` rather than a number. The listener itself
|
|
was healthy.
|
|
**Cause:** Quoting error.
|
|
**Correction:** Identified PID `12102` by inspection.
|
|
**Consequence:** Superseded by F-019.
|
|
|
|
---
|
|
|
|
## Phase 3 — network isolation
|
|
|
|
### F-017 — containers were provisioned on the home LAN
|
|
`srv-b`, CT 100, CT 101.
|
|
|
|
**Observed:** Both containers had `net0` on `vmbr0` at `10.0.0.20/24` and
|
|
`10.0.0.21/24`, gateway `10.0.0.1` — the ISP router's network. Anything on the
|
|
home LAN could reach them.
|
|
**Cause:** **Specification defect in `ENVIRONMENT.md` revision 4.** The
|
|
document placed containers on both bridges, giving each a management address on
|
|
the LAN, on the unstated assumption that a LAN workstation would browse
|
|
staging. That assumption was never confirmed and was wrong.
|
|
**Correction:** Both containers moved to `vmbr1` only, gateway `10.20.0.1`.
|
|
`net1` deleted. Host NAT extended for `10.20.0.0/24`.
|
|
**Consequence:** Containers are on the portless service bridge only. `srv-b` is
|
|
router and bastion; access arrives via WireGuard. Promoted into the
|
|
specification.
|
|
|
|
---
|
|
|
|
### F-018 — removing the LAN interface did not remove LAN reachability
|
|
CT 100, CT 101.
|
|
|
|
**Observed:** After F-017, both containers still reached `10.0.0.1`
|
|
successfully.
|
|
**Cause:** Removing an interface removes an address, not a route.
|
|
`ip_forward=1` plus the new `MASQUERADE -s 10.20.0.0/24 -o vmbr0` — added to
|
|
give the containers internet access — also gave them the LAN, translated to the
|
|
host's address.
|
|
**Correction:** A `RETURN` in `nat POSTROUTING` ahead of both masquerades, and
|
|
a `DROP` in `FORWARD`, both scoped `-s 10.20.0.0/24 -d 10.0.0.0/24`, plus an
|
|
explicit `ACCEPT` for container-to-container traffic.
|
|
**Consequence:** **Interface removal is not isolation.** Automation must assert
|
|
the negative — that the LAN is unreachable — not merely that the interface is
|
|
gone. Note that the containers still reach `srv-b` itself at `10.0.0.12`,
|
|
because a container addressing its gateway takes the INPUT path and never
|
|
enters FORWARD. That is required, not a leak.
|
|
|
|
---
|
|
|
|
### F-019 — placeholder backend did not survive reboot
|
|
CT 100.
|
|
|
|
**Observed:** After the F-017 reboot, nothing listened on `10.20.0.10:8770`.
|
|
**Cause:** The placeholder was a bare foreground process with a PID file. No
|
|
supervision.
|
|
**Correction:** Recreated as `mechcomp-placeholder.service`.
|
|
**Consequence:** Anything a proof depends on must be supervised, or the proof
|
|
expires silently at the next reboot and the following session diagnoses a
|
|
proxy fault that does not exist. Applies to temporary scaffolding as much as to
|
|
real services.
|
|
**Resolution (2026-08-16):** Installed `/usr/local/libexec/mechcomp-placeholder.py`
|
|
and `/etc/systemd/system/mechcomp-placeholder.service`, enabled at
|
|
`multi-user.target`, running as `mechcomp:mechcomp`, reading
|
|
`/etc/mechcomp/mechcomp.env`, binding `10.20.0.10:8770`, with `Restart=on-failure`
|
|
and the standard hardening set. CT 100 was rebooted: the unit restarted
|
|
automatically, the listener returned on the service address only, nginx proxied
|
|
successfully, `X-Forwarded-Proto: https` was re-observed at the backend, and the
|
|
container settled to `running` with zero failed units.
|
|
|
|
---
|
|
|
|
### F-020 — service FQDN mapping changed twice
|
|
`srv-b`, CT 100, CT 101.
|
|
|
|
**Observed:** `mechanical-compiler.dev.infra` was mapped to `10.20.0.10`, then
|
|
corrected to `10.0.0.21`, then corrected again to `10.20.0.11`.
|
|
**Cause:** Three different states, each correct for its moment. `10.20.0.10`
|
|
was wrong — it pointed at the application rather than the proxy. `10.0.0.21`
|
|
was right while CT 101 had a LAN address and workstation access was assumed.
|
|
`10.20.0.11` became right once F-017 removed the LAN interface and access moved
|
|
to WireGuard through `srv-b`.
|
|
**Correction:** `10.20.0.11` on all three hosts.
|
|
**Consequence:** Not an error in either direction — a value that tracked a
|
|
topology decision. Recorded because the reasoning matters more than the value:
|
|
**the service FQDN names the proxy, and the proxy has exactly one address.**
|
|
|
|
---
|
|
|
|
### F-021 — deleting the Proxmox interface left a stale guest `eth1`
|
|
CT 100, CT 101. Network isolation.
|
|
|
|
**Observed:** After the F-017 topology change and reboot, both containers
|
|
reported `systemctl is-system-running` → `degraded`, with
|
|
`networking.service` and `ifupdown-wait-online.service` failed. Both guests'
|
|
`/etc/network/interfaces` still carried an `auto eth1` / static `iface eth1`
|
|
stanza although neither had an `eth1` link. The boot journal on both:
|
|
|
|
```
|
|
Cannot find device "eth1"
|
|
ifup: failed to bring up eth1
|
|
```
|
|
|
|
**Cause:** **Proven.** `pct set --delete net1` removes the LXC interface but
|
|
does not remove the stanza already written into the guest. `ifup -a` therefore
|
|
exited non-zero at boot even though `eth0` came up correctly.
|
|
`ifupdown-wait-online` failed as a consequence, not independently.
|
|
`systemd-networkd` was investigated and explicitly ruled out — its units were
|
|
disabled on both containers.
|
|
**Correction:** Preserved a pre-correction copy, removed only the stale `eth1`
|
|
stanza on each guest, restarted the affected units. Both containers then
|
|
rebooted: the stanza did not return, both units succeeded, both settled to
|
|
`running` with zero failed units.
|
|
**Consequence:** **Removing a container interface is not proof that the guest
|
|
converged.** After any topology mutation, automation must inspect the guest
|
|
interface file and assert `systemctl is-system-running = running` with zero
|
|
failed units *after a reboot*. Connectivity alone is insufficient — the
|
|
surviving interface works fine while boot remains degraded, which is precisely
|
|
how this went unnoticed through an entire verification pass.
|
|
|
|
---
|
|
|
|
### F-022 — transient DNS resolution timeout in CT 101
|
|
CT 101. Formal acceptance.
|
|
|
|
**Observed:** During the first acceptance pass, two unrelated public HTTPS
|
|
tests both failed at name resolution:
|
|
|
|
```
|
|
deb.debian.org: curl: (28) Resolving timed out after 5000 ms
|
|
gitea.barternetwork.us: curl: (28) Resolving timed out after 5000 ms
|
|
```
|
|
|
|
IP routing to `10.110.0.1`, `10.0.0.12` and CT 100 remained working throughout.
|
|
**Cause:** **Unproven.** Diagnostics showed resolver configuration identical to
|
|
CT 100 and the host (`nameserver 75.75.75.75`, `hosts: files dns`),
|
|
`systemd-resolved` absent, `75.75.75.75` reachable, and `getent ahostsv4`
|
|
resolving both names immediately afterward. Repeat HTTPS tests returned 200.
|
|
|
|
Worth noting as context, not as cause: since F-017, container DNS traverses the
|
|
host's masquerade to an external resolver. That dependency is new.
|
|
**Correction:** None. No configuration was changed.
|
|
**Consequence:** Do not convert a one-shot resolver timeout into a
|
|
configuration change without evidence. On recurrence, capture resolver state
|
|
and DNS traffic at the moment of failure before touching anything. A candidate
|
|
mitigation — a second `nameserver` line, so a single hiccup retries rather than
|
|
fails — is recorded but deliberately not applied on one unexplained event.
|
|
|
|
---
|
|
|
|
### F-023 — relay accepted the mail; final delivery failed
|
|
`srv-b`. Alerting.
|
|
|
|
**Observed:** The relay was discovered at `10.110.0.1:25` over `wg0`, banner
|
|
`wg-pk.diagnostics.kane-il.us`, offering STARTTLS with a self-signed
|
|
`CN = wg-pk`. It accepted unauthenticated SMTP from `10.110.0.12` through
|
|
`RCPT TO`, before and after STARTTLS. Ports 465 and 587 were unavailable, and
|
|
the public address `198.58.111.109` did not expose SMTP on this path.
|
|
|
|
With `relayhost = [10.110.0.1]:25` and `root: sandor@kane-il.us`, `srv-b`
|
|
recorded successful handoff:
|
|
|
|
```
|
|
relay=10.110.0.1[10.110.0.1]:25 dsn=2.0.0
|
|
status=sent (250 2.0.0 Ok: queued as 87E446243A)
|
|
```
|
|
|
|
The local queue emptied. The operator then received a **delivery-failure**
|
|
message at `sandor@kane-il.us`.
|
|
**Cause:** **Unproven**, downstream of the demonstrated handoff. The bounce
|
|
notice itself arriving at `sandor@kane-il.us` establishes that the relay can
|
|
deliver to that address, which narrows the problem to the failing message
|
|
rather than the destination. Leading hypothesis, untested: the envelope sender
|
|
is `root@srv-b.dev.infra`, and `dev.infra` does not resolve publicly, so a
|
|
downstream MTA rejects on sender-domain verification. Candidate remedies are
|
|
`myorigin` or `smtp_generic_maps`.
|
|
**Correction:** None. Mail alerting was deferred by operator decision.
|
|
**Consequence:** **SMTP 250 from the relay and an empty local queue prove
|
|
handoff, not delivery.** Alerting acceptance requires demonstrated end-to-end
|
|
receipt. Until then mail, `smartd` alerting, and any mail-dependent backup
|
|
alerting are unaccepted. Note separately that `postfix check` reports
|
|
divergence between `/var/spool/postfix` copies and their host originals,
|
|
including `/etc/hosts` and NSS libraries — a known cause of resolution failure
|
|
inside the chroot, and adjacent enough to this failure to be checked first.
|
|
|
|
---
|
|
|
|
## Open, not closed
|
|
|
|
| # | Status |
|
|
|---|---|
|
|
| F-006 | **Open** — cause unproven. Recurrence should capture `dpkg` lock state. |
|
|
| F-012 | **Open** — cause unproven. Leading candidate ruled out by inspection. `default_server` was added as independent hardening and does **not** close this. |
|
|
| F-019 | **Corrected** 2026-08-16. Reboot persistence proven. |
|
|
| F-021 | **Corrected** 2026-08-16. Reboot persistence proven. |
|
|
| F-022 | **Open** — cause unproven, no correction applied. |
|
|
| F-023 | **Deferred** — downstream mail failure, cause unproven. Hypothesis recorded. |
|
|
|
|
Everything else is closed with a proven cause and a proven correction.
|