752 lines
33 KiB
Markdown
752 lines
33 KiB
Markdown
# FAILURES.md
|
|
|
|
Append-only record of every failure encountered building the Mechanical
|
|
Compiler environment.
|
|
|
|
| | |
|
|
|---|---|
|
|
| Scope | All instances. Staging entries are marked `srv-b`. |
|
|
| Updated | 2026-08-18, after repository seeding and container standardisation |
|
|
| Method | `PROCESS.md` section 7 |
|
|
| Rule | Append only. Never edit an entry except to add a `Resolution` line. |
|
|
| Numbering | Sequential, never reused. See §0 on the renumbering. |
|
|
|
|
---
|
|
|
|
## 0. Why this file exists separately
|
|
|
|
This is the primary input to provisioning automation.
|
|
|
|
Every entry below is something a script written from `ENVIRONMENT.md` alone
|
|
would have got wrong. `F-003` and `F-004` in particular would have produced a
|
|
container that booted, reported success, and was quietly broken. The
|
|
specification cannot anticipate them; only contact with a host reveals them.
|
|
|
|
When the production automation is written, it should be read **before**
|
|
`ENVIRONMENT.md`, not after.
|
|
|
|
### Renumbering note, 2026-08-15
|
|
|
|
The staging checkpoint and the reconciliation checkpoint were written
|
|
independently and both allocated `F-008` and `F-009`. The reconciliation
|
|
entries have been renumbered to `F-015` and `F-016`. Original phase-1 numbering
|
|
is preserved because it is cited elsewhere.
|
|
|
|
### Entry format
|
|
|
|
```
|
|
### F-nnn — one-line summary
|
|
Host / container. Phase.
|
|
**Observed:** what was seen, verbatim where possible.
|
|
**Cause:** proven cause, or explicitly "unproven".
|
|
**Correction:** the smallest change that fixed it.
|
|
**Consequence:** what the specification or automation must do differently.
|
|
```
|
|
|
|
A failure with no `Consequence` line is either not yet understood or not worth
|
|
recording.
|
|
|
|
---
|
|
|
|
## Phase 1 — initial container build
|
|
|
|
### F-001 — `arping` absent on the Proxmox host
|
|
`srv-b`.
|
|
|
|
**Observed:** `command -v arping` → not installed.
|
|
**Cause:** Not part of the base PVE install.
|
|
**Correction:** `apt install iputils-arping`.
|
|
**Consequence:** Automation must install `iputils-arping` before duplicate
|
|
address detection, or the DAD check silently does not happen.
|
|
|
|
---
|
|
|
|
### F-002 — invalid Proxmox inspection command
|
|
`srv-b`.
|
|
|
|
**Observed:** `pvesm config local-lvm` returned CLI usage text.
|
|
**Cause:** Operator command error. Not a host fault.
|
|
**Correction:** Read `/etc/pve/storage.cfg` directly.
|
|
**Consequence:** Read storage configuration from the file, not from a
|
|
subcommand that does not exist.
|
|
|
|
---
|
|
|
|
### F-003 — CT 101 booted with degraded systemd
|
|
CT 101.
|
|
|
|
**Observed:**
|
|
|
|
```
|
|
status=226/NAMESPACE
|
|
Failed to set up mount namespacing: Permission denied
|
|
systemd-logind.service failed
|
|
systemd-networkd.service failed
|
|
systemd-timedated.service failed
|
|
systemd-networkd.socket failed
|
|
```
|
|
|
|
**Cause:** Debian 12 systemd requires mount namespacing that an unprivileged
|
|
LXC denies without `nesting`.
|
|
**Correction:** `features: nesting=1`. Deliberately tested alone rather than
|
|
copying `nesting=1,keyctl=1` from CT 100 — `keyctl=1` proved unnecessary.
|
|
**Consequence:** **Every Debian 12 unprivileged container needs
|
|
`nesting=1`**, including ones running nothing but nginx. `keyctl=1` is needed
|
|
only where Docker runs. Promoted into the specification.
|
|
|
|
This is the clearest example of why the manual build was correct: the container
|
|
started, appeared healthy at a glance, and was broken in four services.
|
|
|
|
---
|
|
|
|
### F-004 — locale configured but never generated
|
|
CT 100, CT 101.
|
|
|
|
**Observed:** `LANG=en_US.UTF-8` set, but `locale -a` listed only `C`,
|
|
`C.utf8`, `POSIX`. `locale` emitted `Cannot set LC_*` warnings.
|
|
**Cause:** Setting `LANG` does not generate a locale. The container template
|
|
ships none.
|
|
**Correction:** install `locales`, enable `en_US.UTF-8 UTF-8`, `locale-gen`,
|
|
`update-locale`.
|
|
**Consequence:** Specification must state the generation step, not just the
|
|
value. Silent until something depends on collation or encoding.
|
|
|
|
---
|
|
|
|
### F-005 — locale repair exposed a pending upgrade backlog
|
|
CT 100, CT 101.
|
|
|
|
**Observed:** Installing `locales` pulled glibc-related packages and revealed
|
|
52 further pending upgrades.
|
|
**Cause:** Template was not current.
|
|
**Correction:** Full upgrade of each container, one at a time. Both reached
|
|
Debian 12.15, zero pending, no reboot required.
|
|
**Consequence:** Automation must `apt full-upgrade` immediately after first
|
|
boot, before anything else is installed.
|
|
|
|
---
|
|
|
|
### F-006 — transient `curl` / `git` command-not-found
|
|
CT 100.
|
|
|
|
**Observed:** An egress test returned `curl: command not found` and
|
|
`git: command not found`. Immediately afterward both packages were installed
|
|
and current, `PATH` normal, both executables present and working.
|
|
**Cause:** **Unproven.** Hypothesis only: the test ran during the F-005
|
|
upgrade, while `dpkg` had the binaries briefly unlinked mid-transaction.
|
|
Adjacent in time, consistent with the symptom, not demonstrated.
|
|
**Correction:** None required. Repeat tests with absolute paths passed.
|
|
**Consequence:** Do not close. If it recurs, capture `dpkg` lock state at the
|
|
moment of failure.
|
|
|
|
---
|
|
|
|
### F-007 — recursive `chown` failed on ext4 `lost+found`
|
|
CT 100.
|
|
|
|
**Observed:** `chown: cannot read directory '/var/lib/mechcomp/lost+found':
|
|
Permission denied`. That directory is `nobody:nogroup`, mode `0700`.
|
|
**Cause:** `/var/lib/mechcomp` is a real filesystem root, not a directory.
|
|
`lost+found` is created by `mkfs` and is not ours to manage.
|
|
**Correction:** Own explicit application paths only. Leave `lost+found` alone.
|
|
**Consequence:** **Never `chown -R` a mount point root.** Enumerate the
|
|
directories the application actually uses. Promoted into the specification.
|
|
|
|
---
|
|
|
|
### F-008 — Git "dubious ownership"
|
|
CT 100.
|
|
|
|
**Observed:** `fatal: detected dubious ownership in repository at
|
|
'/var/www/mechcomp'` when verification ran as root.
|
|
**Cause:** Repository cloned as `mechcomp`, inspected as `root`.
|
|
**Correction:** Re-ran verification as the owning user. **No root
|
|
`safe.directory` exception was added** — the correct fix, since the exception
|
|
would have masked every future instance of the same mistake.
|
|
**Consequence:** All repository operations run as the service user.
|
|
|
|
---
|
|
|
|
### F-009 — dependency manifests absent from the application repository
|
|
CT 100.
|
|
|
|
**Observed:** `requirements-base.txt` and `requirements-cad.txt` do not exist;
|
|
the repository is `LICENSE` and `README.md` at commit `e85c4f4e`.
|
|
**Cause:** Application implementation has not started. Not a provisioning
|
|
failure.
|
|
**Correction:** None. Do not invent pins.
|
|
**Consequence:** Infrastructure and application acceptance are separately
|
|
gated. Infrastructure may complete without the application existing.
|
|
|
|
---
|
|
|
|
## Phase 2 — architect corrections
|
|
|
|
### F-010 — placeholder PID file rewrite failed to substitute
|
|
CT 100.
|
|
|
|
**Observed:** An automated rewrite left the literal `$!` in the PID file.
|
|
**Cause:** Quoting error in the rewriting command.
|
|
**Correction:** Identified the real PID by inspection.
|
|
**Consequence:** Superseded by F-019 — the placeholder should be a systemd
|
|
unit and have no PID file at all.
|
|
|
|
---
|
|
|
|
### F-011 — root could not overwrite a `mechcomp`-owned file in `/tmp`
|
|
CT 100.
|
|
|
|
**Observed:** `Permission denied` writing a file owned by `mechcomp:mechcomp`
|
|
in `/tmp`, as root.
|
|
**Cause:** `fs.protected_regular = 2` with `/tmp` mode `1777`. Expected kernel
|
|
behaviour, not a fault.
|
|
**Correction:** Rewrote the file as the owning user.
|
|
**Consequence:** Do not use `/tmp` for state shared across users. Moot once
|
|
services run under systemd with `PrivateTmp=true`.
|
|
|
|
---
|
|
|
|
### F-012 — first request after nginx reload returned the Debian welcome page
|
|
CT 101.
|
|
|
|
**Observed:** `nginx -t` passed; the immediately following request returned the
|
|
615-byte Debian default page. All subsequent requests proxied correctly.
|
|
**Cause:** **Unproven.** The obvious candidate — Debian's `default` site still
|
|
enabled and acting as `default_server` — was tested and ruled out:
|
|
`sites-enabled` contains only the project vhost. Remaining hypothesis is a
|
|
graceful worker transition. Not demonstrated.
|
|
**Correction:** None required.
|
|
**Consequence:** Do not close. Independently, declare `default_server`
|
|
explicitly on the project vhost so that adding a second server block cannot
|
|
make catch-all behaviour depend on file ordering.
|
|
|
|
---
|
|
|
|
### F-013 — TLS test validated against an IP literal
|
|
CT 101.
|
|
|
|
**Observed:** `SSL: no alternative certificate subject name matches target
|
|
host name '10.0.0.21'`.
|
|
**Cause:** Test used an address; the certificate SAN contains a DNS name.
|
|
Correct behaviour by both curl and OpenSSL.
|
|
**Correction:** Test by hostname.
|
|
**Consequence:** Acceptance checks must use the service FQDN. An IP-literal
|
|
TLS test is always wrong unless the certificate carries an IP SAN.
|
|
|
|
---
|
|
|
|
### F-014 — CT 101 had no `mechcomp` group
|
|
CT 101.
|
|
|
|
**Observed:** Adding `sandor` to `mechcomp` failed; the group did not exist.
|
|
**Cause:** Specification said the admin joins the `mechcomp` group in *both*
|
|
containers, but the group is created as a side effect of creating the service
|
|
user, which exists only in CT 100.
|
|
**Correction:** Created the system group and added `sandor`.
|
|
**Consequence:** Specification defect. Either create the group explicitly where
|
|
it is required, or scope the group membership to CT 100. GID 996 now matches
|
|
across both containers, which is harmless and mildly useful.
|
|
|
|
---
|
|
|
|
### F-015 — CT 101 verification lacked `curl`
|
|
CT 101. *(Renumbered from a duplicate `F-008`.)*
|
|
|
|
**Observed:** `curl: command not found`.
|
|
**Cause:** Not in the base template; CT 101's package list did not include it.
|
|
**Correction:** Installed `curl` and `ca-certificates`.
|
|
**Consequence:** The proxy container needs its own minimal toolset. Do not
|
|
assume packages installed in CT 100 exist in CT 101.
|
|
|
|
---
|
|
|
|
### F-016 — placeholder PID capture stored a literal `$!`
|
|
CT 100. *(Renumbered from a duplicate `F-009`.)*
|
|
|
|
**Observed:** PID file contained `$!` rather than a number. The listener itself
|
|
was healthy.
|
|
**Cause:** Quoting error.
|
|
**Correction:** Identified PID `12102` by inspection.
|
|
**Consequence:** Superseded by F-019.
|
|
|
|
---
|
|
|
|
## Phase 3 — network isolation
|
|
|
|
### F-017 — containers were provisioned on the home LAN
|
|
`srv-b`, CT 100, CT 101.
|
|
|
|
**Observed:** Both containers had `net0` on `vmbr0` at `10.0.0.20/24` and
|
|
`10.0.0.21/24`, gateway `10.0.0.1` — the ISP router's network. Anything on the
|
|
home LAN could reach them.
|
|
**Cause:** **Specification defect in `ENVIRONMENT.md` revision 4.** The
|
|
document placed containers on both bridges, giving each a management address on
|
|
the LAN, on the unstated assumption that a LAN workstation would browse
|
|
staging. That assumption was never confirmed and was wrong.
|
|
**Correction:** Both containers moved to `vmbr1` only, gateway `10.20.0.1`.
|
|
`net1` deleted. Host NAT extended for `10.20.0.0/24`.
|
|
**Consequence:** Containers are on the portless service bridge only. `srv-b` is
|
|
router and bastion; access arrives via WireGuard. Promoted into the
|
|
specification.
|
|
|
|
---
|
|
|
|
### F-018 — removing the LAN interface did not remove LAN reachability
|
|
CT 100, CT 101.
|
|
|
|
**Observed:** After F-017, both containers still reached `10.0.0.1`
|
|
successfully.
|
|
**Cause:** Removing an interface removes an address, not a route.
|
|
`ip_forward=1` plus the new `MASQUERADE -s 10.20.0.0/24 -o vmbr0` — added to
|
|
give the containers internet access — also gave them the LAN, translated to the
|
|
host's address.
|
|
**Correction:** A `RETURN` in `nat POSTROUTING` ahead of both masquerades, and
|
|
a `DROP` in `FORWARD`, both scoped `-s 10.20.0.0/24 -d 10.0.0.0/24`, plus an
|
|
explicit `ACCEPT` for container-to-container traffic.
|
|
**Consequence:** **Interface removal is not isolation.** Automation must assert
|
|
the negative — that the LAN is unreachable — not merely that the interface is
|
|
gone. Note that the containers still reach `srv-b` itself at `10.0.0.12`,
|
|
because a container addressing its gateway takes the INPUT path and never
|
|
enters FORWARD. That is required, not a leak.
|
|
|
|
---
|
|
|
|
### F-019 — placeholder backend did not survive reboot
|
|
CT 100.
|
|
|
|
**Observed:** After the F-017 reboot, nothing listened on `10.20.0.10:8770`.
|
|
**Cause:** The placeholder was a bare foreground process with a PID file. No
|
|
supervision.
|
|
**Correction:** Recreated as `mechcomp-placeholder.service`.
|
|
**Consequence:** Anything a proof depends on must be supervised, or the proof
|
|
expires silently at the next reboot and the following session diagnoses a
|
|
proxy fault that does not exist. Applies to temporary scaffolding as much as to
|
|
real services.
|
|
**Resolution (2026-08-16):** Installed `/usr/local/libexec/mechcomp-placeholder.py`
|
|
and `/etc/systemd/system/mechcomp-placeholder.service`, enabled at
|
|
`multi-user.target`, running as `mechcomp:mechcomp`, reading
|
|
`/etc/mechcomp/mechcomp.env`, binding `10.20.0.10:8770`, with `Restart=on-failure`
|
|
and the standard hardening set. CT 100 was rebooted: the unit restarted
|
|
automatically, the listener returned on the service address only, nginx proxied
|
|
successfully, `X-Forwarded-Proto: https` was re-observed at the backend, and the
|
|
container settled to `running` with zero failed units.
|
|
|
|
---
|
|
|
|
### F-020 — service FQDN mapping changed twice
|
|
`srv-b`, CT 100, CT 101.
|
|
|
|
**Observed:** `mechanical-compiler.dev.infra` was mapped to `10.20.0.10`, then
|
|
corrected to `10.0.0.21`, then corrected again to `10.20.0.11`.
|
|
**Cause:** Three different states, each correct for its moment. `10.20.0.10`
|
|
was wrong — it pointed at the application rather than the proxy. `10.0.0.21`
|
|
was right while CT 101 had a LAN address and workstation access was assumed.
|
|
`10.20.0.11` became right once F-017 removed the LAN interface and access moved
|
|
to WireGuard through `srv-b`.
|
|
**Correction:** `10.20.0.11` on all three hosts.
|
|
**Consequence:** Not an error in either direction — a value that tracked a
|
|
topology decision. Recorded because the reasoning matters more than the value:
|
|
**the service FQDN names the proxy, and the proxy has exactly one address.**
|
|
|
|
---
|
|
|
|
### F-021 — deleting the Proxmox interface left a stale guest `eth1`
|
|
CT 100, CT 101. Network isolation.
|
|
|
|
**Observed:** After the F-017 topology change and reboot, both containers
|
|
reported `systemctl is-system-running` → `degraded`, with
|
|
`networking.service` and `ifupdown-wait-online.service` failed. Both guests'
|
|
`/etc/network/interfaces` still carried an `auto eth1` / static `iface eth1`
|
|
stanza although neither had an `eth1` link. The boot journal on both:
|
|
|
|
```
|
|
Cannot find device "eth1"
|
|
ifup: failed to bring up eth1
|
|
```
|
|
|
|
**Cause:** **Proven.** `pct set --delete net1` removes the LXC interface but
|
|
does not remove the stanza already written into the guest. `ifup -a` therefore
|
|
exited non-zero at boot even though `eth0` came up correctly.
|
|
`ifupdown-wait-online` failed as a consequence, not independently.
|
|
`systemd-networkd` was investigated and explicitly ruled out — its units were
|
|
disabled on both containers.
|
|
**Correction:** Preserved a pre-correction copy, removed only the stale `eth1`
|
|
stanza on each guest, restarted the affected units. Both containers then
|
|
rebooted: the stanza did not return, both units succeeded, both settled to
|
|
`running` with zero failed units.
|
|
**Consequence:** **Removing a container interface is not proof that the guest
|
|
converged.** After any topology mutation, automation must inspect the guest
|
|
interface file and assert `systemctl is-system-running = running` with zero
|
|
failed units *after a reboot*. Connectivity alone is insufficient — the
|
|
surviving interface works fine while boot remains degraded, which is precisely
|
|
how this went unnoticed through an entire verification pass.
|
|
|
|
---
|
|
|
|
### F-022 — transient DNS resolution timeout in CT 101
|
|
CT 101. Formal acceptance.
|
|
|
|
**Observed:** During the first acceptance pass, two unrelated public HTTPS
|
|
tests both failed at name resolution:
|
|
|
|
```
|
|
deb.debian.org: curl: (28) Resolving timed out after 5000 ms
|
|
gitea.barternetwork.us: curl: (28) Resolving timed out after 5000 ms
|
|
```
|
|
|
|
IP routing to `10.110.0.1`, `10.0.0.12` and CT 100 remained working throughout.
|
|
**Cause:** **Unproven.** Diagnostics showed resolver configuration identical to
|
|
CT 100 and the host (`nameserver 75.75.75.75`, `hosts: files dns`),
|
|
`systemd-resolved` absent, `75.75.75.75` reachable, and `getent ahostsv4`
|
|
resolving both names immediately afterward. Repeat HTTPS tests returned 200.
|
|
|
|
Worth noting as context, not as cause: since F-017, container DNS traverses the
|
|
host's masquerade to an external resolver. That dependency is new.
|
|
**Correction:** None. No configuration was changed.
|
|
**Consequence:** Do not convert a one-shot resolver timeout into a
|
|
configuration change without evidence. On recurrence, capture resolver state
|
|
and DNS traffic at the moment of failure before touching anything. A candidate
|
|
mitigation — a second `nameserver` line, so a single hiccup retries rather than
|
|
fails — is recorded but deliberately not applied on one unexplained event.
|
|
|
|
---
|
|
|
|
### F-023 — relay accepted the mail; final delivery failed
|
|
`srv-b`. Alerting.
|
|
|
|
**Observed:** The relay was discovered at `10.110.0.1:25` over `wg0`, banner
|
|
`wg-pk.diagnostics.kane-il.us`, offering STARTTLS with a self-signed
|
|
`CN = wg-pk`. It accepted unauthenticated SMTP from `10.110.0.12` through
|
|
`RCPT TO`, before and after STARTTLS. Ports 465 and 587 were unavailable, and
|
|
the public address `198.58.111.109` did not expose SMTP on this path.
|
|
|
|
With `relayhost = [10.110.0.1]:25` and `root: sandor@kane-il.us`, `srv-b`
|
|
recorded successful handoff:
|
|
|
|
```
|
|
relay=10.110.0.1[10.110.0.1]:25 dsn=2.0.0
|
|
status=sent (250 2.0.0 Ok: queued as 87E446243A)
|
|
```
|
|
|
|
The local queue emptied. The operator then received a **delivery-failure**
|
|
message at `sandor@kane-il.us`.
|
|
**Cause:** **Unproven**, downstream of the demonstrated handoff. The bounce
|
|
notice itself arriving at `sandor@kane-il.us` establishes that the relay can
|
|
deliver to that address, which narrows the problem to the failing message
|
|
rather than the destination. Leading hypothesis, untested: the envelope sender
|
|
is `root@srv-b.dev.infra`, and `dev.infra` does not resolve publicly, so a
|
|
downstream MTA rejects on sender-domain verification. Candidate remedies are
|
|
`myorigin` or `smtp_generic_maps`.
|
|
**Correction:** None. Mail alerting was deferred by operator decision.
|
|
**Consequence:** **SMTP 250 from the relay and an empty local queue prove
|
|
handoff, not delivery.** Alerting acceptance requires demonstrated end-to-end
|
|
receipt. Until then mail, `smartd` alerting, and any mail-dependent backup
|
|
alerting are unaccepted. Note separately that `postfix check` reports
|
|
divergence between `/var/spool/postfix` copies and their host originals,
|
|
including `/etc/hosts` and NSS libraries — a known cause of resolution failure
|
|
inside the chroot, and adjacent enough to this failure to be checked first.
|
|
|
|
**Resolution (2026-08-17):** Cause proven, three hops downstream of the origin.
|
|
`mx1` runs `permit_mynetworks, permit_auth_destination, reject` on both relay
|
|
and recipient restrictions. `wg-pk` connected from public addresses absent from
|
|
`mx1`'s `mynetworks`, and `kane-il.us` is not an authorised destination there,
|
|
so `RCPT TO <sandor@kane-il.us>` was rejected `554 5.7.1 Access denied` over
|
|
both IPv4 and IPv6. Added only `198.58.111.109/32` and
|
|
`[2600:3c00::f03c:92ff:fe42:43d7]/128` to `mx1`'s `mynetworks` and reloaded.
|
|
|
|
Four hypotheses were tested and ruled out with evidence before any change: the
|
|
`srv-b` alias (`postalias -q root` resolved correctly and the journal showed
|
|
`orig_to=<root>` forwarded), sender-domain rejection (`wg-pk` rewrites the
|
|
sender to `postmaster@diagnostics.kane-il.us` and `MAIL FROM` was accepted),
|
|
address-family asymmetry (both families failed identically), and routing on
|
|
`wg-pk` (it connected and received the rejection).
|
|
|
|
Method worth reusing: **RCPT-only probes before and after the change, over both
|
|
address families, with no message body.** `554` before, `250` after. That
|
|
proves the change caused the fix rather than coinciding with it — the standard
|
|
F-012 was written to enforce. Two messages then delivered end to end with full
|
|
headers captured, empty queue, no deferred entries. **Closed.**
|
|
|
|
---
|
|
|
|
### F-024 — work order specified a log file that does not exist
|
|
`srv-b`. Work order 002.
|
|
|
|
**Observed:** WORK ORDER 002 instructed `grep /var/log/mail.log`. The file does
|
|
not exist on `srv-b`.
|
|
**Cause:** **Proven.** Proxmox VE ships without `rsyslog`. Postfix logs to
|
|
journald only.
|
|
**Correction:** The operator used `journalctl -u postfix@-` and completed the
|
|
diagnosis. No configuration was changed; installing `rsyslog` to satisfy a
|
|
document would have been the wrong direction.
|
|
**Consequence:** **Specification defect.** All log inspection on a Proxmox host
|
|
uses `journalctl`, not files under `/var/log`. Corrected in `ENVIRONMENT.md`
|
|
revision 5. Automation that greps a logfile path will silently find nothing,
|
|
which is worse than failing.
|
|
|
|
---
|
|
|
|
### F-025 — the F-023 correction granted wider relay than required
|
|
`mx1`, `wg-pk`. Estate scope.
|
|
|
|
**Observed:** Adding `wg-pk`'s two public addresses to `mx1`'s `mynetworks`
|
|
authorises them under `permit_mynetworks`, which appears in **both**
|
|
`smtpd_relay_restrictions` and `smtpd_recipient_restrictions`. That grants relay
|
|
to **any** destination, not only to `kane-il.us`.
|
|
|
|
`wg-pk` in turn carries `mynetworks = 10.110.0.0/22` with
|
|
`permit_mynetworks permit_sasl_authenticated defer_unauth_destination`. The
|
|
resulting chain:
|
|
|
|
```
|
|
any peer on 10.110.0.0/22 (including CT 100 and CT 101, which arrive
|
|
as 10.110.0.12 through the srv-b masquerade)
|
|
-> wg-pk permit_mynetworks
|
|
-> mx1 permit_mynetworks
|
|
-> any destination on the internet, as kane-il.us infrastructure
|
|
```
|
|
|
|
Before the correction `mx1` rejected at the final hop, so the path was closed by
|
|
accident rather than by policy.
|
|
|
|
**Cause:** **Proven.** `permit_mynetworks` is destination-agnostic by design.
|
|
**Correction:** **Pending — operator decision.** Two independent parts:
|
|
|
|
- *Ours:* block SMTP egress from the container network at `srv-b`. The
|
|
containers have no reason to originate mail; if they ever should, that is a
|
|
deliberate decision rather than something inherited from a masquerade. Local,
|
|
precise, touches no estate policy.
|
|
- *Estate:* whether `wg-pk`'s `mynetworks` should be explicit `/32` entries for
|
|
the hosts that legitimately originate mail rather than the whole tunnel range.
|
|
|
|
**Constraint:** `wg-pk` must continue to relay to arbitrary external
|
|
destinations — that is its purpose, since Hubzilla registration and notification
|
|
mail depends on it and the ISP blocks port 25. Restricting it by *recipient*
|
|
would break that. The question is which clients may ask, not where it may send.
|
|
|
|
**Consequence:** A correction that fixes the observed failure may widen an
|
|
adjacent boundary. Record what a trust change grants, not only what it repairs.
|
|
|
|
**Project-local correction (2026-08-17, WO-003):** Complete. A persistent
|
|
`srv-b` FORWARD rule drops TCP 25, 465 and 587 from `10.20.0.0/24` to
|
|
`10.110.0.0/22`. Reachability was proven **before** the change, both containers
|
|
are blocked after it, ordinary internet egress is intact, host-originated mail
|
|
still delivers end to end, and the rule survived a host reboot.
|
|
|
|
**Estate half: still open.** Whether `wg-pk` should narrow `mynetworks` from
|
|
`10.110.0.0/22` to the explicit hosts that legitimately originate mail is a
|
|
CIVICVS decision. No change was made to `wg-pk`, `mx1` or `kane-il.us`.
|
|
|
|
F-025 is therefore **partially corrected**, not closed. The two halves are kept
|
|
in one entry rather than split, because the exposure is a single chain and
|
|
splitting it would let one half be closed while the other is forgotten.
|
|
|
|
---
|
|
|
|
### F-026 — containers were not configured to start with the host
|
|
`srv-b`, CT 100, CT 101. Work order 003.
|
|
|
|
**Observed:** After the WO-003 persistence reboot — the first true **host**
|
|
reboot since the containers were built — both were stopped:
|
|
|
|
```
|
|
pct status 100 -> status: stopped
|
|
pct status 101 -> status: stopped
|
|
```
|
|
|
|
**Cause:** **Proven.** Neither container configuration contained `onboot: 1`, so
|
|
Proxmox had no instruction to start either guest. `pve-guests.service` was
|
|
enabled, active, and its `startall` task completed successfully — proving the
|
|
startup machinery worked and simply had nothing to do.
|
|
**Correction:** `pct set 100 -onboot 1` and `pct set 101 -onboot 1`. The
|
|
containers were deliberately left stopped so the correction could be proven by
|
|
reboot rather than masked by a manual start. After a second host reboot both
|
|
returned automatically and host plus both guests reported `running`.
|
|
**Consequence:** **Specification defect.** `ENVIRONMENT.md` never required
|
|
`onboot: 1`, and every earlier reboot in this project was `pct reboot` — which
|
|
restarts a guest without ever exercising host-boot autostart. A container that
|
|
survives `pct reboot` is not thereby proven to survive a host reboot. Gate 1
|
|
must assert guest state after a **host** reboot specifically.
|
|
|
|
---
|
|
|
|
### F-027 — a verification test conflated tool failure with the tested condition
|
|
`srv-b`. Work order 003.
|
|
|
|
**Observed:** WO-003 B3 and B4 used this pattern to test SMTP reachability from
|
|
a container:
|
|
|
|
```bash
|
|
pct exec $c -- timeout 5 bash -c '...' \
|
|
&& echo "REACHABLE — WRONG" || echo "blocked — correct"
|
|
```
|
|
|
|
With the containers stopped (F-026) it produced:
|
|
|
|
```
|
|
-- CT 100: container '100' not running!
|
|
blocked — correct
|
|
```
|
|
|
|
**Cause:** **Proven.** The shell `||` branch catches *any* non-zero exit from
|
|
`pct exec`, including failures of `pct exec` itself. The test could not
|
|
distinguish "the firewall blocked the connection" from "the command never ran."
|
|
**Correction:** Assert guest state before interpreting any network result:
|
|
|
|
```bash
|
|
if [ "$(pct status "$c" | awk '{print $2}')" != "running" ]; then
|
|
echo "CT $c NOT RUNNING — TEST INVALID"
|
|
elif pct exec "$c" -- timeout 5 bash -c '...'; then
|
|
echo "REACHABLE — WRONG"
|
|
else
|
|
echo "blocked — correct"
|
|
fi
|
|
```
|
|
|
|
**Consequence:** **A test whose failure mode is indistinguishable from success
|
|
is worse than no test**, because it manufactures confidence. This one printed a
|
|
pass while the thing under test did not exist. Only a separate health assertion
|
|
caught it, and that was ordering luck rather than design. Every negative
|
|
assertion — "X is unreachable", "Y is absent", "Z is refused" — must first
|
|
establish that the test could have observed the positive case.
|
|
|
|
Applies retroactively: the isolation proofs in F-018 used the same shape. They
|
|
happened to be run against containers that were up, and the result was
|
|
independently corroborated, but the pattern was unsafe there too.
|
|
|
|
---
|
|
|
|
### F-028 — three command defects in work order 003
|
|
`srv-b`. Work order 003. Minor, grouped.
|
|
|
|
**Observed and corrected in place by the operator:**
|
|
|
|
| Defect | Cause | Correction |
|
|
|---|---|---|
|
|
| Serial numbers absent from A1 output | `grep -E "Serial Number"` is case-sensitive; SCSI output says `Serial number:` | `grep -iE` with a case-insensitive alternation on product, model, serial and SMART support |
|
|
| `journalctl -u smartd -b` returned `-- No entries --` | `smartd.service` is an alias; the real unit is `smartmontools.service` and owns the journal | Query `smartmontools` |
|
|
| A3 `sed` would have produced a malformed directive | Prefix substitution left the trailing `-M exec ...` in place and duplicated `-m root` | Replace the whole `DEFAULT` line, in both A3 and A4 |
|
|
|
|
**Consequence:** Match on the canonical unit name, not a convenience alias.
|
|
Prefer whole-line replacement over prefix patching in configuration edits. Use
|
|
case-insensitive matching when parsing tool output whose field names vary by
|
|
device class.
|
|
|
|
---
|
|
|
|
### F-029 — pip cache written into the repository working tree
|
|
CT 100. Repository seeding.
|
|
|
|
**Observed:** `git add -A` staged 465 files where 28 were expected. The extra
|
|
437 were `.cache/` — pip's download cache.
|
|
**Cause:** **Proven.** The service user's home is the install directory, so
|
|
`$HOME/.cache` **is** `/var/www/mechcomp/.cache`. `make deps` wrote there.
|
|
**Correction:** `.cache/`, `.local/`, `.ssh/`, `.gitconfig` and `.lesshst`
|
|
added to `.gitignore`; the cache unstaged before committing.
|
|
**Consequence:** Setting the service user's home to `install_dir` is correct
|
|
and matches YunoHost convention, but it means **any tool writing to `$HOME`
|
|
writes into the working tree**. `.gitignore` must anticipate that from the
|
|
first commit.
|
|
|
|
Method note: reading the staged list by eye showed nothing wrong. Counting it
|
|
and grouping by top-level directory found the problem immediately. **Verify by
|
|
counting, not by scanning** — a long correct-looking list is exactly where an
|
|
extra 437 files hide.
|
|
|
|
---
|
|
|
|
### F-030 — a network command in a work order could hang indefinitely
|
|
`srv-b`. Repository seeding.
|
|
|
|
**Observed:** An SSH authentication test was issued with neither
|
|
`ConnectTimeout` nor `BatchMode`. Gitea's SSH is on port 42022, so the attempt
|
|
on 22 hung and the operator had to interrupt it.
|
|
**Cause:** **Proven.** Missing timeout options, and an incorrect assumption
|
|
about the port.
|
|
**Correction:** Re-issued with `-o ConnectTimeout=10 -o BatchMode=yes -p 42022`,
|
|
after a `timeout`-wrapped reachability probe that could not hang.
|
|
**Consequence:** **Every network command in a work order carries an explicit
|
|
timeout.** A command that can hang strands the session, and the operator cannot
|
|
tell a hang from slow progress. Reachability is probed with `timeout` before any
|
|
client is invoked.
|
|
|
|
Also recorded, having been got wrong twice: Gitea **deploy tokens** are
|
|
account-level under user Settings, not repository settings. Repository-scoped
|
|
credentials are **deploy keys**. SSH here is on **42022**, so remotes need
|
|
`ssh://git@host:42022/owner/repo.git` — the `git@host:path` shorthand cannot
|
|
carry a port.
|
|
|
|
---
|
|
|
|
### F-031 — a standard was asserted from a check that never examined the thing
|
|
CT 100, CT 101, CT 102. Container standardisation.
|
|
|
|
**Observed:** The architect stated that CT 100 and CT 101 had no mail agent, and
|
|
standardised CT 102 by purging Postfix to match. The first run of
|
|
`ct-baseline.sh` reported a mail agent installed on **both** CT 100 and CT 101.
|
|
CT 102 — the container just "corrected" — was the only one conforming.
|
|
|
|
**Cause:** **Proven.** The earlier assertion came from inspecting Postfix
|
|
configuration and `/etc/aliases` **on `srv-b`**, during work order 002. That
|
|
established what the host does. It said nothing about the containers, and the
|
|
containers were never examined.
|
|
|
|
**Correction:** Postfix purged from CT 100 and CT 101. Queues were checked first
|
|
and both were empty, so nothing was lost. Re-run: 62 passed, 0 failed.
|
|
|
|
**Consequence:** **A standard asserted from a proxy observation is not a
|
|
standard.** The check must examine the thing being standardised. This is the
|
|
same class as F-027 — a result that cannot distinguish the case it claims to
|
|
test — but reached through inference rather than through shell semantics.
|
|
|
|
It is also the justification for `ct-baseline.sh` existing. Divergence had been
|
|
discovered by asking, one property at a time, whenever something behaved oddly.
|
|
That does not terminate: every check finds a new difference because nothing
|
|
states what "the same" means. The script is that statement, and it disagreed
|
|
with the architect on its first run.
|
|
|
|
---
|
|
|
|
### F-032 — the same host defects were discovered independently by two projects
|
|
`srv-b`. Cross-project.
|
|
|
|
**Observed:** Kane Fabric's `INFRASTRUCTURE_BASELINE.md`, written independently,
|
|
records `nesting=1,keyctl=1` against `226/NAMESPACE` systemd failures, and that
|
|
minimal Debian CTs lack a generated `en_US.UTF-8`. These are F-003 and F-004,
|
|
found again on the same host, in different words, at a different time.
|
|
**Cause:** **Proven.** Both are properties of the Proxmox/Debian 12 host, not of
|
|
either application. Neither project's documentation was visible to the other.
|
|
**Correction:** None required — both projects reached the correct conclusion.
|
|
**Consequence:** **Host-level requirements belong to the host, not to whichever
|
|
project discovered them first.** `ct-baseline.sh` encodes them once and checks
|
|
every container on `srv-b` regardless of owner, so the third project does not
|
|
have to discover them a third time.
|
|
|
|
Note that `keyctl=1` is required **only where Docker runs**. CT 101 was proven
|
|
to work with `nesting=1` alone (F-003). Kane Fabric's baseline currently
|
|
specifies both; if CT 102 runs no containers, `keyctl` can be dropped.
|
|
|
|
---
|
|
|
|
## Open, not closed
|
|
|
|
| # | Status |
|
|
|---|---|
|
|
| F-006 | **Open** — cause unproven. Recurrence should capture `dpkg` lock state. |
|
|
| F-012 | **Open** — cause unproven. Leading candidate ruled out by inspection. `default_server` was added as independent hardening and does **not** close this. |
|
|
| F-019 | **Corrected** 2026-08-16. Reboot persistence proven. |
|
|
| F-021 | **Corrected** 2026-08-16. Reboot persistence proven. |
|
|
| F-022 | **Open** — cause unproven, no correction applied. |
|
|
| F-023 | **Closed** 2026-08-17. Cause proven at `mx1`; delivery proven twice. |
|
|
| F-024 | **Closed** 2026-08-17. Specification corrected. |
|
|
| F-025 | **Partially corrected** 2026-08-17. Project-local half closed and reboot-proven; **estate half open**, awaiting a CIVICVS decision on `wg-pk`. |
|
|
| F-026 | **Corrected** 2026-08-17. Reboot persistence proven. |
|
|
| F-027 | **Corrected** 2026-08-17. Applies retroactively to F-018's proofs. |
|
|
| F-028 | **Corrected** 2026-08-17. |
|
|
| F-029 | **Corrected** 2026-08-18. |
|
|
| F-030 | **Corrected** 2026-08-18. |
|
|
| F-031 | **Corrected** 2026-08-18. Root cause of `ct-baseline.sh`. |
|
|
| F-032 | **Closed** 2026-08-18. No correction required; encoded in `ct-baseline.sh`. |
|
|
|
|
Everything else is closed with a proven cause and a proven correction.
|