This commit is contained in:
2026-08-16 20:25:18 -04:00
parent eac601ed99
commit e67988eef5
4 changed files with 727 additions and 828 deletions
+116 -4
View File
@@ -6,6 +6,7 @@ Compiler environment.
| | |
|---|---|
| Scope | All instances. Staging entries are marked `srv-b`. |
| Updated | 2026-08-16, after staging acceptance |
| Rule | Append only. Never edit an entry except to add a `Resolution` line. |
| Numbering | Sequential, never reused. See §0 on the renumbering. |
@@ -315,11 +316,19 @@ CT 100.
**Observed:** After the F-017 reboot, nothing listened on `10.20.0.10:8770`.
**Cause:** The placeholder was a bare foreground process with a PID file. No
supervision.
**Correction:** Pending — to be recreated as `mechcomp-placeholder.service`.
**Correction:** Recreated as `mechcomp-placeholder.service`.
**Consequence:** Anything a proof depends on must be supervised, or the proof
expires silently at the next reboot and the following session diagnoses a
proxy fault that does not exist. Applies to temporary scaffolding as much as to
real services.
**Resolution (2026-08-16):** Installed `/usr/local/libexec/mechcomp-placeholder.py`
and `/etc/systemd/system/mechcomp-placeholder.service`, enabled at
`multi-user.target`, running as `mechcomp:mechcomp`, reading
`/etc/mechcomp/mechcomp.env`, binding `10.20.0.10:8770`, with `Restart=on-failure`
and the standard hardening set. CT 100 was rebooted: the unit restarted
automatically, the listener returned on the service address only, nginx proxied
successfully, `X-Forwarded-Proto: https` was re-observed at the backend, and the
container settled to `running` with zero failed units.
---
@@ -340,10 +349,113 @@ topology decision. Recorded because the reasoning matters more than the value:
---
### F-021 — deleting the Proxmox interface left a stale guest `eth1`
CT 100, CT 101. Network isolation.
**Observed:** After the F-017 topology change and reboot, both containers
reported `systemctl is-system-running` → `degraded`, with
`networking.service` and `ifupdown-wait-online.service` failed. Both guests'
`/etc/network/interfaces` still carried an `auto eth1` / static `iface eth1`
stanza although neither had an `eth1` link. The boot journal on both:
```
Cannot find device "eth1"
ifup: failed to bring up eth1
```
**Cause:** **Proven.** `pct set --delete net1` removes the LXC interface but
does not remove the stanza already written into the guest. `ifup -a` therefore
exited non-zero at boot even though `eth0` came up correctly.
`ifupdown-wait-online` failed as a consequence, not independently.
`systemd-networkd` was investigated and explicitly ruled out — its units were
disabled on both containers.
**Correction:** Preserved a pre-correction copy, removed only the stale `eth1`
stanza on each guest, restarted the affected units. Both containers then
rebooted: the stanza did not return, both units succeeded, both settled to
`running` with zero failed units.
**Consequence:** **Removing a container interface is not proof that the guest
converged.** After any topology mutation, automation must inspect the guest
interface file and assert `systemctl is-system-running = running` with zero
failed units *after a reboot*. Connectivity alone is insufficient — the
surviving interface works fine while boot remains degraded, which is precisely
how this went unnoticed through an entire verification pass.
---
### F-022 — transient DNS resolution timeout in CT 101
CT 101. Formal acceptance.
**Observed:** During the first acceptance pass, two unrelated public HTTPS
tests both failed at name resolution:
```
deb.debian.org: curl: (28) Resolving timed out after 5000 ms
gitea.barternetwork.us: curl: (28) Resolving timed out after 5000 ms
```
IP routing to `10.110.0.1`, `10.0.0.12` and CT 100 remained working throughout.
**Cause:** **Unproven.** Diagnostics showed resolver configuration identical to
CT 100 and the host (`nameserver 75.75.75.75`, `hosts: files dns`),
`systemd-resolved` absent, `75.75.75.75` reachable, and `getent ahostsv4`
resolving both names immediately afterward. Repeat HTTPS tests returned 200.
Worth noting as context, not as cause: since F-017, container DNS traverses the
host's masquerade to an external resolver. That dependency is new.
**Correction:** None. No configuration was changed.
**Consequence:** Do not convert a one-shot resolver timeout into a
configuration change without evidence. On recurrence, capture resolver state
and DNS traffic at the moment of failure before touching anything. A candidate
mitigation — a second `nameserver` line, so a single hiccup retries rather than
fails — is recorded but deliberately not applied on one unexplained event.
---
### F-023 — relay accepted the mail; final delivery failed
`srv-b`. Alerting.
**Observed:** The relay was discovered at `10.110.0.1:25` over `wg0`, banner
`wg-pk.diagnostics.kane-il.us`, offering STARTTLS with a self-signed
`CN = wg-pk`. It accepted unauthenticated SMTP from `10.110.0.12` through
`RCPT TO`, before and after STARTTLS. Ports 465 and 587 were unavailable, and
the public address `198.58.111.109` did not expose SMTP on this path.
With `relayhost = [10.110.0.1]:25` and `root: sandor@kane-il.us`, `srv-b`
recorded successful handoff:
```
relay=10.110.0.1[10.110.0.1]:25 dsn=2.0.0
status=sent (250 2.0.0 Ok: queued as 87E446243A)
```
The local queue emptied. The operator then received a **delivery-failure**
message at `sandor@kane-il.us`.
**Cause:** **Unproven**, downstream of the demonstrated handoff. The bounce
notice itself arriving at `sandor@kane-il.us` establishes that the relay can
deliver to that address, which narrows the problem to the failing message
rather than the destination. Leading hypothesis, untested: the envelope sender
is `root@srv-b.dev.infra`, and `dev.infra` does not resolve publicly, so a
downstream MTA rejects on sender-domain verification. Candidate remedies are
`myorigin` or `smtp_generic_maps`.
**Correction:** None. Mail alerting was deferred by operator decision.
**Consequence:** **SMTP 250 from the relay and an empty local queue prove
handoff, not delivery.** Alerting acceptance requires demonstrated end-to-end
receipt. Until then mail, `smartd` alerting, and any mail-dependent backup
alerting are unaccepted. Note separately that `postfix check` reports
divergence between `/var/spool/postfix` copies and their host originals,
including `/etc/hosts` and NSS libraries — a known cause of resolution failure
inside the chroot, and adjacent enough to this failure to be checked first.
---
## Open, not closed
| # | Status |
|---|---|
| F-006 | Cause unproven. Recurrence should capture `dpkg` lock state. |
| F-012 | Cause unproven. Leading candidate ruled out by inspection. |
| F-019 | Correction pending. |
| F-006 | **Open** — cause unproven. Recurrence should capture `dpkg` lock state. |
| F-012 | **Open** — cause unproven. Leading candidate ruled out by inspection. `default_server` was added as independent hardening and does **not** close this. |
| F-019 | **Corrected** 2026-08-16. Reboot persistence proven. |
| F-021 | **Corrected** 2026-08-16. Reboot persistence proven. |
| F-022 | **Open** — cause unproven, no correction applied. |
| F-023 | **Deferred** — downstream mail failure, cause unproven. Hypothesis recorded. |
Everything else is closed with a proven cause and a proven correction.