No behaviour change. Comments and a FAILURES entry. The Y profile builds a section area of 135.572973 against a recorded 135.574, out by 0.001027, while every other value for that case matches exactly. Three-Fin matches on everything including area. Isolated by probing the reference inside the pinned toolchain image. The hull cap is identical to nine figures, the bare union is identical, and a single filleted pair is identical at 30 vertices and 125.699057 mm2. The difference appears only when the three filleted pairs are combined, and the three pairs, which are related by 120 degree symmetry and must be identical, come back as 125.699057, 125.698029, 125.699057. Cause proven. The arc segment count is a ceiling on a quantity that is frequently an exact integer: a 60 degree half-angle at $fn=48 gives exactly 8. Floating point delivers that as 8.000000000000004 on one corner and 7.999999999999998 on the others, so one corner gets a whole extra segment. The half-angles come from the merged polygon, whose vertices come from the boolean kernel, and BOSL2 clipper and GEOS disagree in the last bit. Reproduced unguarded because the reference is unguarded. Rounding the count before the ceiling was implemented and reverted: it makes the three pairs identical and fixes Y exactly, and breaks Three-Fin, which had been matching to the digit. Three-Fin has the same asymmetry and the oracle records it. BOSL2 tipped the same way GEOS does there and the opposite way on Y. Two consequences for the project rather than the code. Some recorded values encode float noise rather than geometry, so a port that is geometrically more correct than the reference will fail those cases. And the tolerance model may need revisiting: VOLUME_MM3 is compared at the lengths tolerance of 1e-4 despite being area times 100 mm, so a 1e-3 area difference becomes a 1e-1 volume difference. MASS_G is derived the same way. No decision yet. The number of affected cases is unknown and is the only thing that should drive it, and that is not knowable until build() exists and all 123 cases can run.
862 lines
38 KiB
Markdown
862 lines
38 KiB
Markdown
# FAILURES.md
|
|
|
|
Append-only record of every failure encountered building the Mechanical
|
|
Compiler environment.
|
|
|
|
| | |
|
|
|---|---|
|
|
| Scope | All instances. Staging entries are marked `srv-b`. |
|
|
| Updated | 2026-08-18, after repository seeding and container standardisation |
|
|
| Method | `PROCESS.md` section 7 |
|
|
| Rule | Append only. Never edit an entry except to add a `Resolution` line. |
|
|
| Numbering | Sequential, never reused. See §0 on the renumbering. |
|
|
|
|
---
|
|
|
|
## 0. Why this file exists separately
|
|
|
|
This is the primary input to provisioning automation.
|
|
|
|
Every entry below is something a script written from `ENVIRONMENT.md` alone
|
|
would have got wrong. `F-003` and `F-004` in particular would have produced a
|
|
container that booted, reported success, and was quietly broken. The
|
|
specification cannot anticipate them; only contact with a host reveals them.
|
|
|
|
When the production automation is written, it should be read **before**
|
|
`ENVIRONMENT.md`, not after.
|
|
|
|
### Renumbering note, 2026-08-15
|
|
|
|
The staging checkpoint and the reconciliation checkpoint were written
|
|
independently and both allocated `F-008` and `F-009`. The reconciliation
|
|
entries have been renumbered to `F-015` and `F-016`. Original phase-1 numbering
|
|
is preserved because it is cited elsewhere.
|
|
|
|
### Entry format
|
|
|
|
```
|
|
### F-nnn — one-line summary
|
|
Host / container. Phase.
|
|
**Observed:** what was seen, verbatim where possible.
|
|
**Cause:** proven cause, or explicitly "unproven".
|
|
**Correction:** the smallest change that fixed it.
|
|
**Consequence:** what the specification or automation must do differently.
|
|
```
|
|
|
|
A failure with no `Consequence` line is either not yet understood or not worth
|
|
recording.
|
|
|
|
---
|
|
|
|
## Phase 1 — initial container build
|
|
|
|
### F-001 — `arping` absent on the Proxmox host
|
|
`srv-b`.
|
|
|
|
**Observed:** `command -v arping` → not installed.
|
|
**Cause:** Not part of the base PVE install.
|
|
**Correction:** `apt install iputils-arping`.
|
|
**Consequence:** Automation must install `iputils-arping` before duplicate
|
|
address detection, or the DAD check silently does not happen.
|
|
|
|
---
|
|
|
|
### F-002 — invalid Proxmox inspection command
|
|
`srv-b`.
|
|
|
|
**Observed:** `pvesm config local-lvm` returned CLI usage text.
|
|
**Cause:** Operator command error. Not a host fault.
|
|
**Correction:** Read `/etc/pve/storage.cfg` directly.
|
|
**Consequence:** Read storage configuration from the file, not from a
|
|
subcommand that does not exist.
|
|
|
|
---
|
|
|
|
### F-003 — CT 101 booted with degraded systemd
|
|
CT 101.
|
|
|
|
**Observed:**
|
|
|
|
```
|
|
status=226/NAMESPACE
|
|
Failed to set up mount namespacing: Permission denied
|
|
systemd-logind.service failed
|
|
systemd-networkd.service failed
|
|
systemd-timedated.service failed
|
|
systemd-networkd.socket failed
|
|
```
|
|
|
|
**Cause:** Debian 12 systemd requires mount namespacing that an unprivileged
|
|
LXC denies without `nesting`.
|
|
**Correction:** `features: nesting=1`. Deliberately tested alone rather than
|
|
copying `nesting=1,keyctl=1` from CT 100 — `keyctl=1` proved unnecessary.
|
|
**Consequence:** **Every Debian 12 unprivileged container needs
|
|
`nesting=1`**, including ones running nothing but nginx. `keyctl=1` is needed
|
|
only where Docker runs. Promoted into the specification.
|
|
|
|
This is the clearest example of why the manual build was correct: the container
|
|
started, appeared healthy at a glance, and was broken in four services.
|
|
|
|
---
|
|
|
|
### F-004 — locale configured but never generated
|
|
CT 100, CT 101.
|
|
|
|
**Observed:** `LANG=en_US.UTF-8` set, but `locale -a` listed only `C`,
|
|
`C.utf8`, `POSIX`. `locale` emitted `Cannot set LC_*` warnings.
|
|
**Cause:** Setting `LANG` does not generate a locale. The container template
|
|
ships none.
|
|
**Correction:** install `locales`, enable `en_US.UTF-8 UTF-8`, `locale-gen`,
|
|
`update-locale`.
|
|
**Consequence:** Specification must state the generation step, not just the
|
|
value. Silent until something depends on collation or encoding.
|
|
|
|
---
|
|
|
|
### F-005 — locale repair exposed a pending upgrade backlog
|
|
CT 100, CT 101.
|
|
|
|
**Observed:** Installing `locales` pulled glibc-related packages and revealed
|
|
52 further pending upgrades.
|
|
**Cause:** Template was not current.
|
|
**Correction:** Full upgrade of each container, one at a time. Both reached
|
|
Debian 12.15, zero pending, no reboot required.
|
|
**Consequence:** Automation must `apt full-upgrade` immediately after first
|
|
boot, before anything else is installed.
|
|
|
|
---
|
|
|
|
### F-006 — transient `curl` / `git` command-not-found
|
|
CT 100.
|
|
|
|
**Observed:** An egress test returned `curl: command not found` and
|
|
`git: command not found`. Immediately afterward both packages were installed
|
|
and current, `PATH` normal, both executables present and working.
|
|
**Cause:** **Unproven.** Hypothesis only: the test ran during the F-005
|
|
upgrade, while `dpkg` had the binaries briefly unlinked mid-transaction.
|
|
Adjacent in time, consistent with the symptom, not demonstrated.
|
|
**Correction:** None required. Repeat tests with absolute paths passed.
|
|
**Consequence:** Do not close. If it recurs, capture `dpkg` lock state at the
|
|
moment of failure.
|
|
|
|
---
|
|
|
|
### F-007 — recursive `chown` failed on ext4 `lost+found`
|
|
CT 100.
|
|
|
|
**Observed:** `chown: cannot read directory '/var/lib/mechcomp/lost+found':
|
|
Permission denied`. That directory is `nobody:nogroup`, mode `0700`.
|
|
**Cause:** `/var/lib/mechcomp` is a real filesystem root, not a directory.
|
|
`lost+found` is created by `mkfs` and is not ours to manage.
|
|
**Correction:** Own explicit application paths only. Leave `lost+found` alone.
|
|
**Consequence:** **Never `chown -R` a mount point root.** Enumerate the
|
|
directories the application actually uses. Promoted into the specification.
|
|
|
|
---
|
|
|
|
### F-008 — Git "dubious ownership"
|
|
CT 100.
|
|
|
|
**Observed:** `fatal: detected dubious ownership in repository at
|
|
'/var/www/mechcomp'` when verification ran as root.
|
|
**Cause:** Repository cloned as `mechcomp`, inspected as `root`.
|
|
**Correction:** Re-ran verification as the owning user. **No root
|
|
`safe.directory` exception was added** — the correct fix, since the exception
|
|
would have masked every future instance of the same mistake.
|
|
**Consequence:** All repository operations run as the service user.
|
|
|
|
---
|
|
|
|
### F-009 — dependency manifests absent from the application repository
|
|
CT 100.
|
|
|
|
**Observed:** `requirements-base.txt` and `requirements-cad.txt` do not exist;
|
|
the repository is `LICENSE` and `README.md` at commit `e85c4f4e`.
|
|
**Cause:** Application implementation has not started. Not a provisioning
|
|
failure.
|
|
**Correction:** None. Do not invent pins.
|
|
**Consequence:** Infrastructure and application acceptance are separately
|
|
gated. Infrastructure may complete without the application existing.
|
|
|
|
---
|
|
|
|
## Phase 2 — architect corrections
|
|
|
|
### F-010 — placeholder PID file rewrite failed to substitute
|
|
CT 100.
|
|
|
|
**Observed:** An automated rewrite left the literal `$!` in the PID file.
|
|
**Cause:** Quoting error in the rewriting command.
|
|
**Correction:** Identified the real PID by inspection.
|
|
**Consequence:** Superseded by F-019 — the placeholder should be a systemd
|
|
unit and have no PID file at all.
|
|
|
|
---
|
|
|
|
### F-011 — root could not overwrite a `mechcomp`-owned file in `/tmp`
|
|
CT 100.
|
|
|
|
**Observed:** `Permission denied` writing a file owned by `mechcomp:mechcomp`
|
|
in `/tmp`, as root.
|
|
**Cause:** `fs.protected_regular = 2` with `/tmp` mode `1777`. Expected kernel
|
|
behaviour, not a fault.
|
|
**Correction:** Rewrote the file as the owning user.
|
|
**Consequence:** Do not use `/tmp` for state shared across users. Moot once
|
|
services run under systemd with `PrivateTmp=true`.
|
|
|
|
---
|
|
|
|
### F-012 — first request after nginx reload returned the Debian welcome page
|
|
CT 101.
|
|
|
|
**Observed:** `nginx -t` passed; the immediately following request returned the
|
|
615-byte Debian default page. All subsequent requests proxied correctly.
|
|
**Cause:** **Unproven.** The obvious candidate — Debian's `default` site still
|
|
enabled and acting as `default_server` — was tested and ruled out:
|
|
`sites-enabled` contains only the project vhost. Remaining hypothesis is a
|
|
graceful worker transition. Not demonstrated.
|
|
**Correction:** None required.
|
|
**Consequence:** Do not close. Independently, declare `default_server`
|
|
explicitly on the project vhost so that adding a second server block cannot
|
|
make catch-all behaviour depend on file ordering.
|
|
|
|
---
|
|
|
|
### F-013 — TLS test validated against an IP literal
|
|
CT 101.
|
|
|
|
**Observed:** `SSL: no alternative certificate subject name matches target
|
|
host name '10.0.0.21'`.
|
|
**Cause:** Test used an address; the certificate SAN contains a DNS name.
|
|
Correct behaviour by both curl and OpenSSL.
|
|
**Correction:** Test by hostname.
|
|
**Consequence:** Acceptance checks must use the service FQDN. An IP-literal
|
|
TLS test is always wrong unless the certificate carries an IP SAN.
|
|
|
|
---
|
|
|
|
### F-014 — CT 101 had no `mechcomp` group
|
|
CT 101.
|
|
|
|
**Observed:** Adding `sandor` to `mechcomp` failed; the group did not exist.
|
|
**Cause:** Specification said the admin joins the `mechcomp` group in *both*
|
|
containers, but the group is created as a side effect of creating the service
|
|
user, which exists only in CT 100.
|
|
**Correction:** Created the system group and added `sandor`.
|
|
**Consequence:** Specification defect. Either create the group explicitly where
|
|
it is required, or scope the group membership to CT 100. GID 996 now matches
|
|
across both containers, which is harmless and mildly useful.
|
|
|
|
---
|
|
|
|
### F-015 — CT 101 verification lacked `curl`
|
|
CT 101. *(Renumbered from a duplicate `F-008`.)*
|
|
|
|
**Observed:** `curl: command not found`.
|
|
**Cause:** Not in the base template; CT 101's package list did not include it.
|
|
**Correction:** Installed `curl` and `ca-certificates`.
|
|
**Consequence:** The proxy container needs its own minimal toolset. Do not
|
|
assume packages installed in CT 100 exist in CT 101.
|
|
|
|
---
|
|
|
|
### F-016 — placeholder PID capture stored a literal `$!`
|
|
CT 100. *(Renumbered from a duplicate `F-009`.)*
|
|
|
|
**Observed:** PID file contained `$!` rather than a number. The listener itself
|
|
was healthy.
|
|
**Cause:** Quoting error.
|
|
**Correction:** Identified PID `12102` by inspection.
|
|
**Consequence:** Superseded by F-019.
|
|
|
|
---
|
|
|
|
## Phase 3 — network isolation
|
|
|
|
### F-017 — containers were provisioned on the home LAN
|
|
`srv-b`, CT 100, CT 101.
|
|
|
|
**Observed:** Both containers had `net0` on `vmbr0` at `10.0.0.20/24` and
|
|
`10.0.0.21/24`, gateway `10.0.0.1` — the ISP router's network. Anything on the
|
|
home LAN could reach them.
|
|
**Cause:** **Specification defect in `ENVIRONMENT.md` revision 4.** The
|
|
document placed containers on both bridges, giving each a management address on
|
|
the LAN, on the unstated assumption that a LAN workstation would browse
|
|
staging. That assumption was never confirmed and was wrong.
|
|
**Correction:** Both containers moved to `vmbr1` only, gateway `10.20.0.1`.
|
|
`net1` deleted. Host NAT extended for `10.20.0.0/24`.
|
|
**Consequence:** Containers are on the portless service bridge only. `srv-b` is
|
|
router and bastion; access arrives via WireGuard. Promoted into the
|
|
specification.
|
|
|
|
---
|
|
|
|
### F-018 — removing the LAN interface did not remove LAN reachability
|
|
CT 100, CT 101.
|
|
|
|
**Observed:** After F-017, both containers still reached `10.0.0.1`
|
|
successfully.
|
|
**Cause:** Removing an interface removes an address, not a route.
|
|
`ip_forward=1` plus the new `MASQUERADE -s 10.20.0.0/24 -o vmbr0` — added to
|
|
give the containers internet access — also gave them the LAN, translated to the
|
|
host's address.
|
|
**Correction:** A `RETURN` in `nat POSTROUTING` ahead of both masquerades, and
|
|
a `DROP` in `FORWARD`, both scoped `-s 10.20.0.0/24 -d 10.0.0.0/24`, plus an
|
|
explicit `ACCEPT` for container-to-container traffic.
|
|
**Consequence:** **Interface removal is not isolation.** Automation must assert
|
|
the negative — that the LAN is unreachable — not merely that the interface is
|
|
gone. Note that the containers still reach `srv-b` itself at `10.0.0.12`,
|
|
because a container addressing its gateway takes the INPUT path and never
|
|
enters FORWARD. That is required, not a leak.
|
|
|
|
---
|
|
|
|
### F-019 — placeholder backend did not survive reboot
|
|
CT 100.
|
|
|
|
**Observed:** After the F-017 reboot, nothing listened on `10.20.0.10:8770`.
|
|
**Cause:** The placeholder was a bare foreground process with a PID file. No
|
|
supervision.
|
|
**Correction:** Recreated as `mechcomp-placeholder.service`.
|
|
**Consequence:** Anything a proof depends on must be supervised, or the proof
|
|
expires silently at the next reboot and the following session diagnoses a
|
|
proxy fault that does not exist. Applies to temporary scaffolding as much as to
|
|
real services.
|
|
**Resolution (2026-08-16):** Installed `/usr/local/libexec/mechcomp-placeholder.py`
|
|
and `/etc/systemd/system/mechcomp-placeholder.service`, enabled at
|
|
`multi-user.target`, running as `mechcomp:mechcomp`, reading
|
|
`/etc/mechcomp/mechcomp.env`, binding `10.20.0.10:8770`, with `Restart=on-failure`
|
|
and the standard hardening set. CT 100 was rebooted: the unit restarted
|
|
automatically, the listener returned on the service address only, nginx proxied
|
|
successfully, `X-Forwarded-Proto: https` was re-observed at the backend, and the
|
|
container settled to `running` with zero failed units.
|
|
|
|
---
|
|
|
|
### F-020 — service FQDN mapping changed twice
|
|
`srv-b`, CT 100, CT 101.
|
|
|
|
**Observed:** `mechanical-compiler.dev.infra` was mapped to `10.20.0.10`, then
|
|
corrected to `10.0.0.21`, then corrected again to `10.20.0.11`.
|
|
**Cause:** Three different states, each correct for its moment. `10.20.0.10`
|
|
was wrong — it pointed at the application rather than the proxy. `10.0.0.21`
|
|
was right while CT 101 had a LAN address and workstation access was assumed.
|
|
`10.20.0.11` became right once F-017 removed the LAN interface and access moved
|
|
to WireGuard through `srv-b`.
|
|
**Correction:** `10.20.0.11` on all three hosts.
|
|
**Consequence:** Not an error in either direction — a value that tracked a
|
|
topology decision. Recorded because the reasoning matters more than the value:
|
|
**the service FQDN names the proxy, and the proxy has exactly one address.**
|
|
|
|
---
|
|
|
|
### F-021 — deleting the Proxmox interface left a stale guest `eth1`
|
|
CT 100, CT 101. Network isolation.
|
|
|
|
**Observed:** After the F-017 topology change and reboot, both containers
|
|
reported `systemctl is-system-running` → `degraded`, with
|
|
`networking.service` and `ifupdown-wait-online.service` failed. Both guests'
|
|
`/etc/network/interfaces` still carried an `auto eth1` / static `iface eth1`
|
|
stanza although neither had an `eth1` link. The boot journal on both:
|
|
|
|
```
|
|
Cannot find device "eth1"
|
|
ifup: failed to bring up eth1
|
|
```
|
|
|
|
**Cause:** **Proven.** `pct set --delete net1` removes the LXC interface but
|
|
does not remove the stanza already written into the guest. `ifup -a` therefore
|
|
exited non-zero at boot even though `eth0` came up correctly.
|
|
`ifupdown-wait-online` failed as a consequence, not independently.
|
|
`systemd-networkd` was investigated and explicitly ruled out — its units were
|
|
disabled on both containers.
|
|
**Correction:** Preserved a pre-correction copy, removed only the stale `eth1`
|
|
stanza on each guest, restarted the affected units. Both containers then
|
|
rebooted: the stanza did not return, both units succeeded, both settled to
|
|
`running` with zero failed units.
|
|
**Consequence:** **Removing a container interface is not proof that the guest
|
|
converged.** After any topology mutation, automation must inspect the guest
|
|
interface file and assert `systemctl is-system-running = running` with zero
|
|
failed units *after a reboot*. Connectivity alone is insufficient — the
|
|
surviving interface works fine while boot remains degraded, which is precisely
|
|
how this went unnoticed through an entire verification pass.
|
|
|
|
---
|
|
|
|
### F-022 — transient DNS resolution timeout in CT 101
|
|
CT 101. Formal acceptance.
|
|
|
|
**Observed:** During the first acceptance pass, two unrelated public HTTPS
|
|
tests both failed at name resolution:
|
|
|
|
```
|
|
deb.debian.org: curl: (28) Resolving timed out after 5000 ms
|
|
gitea.barternetwork.us: curl: (28) Resolving timed out after 5000 ms
|
|
```
|
|
|
|
IP routing to `10.110.0.1`, `10.0.0.12` and CT 100 remained working throughout.
|
|
**Cause:** **Unproven.** Diagnostics showed resolver configuration identical to
|
|
CT 100 and the host (`nameserver 75.75.75.75`, `hosts: files dns`),
|
|
`systemd-resolved` absent, `75.75.75.75` reachable, and `getent ahostsv4`
|
|
resolving both names immediately afterward. Repeat HTTPS tests returned 200.
|
|
|
|
Worth noting as context, not as cause: since F-017, container DNS traverses the
|
|
host's masquerade to an external resolver. That dependency is new.
|
|
**Correction:** None. No configuration was changed.
|
|
**Consequence:** Do not convert a one-shot resolver timeout into a
|
|
configuration change without evidence. On recurrence, capture resolver state
|
|
and DNS traffic at the moment of failure before touching anything. A candidate
|
|
mitigation — a second `nameserver` line, so a single hiccup retries rather than
|
|
fails — is recorded but deliberately not applied on one unexplained event.
|
|
|
|
---
|
|
|
|
### F-023 — relay accepted the mail; final delivery failed
|
|
`srv-b`. Alerting.
|
|
|
|
**Observed:** The relay was discovered at `10.110.0.1:25` over `wg0`, banner
|
|
`wg-pk.diagnostics.kane-il.us`, offering STARTTLS with a self-signed
|
|
`CN = wg-pk`. It accepted unauthenticated SMTP from `10.110.0.12` through
|
|
`RCPT TO`, before and after STARTTLS. Ports 465 and 587 were unavailable, and
|
|
the public address `198.58.111.109` did not expose SMTP on this path.
|
|
|
|
With `relayhost = [10.110.0.1]:25` and `root: sandor@kane-il.us`, `srv-b`
|
|
recorded successful handoff:
|
|
|
|
```
|
|
relay=10.110.0.1[10.110.0.1]:25 dsn=2.0.0
|
|
status=sent (250 2.0.0 Ok: queued as 87E446243A)
|
|
```
|
|
|
|
The local queue emptied. The operator then received a **delivery-failure**
|
|
message at `sandor@kane-il.us`.
|
|
**Cause:** **Unproven**, downstream of the demonstrated handoff. The bounce
|
|
notice itself arriving at `sandor@kane-il.us` establishes that the relay can
|
|
deliver to that address, which narrows the problem to the failing message
|
|
rather than the destination. Leading hypothesis, untested: the envelope sender
|
|
is `root@srv-b.dev.infra`, and `dev.infra` does not resolve publicly, so a
|
|
downstream MTA rejects on sender-domain verification. Candidate remedies are
|
|
`myorigin` or `smtp_generic_maps`.
|
|
**Correction:** None. Mail alerting was deferred by operator decision.
|
|
**Consequence:** **SMTP 250 from the relay and an empty local queue prove
|
|
handoff, not delivery.** Alerting acceptance requires demonstrated end-to-end
|
|
receipt. Until then mail, `smartd` alerting, and any mail-dependent backup
|
|
alerting are unaccepted. Note separately that `postfix check` reports
|
|
divergence between `/var/spool/postfix` copies and their host originals,
|
|
including `/etc/hosts` and NSS libraries — a known cause of resolution failure
|
|
inside the chroot, and adjacent enough to this failure to be checked first.
|
|
|
|
**Resolution (2026-08-17):** Cause proven, three hops downstream of the origin.
|
|
`mx1` runs `permit_mynetworks, permit_auth_destination, reject` on both relay
|
|
and recipient restrictions. `wg-pk` connected from public addresses absent from
|
|
`mx1`'s `mynetworks`, and `kane-il.us` is not an authorised destination there,
|
|
so `RCPT TO <sandor@kane-il.us>` was rejected `554 5.7.1 Access denied` over
|
|
both IPv4 and IPv6. Added only `198.58.111.109/32` and
|
|
`[2600:3c00::f03c:92ff:fe42:43d7]/128` to `mx1`'s `mynetworks` and reloaded.
|
|
|
|
Four hypotheses were tested and ruled out with evidence before any change: the
|
|
`srv-b` alias (`postalias -q root` resolved correctly and the journal showed
|
|
`orig_to=<root>` forwarded), sender-domain rejection (`wg-pk` rewrites the
|
|
sender to `postmaster@diagnostics.kane-il.us` and `MAIL FROM` was accepted),
|
|
address-family asymmetry (both families failed identically), and routing on
|
|
`wg-pk` (it connected and received the rejection).
|
|
|
|
Method worth reusing: **RCPT-only probes before and after the change, over both
|
|
address families, with no message body.** `554` before, `250` after. That
|
|
proves the change caused the fix rather than coinciding with it — the standard
|
|
F-012 was written to enforce. Two messages then delivered end to end with full
|
|
headers captured, empty queue, no deferred entries. **Closed.**
|
|
|
|
---
|
|
|
|
### F-024 — work order specified a log file that does not exist
|
|
`srv-b`. Work order 002.
|
|
|
|
**Observed:** WORK ORDER 002 instructed `grep /var/log/mail.log`. The file does
|
|
not exist on `srv-b`.
|
|
**Cause:** **Proven.** Proxmox VE ships without `rsyslog`. Postfix logs to
|
|
journald only.
|
|
**Correction:** The operator used `journalctl -u postfix@-` and completed the
|
|
diagnosis. No configuration was changed; installing `rsyslog` to satisfy a
|
|
document would have been the wrong direction.
|
|
**Consequence:** **Specification defect.** All log inspection on a Proxmox host
|
|
uses `journalctl`, not files under `/var/log`. Corrected in `ENVIRONMENT.md`
|
|
revision 5. Automation that greps a logfile path will silently find nothing,
|
|
which is worse than failing.
|
|
|
|
---
|
|
|
|
### F-025 — the F-023 correction granted wider relay than required
|
|
`mx1`, `wg-pk`. Estate scope.
|
|
|
|
**Observed:** Adding `wg-pk`'s two public addresses to `mx1`'s `mynetworks`
|
|
authorises them under `permit_mynetworks`, which appears in **both**
|
|
`smtpd_relay_restrictions` and `smtpd_recipient_restrictions`. That grants relay
|
|
to **any** destination, not only to `kane-il.us`.
|
|
|
|
`wg-pk` in turn carries `mynetworks = 10.110.0.0/22` with
|
|
`permit_mynetworks permit_sasl_authenticated defer_unauth_destination`. The
|
|
resulting chain:
|
|
|
|
```
|
|
any peer on 10.110.0.0/22 (including CT 100 and CT 101, which arrive
|
|
as 10.110.0.12 through the srv-b masquerade)
|
|
-> wg-pk permit_mynetworks
|
|
-> mx1 permit_mynetworks
|
|
-> any destination on the internet, as kane-il.us infrastructure
|
|
```
|
|
|
|
Before the correction `mx1` rejected at the final hop, so the path was closed by
|
|
accident rather than by policy.
|
|
|
|
**Cause:** **Proven.** `permit_mynetworks` is destination-agnostic by design.
|
|
**Correction:** **Pending — operator decision.** Two independent parts:
|
|
|
|
- *Ours:* block SMTP egress from the container network at `srv-b`. The
|
|
containers have no reason to originate mail; if they ever should, that is a
|
|
deliberate decision rather than something inherited from a masquerade. Local,
|
|
precise, touches no estate policy.
|
|
- *Estate:* whether `wg-pk`'s `mynetworks` should be explicit `/32` entries for
|
|
the hosts that legitimately originate mail rather than the whole tunnel range.
|
|
|
|
**Constraint:** `wg-pk` must continue to relay to arbitrary external
|
|
destinations — that is its purpose, since Hubzilla registration and notification
|
|
mail depends on it and the ISP blocks port 25. Restricting it by *recipient*
|
|
would break that. The question is which clients may ask, not where it may send.
|
|
|
|
**Consequence:** A correction that fixes the observed failure may widen an
|
|
adjacent boundary. Record what a trust change grants, not only what it repairs.
|
|
|
|
**Project-local correction (2026-08-17, WO-003):** Complete. A persistent
|
|
`srv-b` FORWARD rule drops TCP 25, 465 and 587 from `10.20.0.0/24` to
|
|
`10.110.0.0/22`. Reachability was proven **before** the change, both containers
|
|
are blocked after it, ordinary internet egress is intact, host-originated mail
|
|
still delivers end to end, and the rule survived a host reboot.
|
|
|
|
**Estate half: still open.** Whether `wg-pk` should narrow `mynetworks` from
|
|
`10.110.0.0/22` to the explicit hosts that legitimately originate mail is a
|
|
CIVICVS decision. No change was made to `wg-pk`, `mx1` or `kane-il.us`.
|
|
|
|
F-025 is therefore **partially corrected**, not closed. The two halves are kept
|
|
in one entry rather than split, because the exposure is a single chain and
|
|
splitting it would let one half be closed while the other is forgotten.
|
|
|
|
---
|
|
|
|
### F-026 — containers were not configured to start with the host
|
|
`srv-b`, CT 100, CT 101. Work order 003.
|
|
|
|
**Observed:** After the WO-003 persistence reboot — the first true **host**
|
|
reboot since the containers were built — both were stopped:
|
|
|
|
```
|
|
pct status 100 -> status: stopped
|
|
pct status 101 -> status: stopped
|
|
```
|
|
|
|
**Cause:** **Proven.** Neither container configuration contained `onboot: 1`, so
|
|
Proxmox had no instruction to start either guest. `pve-guests.service` was
|
|
enabled, active, and its `startall` task completed successfully — proving the
|
|
startup machinery worked and simply had nothing to do.
|
|
**Correction:** `pct set 100 -onboot 1` and `pct set 101 -onboot 1`. The
|
|
containers were deliberately left stopped so the correction could be proven by
|
|
reboot rather than masked by a manual start. After a second host reboot both
|
|
returned automatically and host plus both guests reported `running`.
|
|
**Consequence:** **Specification defect.** `ENVIRONMENT.md` never required
|
|
`onboot: 1`, and every earlier reboot in this project was `pct reboot` — which
|
|
restarts a guest without ever exercising host-boot autostart. A container that
|
|
survives `pct reboot` is not thereby proven to survive a host reboot. Gate 1
|
|
must assert guest state after a **host** reboot specifically.
|
|
|
|
---
|
|
|
|
### F-027 — a verification test conflated tool failure with the tested condition
|
|
`srv-b`. Work order 003.
|
|
|
|
**Observed:** WO-003 B3 and B4 used this pattern to test SMTP reachability from
|
|
a container:
|
|
|
|
```bash
|
|
pct exec $c -- timeout 5 bash -c '...' \
|
|
&& echo "REACHABLE — WRONG" || echo "blocked — correct"
|
|
```
|
|
|
|
With the containers stopped (F-026) it produced:
|
|
|
|
```
|
|
-- CT 100: container '100' not running!
|
|
blocked — correct
|
|
```
|
|
|
|
**Cause:** **Proven.** The shell `||` branch catches *any* non-zero exit from
|
|
`pct exec`, including failures of `pct exec` itself. The test could not
|
|
distinguish "the firewall blocked the connection" from "the command never ran."
|
|
**Correction:** Assert guest state before interpreting any network result:
|
|
|
|
```bash
|
|
if [ "$(pct status "$c" | awk '{print $2}')" != "running" ]; then
|
|
echo "CT $c NOT RUNNING — TEST INVALID"
|
|
elif pct exec "$c" -- timeout 5 bash -c '...'; then
|
|
echo "REACHABLE — WRONG"
|
|
else
|
|
echo "blocked — correct"
|
|
fi
|
|
```
|
|
|
|
**Consequence:** **A test whose failure mode is indistinguishable from success
|
|
is worse than no test**, because it manufactures confidence. This one printed a
|
|
pass while the thing under test did not exist. Only a separate health assertion
|
|
caught it, and that was ordering luck rather than design. Every negative
|
|
assertion — "X is unreachable", "Y is absent", "Z is refused" — must first
|
|
establish that the test could have observed the positive case.
|
|
|
|
Applies retroactively: the isolation proofs in F-018 used the same shape. They
|
|
happened to be run against containers that were up, and the result was
|
|
independently corroborated, but the pattern was unsafe there too.
|
|
|
|
---
|
|
|
|
### F-028 — three command defects in work order 003
|
|
`srv-b`. Work order 003. Minor, grouped.
|
|
|
|
**Observed and corrected in place by the operator:**
|
|
|
|
| Defect | Cause | Correction |
|
|
|---|---|---|
|
|
| Serial numbers absent from A1 output | `grep -E "Serial Number"` is case-sensitive; SCSI output says `Serial number:` | `grep -iE` with a case-insensitive alternation on product, model, serial and SMART support |
|
|
| `journalctl -u smartd -b` returned `-- No entries --` | `smartd.service` is an alias; the real unit is `smartmontools.service` and owns the journal | Query `smartmontools` |
|
|
| A3 `sed` would have produced a malformed directive | Prefix substitution left the trailing `-M exec ...` in place and duplicated `-m root` | Replace the whole `DEFAULT` line, in both A3 and A4 |
|
|
|
|
**Consequence:** Match on the canonical unit name, not a convenience alias.
|
|
Prefer whole-line replacement over prefix patching in configuration edits. Use
|
|
case-insensitive matching when parsing tool output whose field names vary by
|
|
device class.
|
|
|
|
---
|
|
|
|
### F-029 — pip cache written into the repository working tree
|
|
CT 100. Repository seeding.
|
|
|
|
**Observed:** `git add -A` staged 465 files where 28 were expected. The extra
|
|
437 were `.cache/` — pip's download cache.
|
|
**Cause:** **Proven.** The service user's home is the install directory, so
|
|
`$HOME/.cache` **is** `/var/www/mechcomp/.cache`. `make deps` wrote there.
|
|
**Correction:** `.cache/`, `.local/`, `.ssh/`, `.gitconfig` and `.lesshst`
|
|
added to `.gitignore`; the cache unstaged before committing.
|
|
**Consequence:** Setting the service user's home to `install_dir` is correct
|
|
and matches YunoHost convention, but it means **any tool writing to `$HOME`
|
|
writes into the working tree**. `.gitignore` must anticipate that from the
|
|
first commit.
|
|
|
|
Method note: reading the staged list by eye showed nothing wrong. Counting it
|
|
and grouping by top-level directory found the problem immediately. **Verify by
|
|
counting, not by scanning** — a long correct-looking list is exactly where an
|
|
extra 437 files hide.
|
|
|
|
---
|
|
|
|
### F-030 — a network command in a work order could hang indefinitely
|
|
`srv-b`. Repository seeding.
|
|
|
|
**Observed:** An SSH authentication test was issued with neither
|
|
`ConnectTimeout` nor `BatchMode`. Gitea's SSH is on port 42022, so the attempt
|
|
on 22 hung and the operator had to interrupt it.
|
|
**Cause:** **Proven.** Missing timeout options, and an incorrect assumption
|
|
about the port.
|
|
**Correction:** Re-issued with `-o ConnectTimeout=10 -o BatchMode=yes -p 42022`,
|
|
after a `timeout`-wrapped reachability probe that could not hang.
|
|
**Consequence:** **Every network command in a work order carries an explicit
|
|
timeout.** A command that can hang strands the session, and the operator cannot
|
|
tell a hang from slow progress. Reachability is probed with `timeout` before any
|
|
client is invoked.
|
|
|
|
Also recorded, having been got wrong twice: Gitea **deploy tokens** are
|
|
account-level under user Settings, not repository settings. Repository-scoped
|
|
credentials are **deploy keys**. SSH here is on **42022**, so remotes need
|
|
`ssh://git@host:42022/owner/repo.git` — the `git@host:path` shorthand cannot
|
|
carry a port.
|
|
|
|
---
|
|
|
|
### F-031 — a standard was asserted from a check that never examined the thing
|
|
CT 100, CT 101, CT 102. Container standardisation.
|
|
|
|
**Observed:** The architect stated that CT 100 and CT 101 had no mail agent, and
|
|
standardised CT 102 by purging Postfix to match. The first run of
|
|
`ct-baseline.sh` reported a mail agent installed on **both** CT 100 and CT 101.
|
|
CT 102 — the container just "corrected" — was the only one conforming.
|
|
|
|
**Cause:** **Proven.** The earlier assertion came from inspecting Postfix
|
|
configuration and `/etc/aliases` **on `srv-b`**, during work order 002. That
|
|
established what the host does. It said nothing about the containers, and the
|
|
containers were never examined.
|
|
|
|
**Correction:** Postfix purged from CT 100 and CT 101. Queues were checked first
|
|
and both were empty, so nothing was lost. Re-run: 62 passed, 0 failed.
|
|
|
|
**Consequence:** **A standard asserted from a proxy observation is not a
|
|
standard.** The check must examine the thing being standardised. This is the
|
|
same class as F-027 — a result that cannot distinguish the case it claims to
|
|
test — but reached through inference rather than through shell semantics.
|
|
|
|
It is also the justification for `ct-baseline.sh` existing. Divergence had been
|
|
discovered by asking, one property at a time, whenever something behaved oddly.
|
|
That does not terminate: every check finds a new difference because nothing
|
|
states what "the same" means. The script is that statement, and it disagreed
|
|
with the architect on its first run.
|
|
|
|
---
|
|
|
|
### F-032 — the same host defects were discovered independently by two projects
|
|
`srv-b`. Cross-project.
|
|
|
|
**Observed:** Kane Fabric's `INFRASTRUCTURE_BASELINE.md`, written independently,
|
|
records `nesting=1,keyctl=1` against `226/NAMESPACE` systemd failures, and that
|
|
minimal Debian CTs lack a generated `en_US.UTF-8`. These are F-003 and F-004,
|
|
found again on the same host, in different words, at a different time.
|
|
**Cause:** **Proven.** Both are properties of the Proxmox/Debian 12 host, not of
|
|
either application. Neither project's documentation was visible to the other.
|
|
**Correction:** None required — both projects reached the correct conclusion.
|
|
**Consequence:** **Host-level requirements belong to the host, not to whichever
|
|
project discovered them first.** `ct-baseline.sh` encodes them once and checks
|
|
every container on `srv-b` regardless of owner, so the third project does not
|
|
have to discover them a third time.
|
|
|
|
Note that `keyctl=1` is required **only where Docker runs**. CT 101 was proven
|
|
to work with `nesting=1` alone (F-003). Kane Fabric's baseline currently
|
|
specifies both; if CT 102 runs no containers, `keyctl` can be dropped.
|
|
|
|
---
|
|
|
|
### F-033 — a restore the tool announced and never performed
|
|
CT 100. Repository tooling, Shapely port gate.
|
|
|
|
**Observed:** `bash tools/reference-toolchain/verify.sh --full` regenerated all 123
|
|
cases identically and diffed in exactly two adjacent lines, `fixtures_sha256` and
|
|
`frozen` — the expected result. The script printed the `DIFFERS` banner and the
|
|
diff body, then exited 1. Its closing line, `The committed oracle has been restored.
|
|
Nothing was overwritten.`, **did not appear in the journal**.
|
|
|
|
Afterwards the working tree carried a modified oracle:
|
|
|
|
```
|
|
M fixtures/strap-beam-8.0.0/strap-beam-fixtures-8.0.0.json
|
|
```
|
|
|
|
and `make_fixtures.py --verify` reported `recorded` and `computed` both
|
|
`577d6575...` — the regenerated hash — and declared `OK - oracle is intact`.
|
|
The committed value is `ddd0f154...`.
|
|
|
|
**Cause:** **Proven.** The script runs under `set -euo pipefail`. On the diff branch
|
|
it executes `diff ... | head -40 >&2`. `diff` exits 1 when files differ, `pipefail`
|
|
propagates that past `head`, and `set -e` aborts the script there. The `cp` that
|
|
restores the pre-run copy is the next statement and never runs. The reassurance is
|
|
printed after the `cp`, so it is unreachable on the only branch that prints it.
|
|
|
|
**Correction:** `|| true` appended to that pipeline. The working tree was restored
|
|
with `git checkout --` on the fixture file. Nothing was lost: the committed bytes
|
|
were in Gitea throughout.
|
|
|
|
**Consequence:** Two things.
|
|
|
|
**A tool's claim that it did something is not evidence that it did.** Verify the
|
|
effect, not the announcement. This is the F-027 class again — the failure mode was
|
|
indistinguishable from success, and it printed the success message.
|
|
|
|
**`make_fixtures.py --verify` cannot detect substitution of the whole file.** It
|
|
recomputes the hash from the document it reads and compares against the value stored
|
|
inside that same document, so a wholly regenerated oracle is self-consistent and
|
|
passes. It proves internal integrity, never identity with the committed oracle. Pair
|
|
it with `git status` whenever the question is *which* oracle is present.
|
|
|
|
---
|
|
|
|
### F-034 — arc segment counts land on exact integers and tip on the last bit
|
|
CT 100. Shapely port, `round_corners`.
|
|
|
|
**Observed:** the Y profile builds a section area of 135.572973 against a
|
|
recorded 135.574 — out by 0.001027, just past the 1e-3 area tolerance — while
|
|
every other recorded value for that case matches exactly: envelope, solved spoke
|
|
radius, minimum wall, channel count. Three-Fin matches on every value including
|
|
area.
|
|
|
|
Isolated by probing the reference directly inside the pinned toolchain image.
|
|
The hull cap is identical to nine figures. The bare union before filleting is
|
|
identical. A single filleted pair is identical, 30 vertices and 125.699057 mm^2.
|
|
The whole difference appears when the three filleted pairs are combined.
|
|
|
|
The three pairs are related by 120 degree symmetry and must be identical. They
|
|
come back as `[125.699057, 125.698029, 125.699057]`.
|
|
|
|
**Cause:** **Proven.** `_circlecorner` sets the arc segment count to
|
|
`max(3, ceil((90 - half_angle)/180 * $fn))`. At a 60 degree half-angle with
|
|
`$fn = 48` that expression is mathematically exactly 8. Floating point delivers
|
|
it as `8.000000000000004` on one corner and `7.999999999999998` on the other
|
|
two, so the ceiling gives 9 on one and 8 on the others — one extra segment,
|
|
slightly less enclosed area.
|
|
|
|
The half-angle is computed from the merged polygon's vertices, and those come
|
|
from the boolean kernel. BOSL2's clipper and GEOS agree to well within any
|
|
meaningful tolerance but not to the last bit, so they tip these ceilings
|
|
differently.
|
|
|
|
**Correction: none. Reproduced unguarded, because the reference is unguarded.**
|
|
|
|
Rounding the count before the ceiling was implemented and reverted. It makes all
|
|
three pairs identical and fixes Y exactly — and breaks Three-Fin, which had been
|
|
matching to the digit. Three-Fin has the same near-integer corners and the same
|
|
one-pair-tips-to-9 asymmetry, and the oracle **records that asymmetry**. BOSL2
|
|
happened to tip the same way GEOS does there, and the opposite way on Y.
|
|
|
|
BOSL2 guards this identical hazard inside `segs()` with a `2e-15` subtraction
|
|
but not at this call site. Adding the guard is better engineering and produces a
|
|
different answer from the reference, which is what matters here.
|
|
|
|
**Consequence:** Three things.
|
|
|
|
**Some recorded values encode float noise, not geometry.** Three-Fin's area
|
|
reflects a spurious extra segment on one of three symmetric corners. A port that
|
|
is geometrically more correct than the reference will fail that case.
|
|
|
|
**Bit-exact agreement across a different boolean kernel is not achievable in
|
|
general.** Not for want of care in the port. Wherever a corner angle lands on an
|
|
exact segment boundary, the result is decided by the kernel's last bit.
|
|
|
|
**The tolerance model may need revisiting, and that is CIVICVS's call.**
|
|
`VOLUME_MM3` ends in `_MM3`, so `test_oracle.py` compares it at the lengths
|
|
tolerance of 1e-4 rather than the areas tolerance of 1e-3. It is section area
|
|
times a 100 mm length, so an area difference of 1e-3 becomes a volume difference
|
|
of 1e-1 — a thousand times the tolerance it is checked against. `MASS_G` is
|
|
derived the same way. Any case affected by this failure mode fails on volume and
|
|
mass long before it fails on area.
|
|
|
|
Do not act on this until `build()` exists and all 123 cases can be run. The
|
|
number of affected cases is unknown and is the only thing that should drive the
|
|
decision.
|
|
|
|
---
|
|
|
|
## Open, not closed
|
|
|
|
| # | Status |
|
|
|---|---|
|
|
| F-006 | **Open** — cause unproven. Recurrence should capture `dpkg` lock state. |
|
|
| F-012 | **Open** — cause unproven. Leading candidate ruled out by inspection. `default_server` was added as independent hardening and does **not** close this. |
|
|
| F-019 | **Corrected** 2026-08-16. Reboot persistence proven. |
|
|
| F-021 | **Corrected** 2026-08-16. Reboot persistence proven. |
|
|
| F-022 | **Open** — cause unproven, no correction applied. |
|
|
| F-023 | **Closed** 2026-08-17. Cause proven at `mx1`; delivery proven twice. |
|
|
| F-024 | **Closed** 2026-08-17. Specification corrected. |
|
|
| F-025 | **Partially corrected** 2026-08-17. Project-local half closed and reboot-proven; **estate half open**, awaiting a CIVICVS decision on `wg-pk`. |
|
|
| F-026 | **Corrected** 2026-08-17. Reboot persistence proven. |
|
|
| F-027 | **Corrected** 2026-08-17. Applies retroactively to F-018's proofs. |
|
|
| F-028 | **Corrected** 2026-08-17. |
|
|
| F-029 | **Corrected** 2026-08-18. |
|
|
| F-030 | **Corrected** 2026-08-18. |
|
|
| F-031 | **Corrected** 2026-08-18. Root cause of `ct-baseline.sh`. |
|
|
| F-032 | **Closed** 2026-08-18. No correction required; encoded in `ct-baseline.sh`. |
|
|
| F-033 | **Corrected** 2026-08-19. Restore path unreachable under `set -e`. |
|
|
| F-034 | **Open** 2026-08-19. Reproduced unguarded; affects an unknown number of oracle cases. |
|
|
|
|
Everything else is closed with a proven cause and a proven correction.
|