Files
mechanical-compiler/docs/FAILURES.md
T
TheRON 6836ece4ff tests: close F-034; bound the derived trio at 8 ULP of the oracle record
The oracle records six significant figures, so the last digit of an area near 200 mm2 is worth 0.001 mm2. Every measured disagreement between port and reference is one, two or three units in that place. SECTION_AREA_MM2, VOLUME_MM3 and MASS_G are now bounded at 8 ULP of the expected value. Everything that positions material keeps the declared 1e-4 mm and remains exact in all 123 cases.

Option 1 scope kept, derivation rejected on measurement. Max |dA|/P over the accepted set is 1.434e-05 mm, one seven-hundredth of the 0.01 mm criterion, so a perimeter x 0.01 bound would have run 1800 to 4500 times the worst real discrepancy and caught nothing. Perimeter also anti-correlates with the error.

Suite 468 passed, 0 failed. The 30 expected failures are resolved, not suppressed. Mutation tested before landing: worst case uses 37.5 percent of its bound, 12 ULP offsets and 1e-4 relative scalings are caught in all 113 cases, 1e-6 and 1e-5 correctly are not.

Adds docs/ACCEPTANCE.md as the specification. Adds F-036, the venv interpreter error, same class as F-035. Corrects the F-034 per-profile distribution to Three-Fin 10, Y 7, A Frame 6, Rectangle 5, T 1, Four-Fin 1, which sums to the stated 30.
2026-08-23 08:19:18 -05:00

48 KiB

FAILURES.md

Append-only record of every failure encountered building the Mechanical Compiler environment.

Scope All instances. Staging entries are marked srv-b.
Updated 2026-08-22, closing F-034
Method PROCESS.md section 7
Rule Append only. Never edit an entry except to add a Resolution line.
Numbering Sequential, never reused. See §0 on the renumbering.

0. Why this file exists separately

This is the primary input to provisioning automation.

Every entry below is something a script written from ENVIRONMENT.md alone would have got wrong. F-003 and F-004 in particular would have produced a container that booted, reported success, and was quietly broken. The specification cannot anticipate them; only contact with a host reveals them.

When the production automation is written, it should be read before ENVIRONMENT.md, not after.

Renumbering note, 2026-08-15

The staging checkpoint and the reconciliation checkpoint were written independently and both allocated F-008 and F-009. The reconciliation entries have been renumbered to F-015 and F-016. Original phase-1 numbering is preserved because it is cited elsewhere.

Entry format

### F-nnn — one-line summary
Host / container. Phase.
**Observed:** what was seen, verbatim where possible.
**Cause:** proven cause, or explicitly "unproven".
**Correction:** the smallest change that fixed it.
**Consequence:** what the specification or automation must do differently.

A failure with no Consequence line is either not yet understood or not worth recording.


Phase 1 — initial container build

F-001 — arping absent on the Proxmox host

srv-b.

Observed: command -v arping → not installed. Cause: Not part of the base PVE install. Correction: apt install iputils-arping. Consequence: Automation must install iputils-arping before duplicate address detection, or the DAD check silently does not happen.


F-002 — invalid Proxmox inspection command

srv-b.

Observed: pvesm config local-lvm returned CLI usage text. Cause: Operator command error. Not a host fault. Correction: Read /etc/pve/storage.cfg directly. Consequence: Read storage configuration from the file, not from a subcommand that does not exist.


F-003 — CT 101 booted with degraded systemd

CT 101.

Observed:

status=226/NAMESPACE
Failed to set up mount namespacing: Permission denied
systemd-logind.service        failed
systemd-networkd.service      failed
systemd-timedated.service     failed
systemd-networkd.socket       failed

Cause: Debian 12 systemd requires mount namespacing that an unprivileged LXC denies without nesting. Correction: features: nesting=1. Deliberately tested alone rather than copying nesting=1,keyctl=1 from CT 100 — keyctl=1 proved unnecessary. Consequence: Every Debian 12 unprivileged container needs nesting=1, including ones running nothing but nginx. keyctl=1 is needed only where Docker runs. Promoted into the specification.

This is the clearest example of why the manual build was correct: the container started, appeared healthy at a glance, and was broken in four services.


F-004 — locale configured but never generated

CT 100, CT 101.

Observed: LANG=en_US.UTF-8 set, but locale -a listed only C, C.utf8, POSIX. locale emitted Cannot set LC_* warnings. Cause: Setting LANG does not generate a locale. The container template ships none. Correction: install locales, enable en_US.UTF-8 UTF-8, locale-gen, update-locale. Consequence: Specification must state the generation step, not just the value. Silent until something depends on collation or encoding.


F-005 — locale repair exposed a pending upgrade backlog

CT 100, CT 101.

Observed: Installing locales pulled glibc-related packages and revealed 52 further pending upgrades. Cause: Template was not current. Correction: Full upgrade of each container, one at a time. Both reached Debian 12.15, zero pending, no reboot required. Consequence: Automation must apt full-upgrade immediately after first boot, before anything else is installed.


F-006 — transient curl / git command-not-found

CT 100.

Observed: An egress test returned curl: command not found and git: command not found. Immediately afterward both packages were installed and current, PATH normal, both executables present and working. Cause: Unproven. Hypothesis only: the test ran during the F-005 upgrade, while dpkg had the binaries briefly unlinked mid-transaction. Adjacent in time, consistent with the symptom, not demonstrated. Correction: None required. Repeat tests with absolute paths passed. Consequence: Do not close. If it recurs, capture dpkg lock state at the moment of failure.


F-007 — recursive chown failed on ext4 lost+found

CT 100.

Observed: chown: cannot read directory '/var/lib/mechcomp/lost+found': Permission denied. That directory is nobody:nogroup, mode 0700. Cause: /var/lib/mechcomp is a real filesystem root, not a directory. lost+found is created by mkfs and is not ours to manage. Correction: Own explicit application paths only. Leave lost+found alone. Consequence: Never chown -R a mount point root. Enumerate the directories the application actually uses. Promoted into the specification.


F-008 — Git "dubious ownership"

CT 100.

Observed: fatal: detected dubious ownership in repository at '/var/www/mechcomp' when verification ran as root. Cause: Repository cloned as mechcomp, inspected as root. Correction: Re-ran verification as the owning user. No root safe.directory exception was added — the correct fix, since the exception would have masked every future instance of the same mistake. Consequence: All repository operations run as the service user.


F-009 — dependency manifests absent from the application repository

CT 100.

Observed: requirements-base.txt and requirements-cad.txt do not exist; the repository is LICENSE and README.md at commit e85c4f4e. Cause: Application implementation has not started. Not a provisioning failure. Correction: None. Do not invent pins. Consequence: Infrastructure and application acceptance are separately gated. Infrastructure may complete without the application existing.


Phase 2 — architect corrections

F-010 — placeholder PID file rewrite failed to substitute

CT 100.

Observed: An automated rewrite left the literal $! in the PID file. Cause: Quoting error in the rewriting command. Correction: Identified the real PID by inspection. Consequence: Superseded by F-019 — the placeholder should be a systemd unit and have no PID file at all.


F-011 — root could not overwrite a mechcomp-owned file in /tmp

CT 100.

Observed: Permission denied writing a file owned by mechcomp:mechcomp in /tmp, as root. Cause: fs.protected_regular = 2 with /tmp mode 1777. Expected kernel behaviour, not a fault. Correction: Rewrote the file as the owning user. Consequence: Do not use /tmp for state shared across users. Moot once services run under systemd with PrivateTmp=true.


F-012 — first request after nginx reload returned the Debian welcome page

CT 101.

Observed: nginx -t passed; the immediately following request returned the 615-byte Debian default page. All subsequent requests proxied correctly. Cause: Unproven. The obvious candidate — Debian's default site still enabled and acting as default_server — was tested and ruled out: sites-enabled contains only the project vhost. Remaining hypothesis is a graceful worker transition. Not demonstrated. Correction: None required. Consequence: Do not close. Independently, declare default_server explicitly on the project vhost so that adding a second server block cannot make catch-all behaviour depend on file ordering.


F-013 — TLS test validated against an IP literal

CT 101.

Observed: SSL: no alternative certificate subject name matches target host name '10.0.0.21'. Cause: Test used an address; the certificate SAN contains a DNS name. Correct behaviour by both curl and OpenSSL. Correction: Test by hostname. Consequence: Acceptance checks must use the service FQDN. An IP-literal TLS test is always wrong unless the certificate carries an IP SAN.


F-014 — CT 101 had no mechcomp group

CT 101.

Observed: Adding sandor to mechcomp failed; the group did not exist. Cause: Specification said the admin joins the mechcomp group in both containers, but the group is created as a side effect of creating the service user, which exists only in CT 100. Correction: Created the system group and added sandor. Consequence: Specification defect. Either create the group explicitly where it is required, or scope the group membership to CT 100. GID 996 now matches across both containers, which is harmless and mildly useful.


F-015 — CT 101 verification lacked curl

CT 101. (Renumbered from a duplicate F-008.)

Observed: curl: command not found. Cause: Not in the base template; CT 101's package list did not include it. Correction: Installed curl and ca-certificates. Consequence: The proxy container needs its own minimal toolset. Do not assume packages installed in CT 100 exist in CT 101.


F-016 — placeholder PID capture stored a literal $!

CT 100. (Renumbered from a duplicate F-009.)

Observed: PID file contained $! rather than a number. The listener itself was healthy. Cause: Quoting error. Correction: Identified PID 12102 by inspection. Consequence: Superseded by F-019.


Phase 3 — network isolation

F-017 — containers were provisioned on the home LAN

srv-b, CT 100, CT 101.

Observed: Both containers had net0 on vmbr0 at 10.0.0.20/24 and 10.0.0.21/24, gateway 10.0.0.1 — the ISP router's network. Anything on the home LAN could reach them. Cause: Specification defect in ENVIRONMENT.md revision 4. The document placed containers on both bridges, giving each a management address on the LAN, on the unstated assumption that a LAN workstation would browse staging. That assumption was never confirmed and was wrong. Correction: Both containers moved to vmbr1 only, gateway 10.20.0.1. net1 deleted. Host NAT extended for 10.20.0.0/24. Consequence: Containers are on the portless service bridge only. srv-b is router and bastion; access arrives via WireGuard. Promoted into the specification.


F-018 — removing the LAN interface did not remove LAN reachability

CT 100, CT 101.

Observed: After F-017, both containers still reached 10.0.0.1 successfully. Cause: Removing an interface removes an address, not a route. ip_forward=1 plus the new MASQUERADE -s 10.20.0.0/24 -o vmbr0 — added to give the containers internet access — also gave them the LAN, translated to the host's address. Correction: A RETURN in nat POSTROUTING ahead of both masquerades, and a DROP in FORWARD, both scoped -s 10.20.0.0/24 -d 10.0.0.0/24, plus an explicit ACCEPT for container-to-container traffic. Consequence: Interface removal is not isolation. Automation must assert the negative — that the LAN is unreachable — not merely that the interface is gone. Note that the containers still reach srv-b itself at 10.0.0.12, because a container addressing its gateway takes the INPUT path and never enters FORWARD. That is required, not a leak.


F-019 — placeholder backend did not survive reboot

CT 100.

Observed: After the F-017 reboot, nothing listened on 10.20.0.10:8770. Cause: The placeholder was a bare foreground process with a PID file. No supervision. Correction: Recreated as mechcomp-placeholder.service. Consequence: Anything a proof depends on must be supervised, or the proof expires silently at the next reboot and the following session diagnoses a proxy fault that does not exist. Applies to temporary scaffolding as much as to real services. Resolution (2026-08-16): Installed /usr/local/libexec/mechcomp-placeholder.py and /etc/systemd/system/mechcomp-placeholder.service, enabled at multi-user.target, running as mechcomp:mechcomp, reading /etc/mechcomp/mechcomp.env, binding 10.20.0.10:8770, with Restart=on-failure and the standard hardening set. CT 100 was rebooted: the unit restarted automatically, the listener returned on the service address only, nginx proxied successfully, X-Forwarded-Proto: https was re-observed at the backend, and the container settled to running with zero failed units.


F-020 — service FQDN mapping changed twice

srv-b, CT 100, CT 101.

Observed: mechanical-compiler.dev.infra was mapped to 10.20.0.10, then corrected to 10.0.0.21, then corrected again to 10.20.0.11. Cause: Three different states, each correct for its moment. 10.20.0.10 was wrong — it pointed at the application rather than the proxy. 10.0.0.21 was right while CT 101 had a LAN address and workstation access was assumed. 10.20.0.11 became right once F-017 removed the LAN interface and access moved to WireGuard through srv-b. Correction: 10.20.0.11 on all three hosts. Consequence: Not an error in either direction — a value that tracked a topology decision. Recorded because the reasoning matters more than the value: the service FQDN names the proxy, and the proxy has exactly one address.


F-021 — deleting the Proxmox interface left a stale guest eth1

CT 100, CT 101. Network isolation.

Observed: After the F-017 topology change and reboot, both containers reported systemctl is-system-running → degraded, with networking.service and ifupdown-wait-online.service failed. Both guests' /etc/network/interfaces still carried an auto eth1 / static iface eth1 stanza although neither had an eth1 link. The boot journal on both:

Cannot find device "eth1"
ifup: failed to bring up eth1

Cause: Proven. pct set --delete net1 removes the LXC interface but does not remove the stanza already written into the guest. ifup -a therefore exited non-zero at boot even though eth0 came up correctly. ifupdown-wait-online failed as a consequence, not independently. systemd-networkd was investigated and explicitly ruled out — its units were disabled on both containers. Correction: Preserved a pre-correction copy, removed only the stale eth1 stanza on each guest, restarted the affected units. Both containers then rebooted: the stanza did not return, both units succeeded, both settled to running with zero failed units. Consequence: Removing a container interface is not proof that the guest converged. After any topology mutation, automation must inspect the guest interface file and assert systemctl is-system-running = running with zero failed units after a reboot. Connectivity alone is insufficient — the surviving interface works fine while boot remains degraded, which is precisely how this went unnoticed through an entire verification pass.


F-022 — transient DNS resolution timeout in CT 101

CT 101. Formal acceptance.

Observed: During the first acceptance pass, two unrelated public HTTPS tests both failed at name resolution:

deb.debian.org:          curl: (28) Resolving timed out after 5000 ms
gitea.barternetwork.us:  curl: (28) Resolving timed out after 5000 ms

IP routing to 10.110.0.1, 10.0.0.12 and CT 100 remained working throughout. Cause: Unproven. Diagnostics showed resolver configuration identical to CT 100 and the host (nameserver 75.75.75.75, hosts: files dns), systemd-resolved absent, 75.75.75.75 reachable, and getent ahostsv4 resolving both names immediately afterward. Repeat HTTPS tests returned 200.

Worth noting as context, not as cause: since F-017, container DNS traverses the host's masquerade to an external resolver. That dependency is new. Correction: None. No configuration was changed. Consequence: Do not convert a one-shot resolver timeout into a configuration change without evidence. On recurrence, capture resolver state and DNS traffic at the moment of failure before touching anything. A candidate mitigation — a second nameserver line, so a single hiccup retries rather than fails — is recorded but deliberately not applied on one unexplained event.


F-023 — relay accepted the mail; final delivery failed

srv-b. Alerting.

Observed: The relay was discovered at 10.110.0.1:25 over wg0, banner wg-pk.diagnostics.kane-il.us, offering STARTTLS with a self-signed CN = wg-pk. It accepted unauthenticated SMTP from 10.110.0.12 through RCPT TO, before and after STARTTLS. Ports 465 and 587 were unavailable, and the public address 198.58.111.109 did not expose SMTP on this path.

With relayhost = [10.110.0.1]:25 and root: sandor@kane-il.us, srv-b recorded successful handoff:

relay=10.110.0.1[10.110.0.1]:25 dsn=2.0.0
status=sent (250 2.0.0 Ok: queued as 87E446243A)

The local queue emptied. The operator then received a delivery-failure message at sandor@kane-il.us. Cause: Unproven, downstream of the demonstrated handoff. The bounce notice itself arriving at sandor@kane-il.us establishes that the relay can deliver to that address, which narrows the problem to the failing message rather than the destination. Leading hypothesis, untested: the envelope sender is root@srv-b.dev.infra, and dev.infra does not resolve publicly, so a downstream MTA rejects on sender-domain verification. Candidate remedies are myorigin or smtp_generic_maps. Correction: None. Mail alerting was deferred by operator decision. Consequence: SMTP 250 from the relay and an empty local queue prove handoff, not delivery. Alerting acceptance requires demonstrated end-to-end receipt. Until then mail, smartd alerting, and any mail-dependent backup alerting are unaccepted. Note separately that postfix check reports divergence between /var/spool/postfix copies and their host originals, including /etc/hosts and NSS libraries — a known cause of resolution failure inside the chroot, and adjacent enough to this failure to be checked first.

Resolution (2026-08-17): Cause proven, three hops downstream of the origin. mx1 runs permit_mynetworks, permit_auth_destination, reject on both relay and recipient restrictions. wg-pk connected from public addresses absent from mx1's mynetworks, and kane-il.us is not an authorised destination there, so RCPT TO <sandor@kane-il.us> was rejected 554 5.7.1 Access denied over both IPv4 and IPv6. Added only 198.58.111.109/32 and [2600:3c00::f03c:92ff:fe42:43d7]/128 to mx1's mynetworks and reloaded.

Four hypotheses were tested and ruled out with evidence before any change: the srv-b alias (postalias -q root resolved correctly and the journal showed orig_to=<root> forwarded), sender-domain rejection (wg-pk rewrites the sender to postmaster@diagnostics.kane-il.us and MAIL FROM was accepted), address-family asymmetry (both families failed identically), and routing on wg-pk (it connected and received the rejection).

Method worth reusing: RCPT-only probes before and after the change, over both address families, with no message body. 554 before, 250 after. That proves the change caused the fix rather than coinciding with it — the standard F-012 was written to enforce. Two messages then delivered end to end with full headers captured, empty queue, no deferred entries. Closed.


F-024 — work order specified a log file that does not exist

srv-b. Work order 002.

Observed: WORK ORDER 002 instructed grep /var/log/mail.log. The file does not exist on srv-b. Cause: Proven. Proxmox VE ships without rsyslog. Postfix logs to journald only. Correction: The operator used journalctl -u postfix@- and completed the diagnosis. No configuration was changed; installing rsyslog to satisfy a document would have been the wrong direction. Consequence: Specification defect. All log inspection on a Proxmox host uses journalctl, not files under /var/log. Corrected in ENVIRONMENT.md revision 5. Automation that greps a logfile path will silently find nothing, which is worse than failing.


F-025 — the F-023 correction granted wider relay than required

mx1, wg-pk. Estate scope.

Observed: Adding wg-pk's two public addresses to mx1's mynetworks authorises them under permit_mynetworks, which appears in both smtpd_relay_restrictions and smtpd_recipient_restrictions. That grants relay to any destination, not only to kane-il.us.

wg-pk in turn carries mynetworks = 10.110.0.0/22 with permit_mynetworks permit_sasl_authenticated defer_unauth_destination. The resulting chain:

any peer on 10.110.0.0/22   (including CT 100 and CT 101, which arrive
                             as 10.110.0.12 through the srv-b masquerade)
  -> wg-pk    permit_mynetworks
  -> mx1      permit_mynetworks
  -> any destination on the internet, as kane-il.us infrastructure

Before the correction mx1 rejected at the final hop, so the path was closed by accident rather than by policy.

Cause: Proven. permit_mynetworks is destination-agnostic by design. Correction: Pending — operator decision. Two independent parts:

  • Ours: block SMTP egress from the container network at srv-b. The containers have no reason to originate mail; if they ever should, that is a deliberate decision rather than something inherited from a masquerade. Local, precise, touches no estate policy.
  • Estate: whether wg-pk's mynetworks should be explicit /32 entries for the hosts that legitimately originate mail rather than the whole tunnel range.

Constraint: wg-pk must continue to relay to arbitrary external destinations — that is its purpose, since Hubzilla registration and notification mail depends on it and the ISP blocks port 25. Restricting it by recipient would break that. The question is which clients may ask, not where it may send.

Consequence: A correction that fixes the observed failure may widen an adjacent boundary. Record what a trust change grants, not only what it repairs.

Project-local correction (2026-08-17, WO-003): Complete. A persistent srv-b FORWARD rule drops TCP 25, 465 and 587 from 10.20.0.0/24 to 10.110.0.0/22. Reachability was proven before the change, both containers are blocked after it, ordinary internet egress is intact, host-originated mail still delivers end to end, and the rule survived a host reboot.

Estate half: still open. Whether wg-pk should narrow mynetworks from 10.110.0.0/22 to the explicit hosts that legitimately originate mail is a CIVICVS decision. No change was made to wg-pk, mx1 or kane-il.us.

F-025 is therefore partially corrected, not closed. The two halves are kept in one entry rather than split, because the exposure is a single chain and splitting it would let one half be closed while the other is forgotten.


F-026 — containers were not configured to start with the host

srv-b, CT 100, CT 101. Work order 003.

Observed: After the WO-003 persistence reboot — the first true host reboot since the containers were built — both were stopped:

pct status 100  ->  status: stopped
pct status 101  ->  status: stopped

Cause: Proven. Neither container configuration contained onboot: 1, so Proxmox had no instruction to start either guest. pve-guests.service was enabled, active, and its startall task completed successfully — proving the startup machinery worked and simply had nothing to do. Correction: pct set 100 -onboot 1 and pct set 101 -onboot 1. The containers were deliberately left stopped so the correction could be proven by reboot rather than masked by a manual start. After a second host reboot both returned automatically and host plus both guests reported running. Consequence: Specification defect. ENVIRONMENT.md never required onboot: 1, and every earlier reboot in this project was pct reboot — which restarts a guest without ever exercising host-boot autostart. A container that survives pct reboot is not thereby proven to survive a host reboot. Gate 1 must assert guest state after a host reboot specifically.


F-027 — a verification test conflated tool failure with the tested condition

srv-b. Work order 003.

Observed: WO-003 B3 and B4 used this pattern to test SMTP reachability from a container:

pct exec $c -- timeout 5 bash -c '...' \
  && echo "REACHABLE — WRONG" || echo "blocked — correct"

With the containers stopped (F-026) it produced:

-- CT 100: container '100' not running!
blocked — correct

Cause: Proven. The shell || branch catches any non-zero exit from pct exec, including failures of pct exec itself. The test could not distinguish "the firewall blocked the connection" from "the command never ran." Correction: Assert guest state before interpreting any network result:

if [ "$(pct status "$c" | awk '{print $2}')" != "running" ]; then
    echo "CT $c NOT RUNNING — TEST INVALID"
elif pct exec "$c" -- timeout 5 bash -c '...'; then
    echo "REACHABLE — WRONG"
else
    echo "blocked — correct"
fi

Consequence: A test whose failure mode is indistinguishable from success is worse than no test, because it manufactures confidence. This one printed a pass while the thing under test did not exist. Only a separate health assertion caught it, and that was ordering luck rather than design. Every negative assertion — "X is unreachable", "Y is absent", "Z is refused" — must first establish that the test could have observed the positive case.

Applies retroactively: the isolation proofs in F-018 used the same shape. They happened to be run against containers that were up, and the result was independently corroborated, but the pattern was unsafe there too.


F-028 — three command defects in work order 003

srv-b. Work order 003. Minor, grouped.

Observed and corrected in place by the operator:

Defect Cause Correction
Serial numbers absent from A1 output grep -E "Serial Number" is case-sensitive; SCSI output says Serial number: grep -iE with a case-insensitive alternation on product, model, serial and SMART support
journalctl -u smartd -b returned -- No entries -- smartd.service is an alias; the real unit is smartmontools.service and owns the journal Query smartmontools
A3 sed would have produced a malformed directive Prefix substitution left the trailing -M exec ... in place and duplicated -m root Replace the whole DEFAULT line, in both A3 and A4

Consequence: Match on the canonical unit name, not a convenience alias. Prefer whole-line replacement over prefix patching in configuration edits. Use case-insensitive matching when parsing tool output whose field names vary by device class.


F-029 — pip cache written into the repository working tree

CT 100. Repository seeding.

Observed: git add -A staged 465 files where 28 were expected. The extra 437 were .cache/ — pip's download cache. Cause: Proven. The service user's home is the install directory, so $HOME/.cache is /var/www/mechcomp/.cache. make deps wrote there. Correction: .cache/, .local/, .ssh/, .gitconfig and .lesshst added to .gitignore; the cache unstaged before committing. Consequence: Setting the service user's home to install_dir is correct and matches YunoHost convention, but it means any tool writing to $HOME writes into the working tree. .gitignore must anticipate that from the first commit.

Method note: reading the staged list by eye showed nothing wrong. Counting it and grouping by top-level directory found the problem immediately. Verify by counting, not by scanning — a long correct-looking list is exactly where an extra 437 files hide.


F-030 — a network command in a work order could hang indefinitely

srv-b. Repository seeding.

Observed: An SSH authentication test was issued with neither ConnectTimeout nor BatchMode. Gitea's SSH is on port 42022, so the attempt on 22 hung and the operator had to interrupt it. Cause: Proven. Missing timeout options, and an incorrect assumption about the port. Correction: Re-issued with -o ConnectTimeout=10 -o BatchMode=yes -p 42022, after a timeout-wrapped reachability probe that could not hang. Consequence: Every network command in a work order carries an explicit timeout. A command that can hang strands the session, and the operator cannot tell a hang from slow progress. Reachability is probed with timeout before any client is invoked.

Also recorded, having been got wrong twice: Gitea deploy tokens are account-level under user Settings, not repository settings. Repository-scoped credentials are deploy keys. SSH here is on 42022, so remotes need ssh://git@host:42022/owner/repo.git — the git@host:path shorthand cannot carry a port.


F-031 — a standard was asserted from a check that never examined the thing

CT 100, CT 101, CT 102. Container standardisation.

Observed: The architect stated that CT 100 and CT 101 had no mail agent, and standardised CT 102 by purging Postfix to match. The first run of ct-baseline.sh reported a mail agent installed on both CT 100 and CT 101. CT 102 — the container just "corrected" — was the only one conforming.

Cause: Proven. The earlier assertion came from inspecting Postfix configuration and /etc/aliases on srv-b, during work order 002. That established what the host does. It said nothing about the containers, and the containers were never examined.

Correction: Postfix purged from CT 100 and CT 101. Queues were checked first and both were empty, so nothing was lost. Re-run: 62 passed, 0 failed.

Consequence: A standard asserted from a proxy observation is not a standard. The check must examine the thing being standardised. This is the same class as F-027 — a result that cannot distinguish the case it claims to test — but reached through inference rather than through shell semantics.

It is also the justification for ct-baseline.sh existing. Divergence had been discovered by asking, one property at a time, whenever something behaved oddly. That does not terminate: every check finds a new difference because nothing states what "the same" means. The script is that statement, and it disagreed with the architect on its first run.


F-032 — the same host defects were discovered independently by two projects

srv-b. Cross-project.

Observed: Kane Fabric's INFRASTRUCTURE_BASELINE.md, written independently, records nesting=1,keyctl=1 against 226/NAMESPACE systemd failures, and that minimal Debian CTs lack a generated en_US.UTF-8. These are F-003 and F-004, found again on the same host, in different words, at a different time. Cause: Proven. Both are properties of the Proxmox/Debian 12 host, not of either application. Neither project's documentation was visible to the other. Correction: None required — both projects reached the correct conclusion. Consequence: Host-level requirements belong to the host, not to whichever project discovered them first. ct-baseline.sh encodes them once and checks every container on srv-b regardless of owner, so the third project does not have to discover them a third time.

Note that keyctl=1 is required only where Docker runs. CT 101 was proven to work with nesting=1 alone (F-003). Kane Fabric's baseline currently specifies both; if CT 102 runs no containers, keyctl can be dropped.


F-033 — a restore the tool announced and never performed

CT 100. Repository tooling, Shapely port gate.

Observed: bash tools/reference-toolchain/verify.sh --full regenerated all 123 cases identically and diffed in exactly two adjacent lines, fixtures_sha256 and frozen — the expected result. The script printed the DIFFERS banner and the diff body, then exited 1. Its closing line, The committed oracle has been restored. Nothing was overwritten., did not appear in the journal.

Afterwards the working tree carried a modified oracle:

 M fixtures/strap-beam-8.0.0/strap-beam-fixtures-8.0.0.json

and make_fixtures.py --verify reported recorded and computed both 577d6575... — the regenerated hash — and declared OK - oracle is intact. The committed value is ddd0f154....

Cause: Proven. The script runs under set -euo pipefail. On the diff branch it executes diff ... | head -40 >&2. diff exits 1 when files differ, pipefail propagates that past head, and set -e aborts the script there. The cp that restores the pre-run copy is the next statement and never runs. The reassurance is printed after the cp, so it is unreachable on the only branch that prints it.

Correction: || true appended to that pipeline. The working tree was restored with git checkout -- on the fixture file. Nothing was lost: the committed bytes were in Gitea throughout.

Consequence: Two things.

A tool's claim that it did something is not evidence that it did. Verify the effect, not the announcement. This is the F-027 class again — the failure mode was indistinguishable from success, and it printed the success message.

make_fixtures.py --verify cannot detect substitution of the whole file. It recomputes the hash from the document it reads and compares against the value stored inside that same document, so a wholly regenerated oracle is self-consistent and passes. It proves internal integrity, never identity with the committed oracle. Pair it with git status whenever the question is which oracle is present.


F-034 — arc segment counts land on exact integers and tip on the last bit

CT 100. Shapely port, round_corners.

Measured 2026-08-20. Open, awaiting a tolerance decision from CIVICVS. The scale of the effect is now known and the port's side is settled: the port is exact and the reference is noisy. Nothing here can be corrected in the port.

Observed, first evidence (60 degree half-angles). The Y profile builds a section area of 135.572973 against a recorded 135.574 — out by 0.001027, just past the 1e-3 area tolerance — while every other recorded value for that case matches exactly: envelope, solved spoke radius, minimum wall, channel count. Three-Fin matches on every value including area.

Isolated by probing the reference directly inside the pinned toolchain image. The hull cap is identical to nine figures. The bare union before filleting is identical. A single filleted pair is identical, 30 vertices and 125.699057 mm^2. The whole difference appears when the three filleted pairs are combined.

The three pairs are related by 120 degree symmetry and must be identical. They come back as [125.699057, 125.698029, 125.699057].

Observed, second evidence (90 degree corners). The ring profiles hit the same mechanism at a different angle. For Rectangle at defaults, the reference builds an outer envelope of 50 vertices and the port builds 48 — 13+13+12+12 against 12x4. RING_SCALE, RING_EDGES_MM, RING_CORNER_WEB_MM and RING_CORNER_R_MAX_MM all agree to six figures, so the fitted polygon is identical and the divergence is entirely in the rounding.

Cause: Proven. _circlecorner sets the arc segment count to max(3, ceil((90 - half_angle)/180 * $fn)). That expression is frequently a mathematically exact integer:

corner half-angle expression at $fn = 48
spoke / fin junction 60 deg exactly 8
ring corner 45 deg exactly 12

A ceiling on an exact integer is a knife edge. Floating point delivers 8 as 8.000000000000004 on one corner and 7.999999999999998 on the others, so the ceiling gives 9 on one and 8 on the rest — one extra segment, slightly less enclosed area.

The half-angle is computed from the merged polygon's vertices, and those come from the boolean kernel. BOSL2's clipper and GEOS agree to well within any meaningful tolerance but not to the last bit, so they tip these ceilings differently.

Which side is noisy is now measured, and it is not the port. Instrumenting _circlecorner and printing at %.17g gives half=45 raw=12 ceil=12 on every call, for both Square and Rectangle — the port lands dead on the integer, deterministically. The reference's arithmetic lands a hair under 45 degrees at some corners, pushing the value fractionally above 12 and the ceiling to 13. Square passes because the reference happens to land on 12 there too.

Correction: none. Reproduced unguarded, because the reference is unguarded.

Rounding the count before the ceiling was implemented and reverted. It makes all three pairs identical and fixes Y exactly — and breaks Three-Fin, which had been matching to the digit. Three-Fin has the same near-integer corners and the same one-pair-tips-to-9 asymmetry, and the oracle records that asymmetry. BOSL2 happened to tip the same way GEOS does there, and the opposite way on Y.

BOSL2 guards this identical hazard inside segs() with a 2e-15 subtraction but not at this call site. Adding the guard is better engineering and produces a different answer from the reference, which is what matters here.

There is nothing left to try on the port's side. The port is already exact; a guard cannot make it match noise it does not have.

Scale of the effect. Measured against all 123 cases once build() existed:

  • 30 of 113 accepted cases affected. 83 are exact.
  • Only three keys ever breach: SECTION_AREA_MM2 (29 cases), VOLUME_MM3 (30), MASS_G (22). Volume is area times 100 mm and mass is volume times density over 1000, so each case carries one discrepancy reported three times.
  • Worst relative error 2.24e-05.
  • By profile: Three-Fin 9, Y 7, A Frame 5, Rectangle 5, T 1, Four-Fin 1. None on Equilateral Triangle, General Triangle, Square, Diamond or Cross.
  • Every dimensional quantity passes, in all 123 cases: ENVELOPE_X_MM, ENVELOPE_Y_MM, MIN_WALL_ACTUAL_MM, and every profile extra (AF_*, FIN_*, SPOKE_*, RING_*, T_*). Every count is exact. All ten rejections fire correctly.

Physical magnitude. At $fn = 48 a chord deviates from its true arc by r(1 - cos 3.75 deg) = r x 2.1413e-3:

corner radius deviation from the true arc
1.25 mm (4x ring corners) 0.0027 mm
2.00 mm (3x ring corners) 0.0043 mm

The port and the reference differ from each other by at most about 0.4 um, and both sit within 0.0043 mm of the exact arc. Against the project's 0.01 mm accuracy criterion that is two orders of margin. The geometry is not in question; only the comparison is.

Consequence: three things, the third now actionable.

Some recorded values encode float noise, not geometry. Three-Fin's area reflects a spurious extra segment on one of three symmetric corners. A port that is geometrically more correct than the reference will fail that case.

Bit-exact agreement across a different boolean kernel is not achievable in general. Not for want of care in the port. Wherever a corner angle lands on an exact segment boundary, the result is decided by the kernel's last bit.

The tolerance model needs revisiting, and that is CIVICVS's call. VOLUME_MM3 ends in _MM3, so test_oracle.py compares it at the lengths tolerance of 1e-4 rather than the areas tolerance of 1e-3 — against a magnitude near 20000, which demands 5e-9 relative agreement from discretised geometry. 3x/Y/steel0.79 breaches on VOLUME_MM3 alone while its area passes: one discrepancy, judged by two wildly different standards by accident of key naming.

Constraint on any fix: do not edit tolerance in the oracle JSON. It sits inside the hashed document — test_integrity_hash covers everything except fixtures_sha256 — so editing it breaks that test by design. The change belongs in test_oracle.py, which is not hashed. Options, in the order recommended:

  1. Scale-aware bounds for the three discretisation-limited keys, derived from the 0.01 mm criterion and the section perimeter. Most faithful to where the error originates.
  2. A relative floor — pass if within the absolute tolerance or within about 1e-4 relative. Simplest. Loosens SECTION_AREA_MM2 from 0.001 mm^2 to about 0.02 mm^2, a real but bounded cost.
  3. Accept 30 known failures and mark them expected. Honest, but make test is never green and a real regression would hide among them.

Whichever is chosen, the checks that catch structural error — SECTION_PARTS, STRAP_CHANNELS, MIN_WALL_ACTUAL_MM, the envelope dimensions — stay exact. A genuine geometry fault runs 1e-2 relative or worse and would still be caught with three orders of margin.

Resolution (2026-08-22): closed. Option 1's scope was kept and its derivation was rejected on measurement.

A read-only pass over all 113 accepted cases re-expressed every area discrepancy as |dA| / P — the uniform boundary displacement that would produce it, directly comparable to the 0.01 mm criterion. Median 0, p95 9.075e-06 mm, max 1.434e-05 mm. A perimeter x 0.01 bound would have run 1.86 mm^2 at the smallest section and 4.54 mm^2 at the largest — 1,800 to 4,500 times the worst real discrepancy — and would have caught nothing. Perimeter is also the wrong normaliser: it anti-correlates with the error, widening the spread from a factor of 3 to a factor of 6.5, because the error is driven by how many corners tip from 12 segments to 13 rather than by boundary length.

The real quantiser is the reference's own six-significant-figure echo. Every discrepancy in the set is 0.001, 0.002 or 0.003 mm^2 — one, two or three units in the last digit the reference ever recorded. The bound adopted is 8 * ulp(expected) where ulp(v) = 10 ** (floor(log10 |v|) - 5), applied to SECTION_AREA_MM2, VOLUME_MM3 and MASS_G only.

Propagation verified rather than assumed: VOLUME/AREA is exactly 100.0 in all 113 cases with maximum deviation 2.8e-14, and MASS/VOLUME is uniform to 1e-8, so one bound governs all three keys honestly.

Correction to the per-profile line above. It records Three-Fin 9, Y 7, A Frame 5, Rectangle 5, T 1, Four-Fin 1 — which sums to 28 against the stated total of 30. Measured: Three-Fin 10, Y 7, A Frame 6, Rectangle 5, T 1, Four-Fin 1 = 30. The entry above is left as written, per the append-only rule; this is the correct distribution.

Result: 468 passed, 0 failed. The 30 expected failures are resolved, not suppressed. Mutation-tested before landing — worst case consumes 37.5% of its bound, offsets of 12 or more last-place units are caught in all 113 cases on all three keys, and a 1e-4 relative scaling is caught everywhere, while 1e-6 and 1e-5 correctly are not. test_tolerance_policy_is_not_vacuous asserts the derived bound never exceeds 1e-4 relative, so a future widening fails a test instead of passing quietly. Specification in docs/ACCEPTANCE.md.


F-035 — su cannot run as a nologin service user

CT 100. Session handover.

Observed: a new assistant opened a session with

pct exec 100 -- su - mechcomp -c 'cd /var/www/mechcomp && git log --oneline -1'
pct exec 100 -- su - mechcomp -c 'cd /var/www/mechcomp && make test'

Both returned This account is currently not available and nothing else. With the same message for the repository check and the test run, the output reads as a broken clone or a broken container. Neither was true — the tree was clean and the suite passed.

Cause: Proven. mechcomp is a service account:

mechcomp:x:999:996::/var/www/mechcomp:/usr/sbin/nologin

su starts the account's login shell, which is nologin, whose entire function is to print that message and exit. The account is fine. runuser -u mechcomp -- executes the command directly without a login shell and works, which is what every command in the porting sessions used.

Correction: use runuser -u mechcomp -- <cmd>. Never su. Recorded as an explicit fact in HANDOFF.md §1 rather than left to be inferred from examples.

Consequence: Two things.

This is the F-027 class again — a tool failure reading as a condition failure. Both commands failed identically and before touching anything, so the message describes the invocation, not the state. When every command in a group fails the same way, suspect the invocation before concluding anything about the system.

Operational facts must be stated, not demonstrated. runuser appeared throughout the previous sessions only inside example commands, so it could be learned by pattern-matching but not by reading. That fails exactly when an assistant composes a command from scratch, which is what happened here. The same applies to bash tools/... over ./tools/..., to Gitea's port 42022, and to running every repository operation as mechcomp. All are now stated as facts in HANDOFF.md §1.


F-036 — a work order used the system interpreter instead of the virtualenv

CT 100. F-034 measurement pass.

Observed: a measurement script delivered by the architect was invoked as runuser -u mechcomp -- python3 /tmp/f034/f034-measure.py and died at import pytest with ModuleNotFoundError, three frames into loading tests/conftest.py.

Cause: Proven. Makefile line 3 sets PY ?= venv/bin/python, and make deps runs pip install -r requirements-base.txt -r requirements-cad.txt followed by pip install -e . into that virtualenv. pytest, Shapely, numpy and mechcomp itself all live in /var/www/mechcomp/venv. System python3 has none of them. Had the script got past the pytest import it would have failed again on Shapely.

Correction: invoke /var/www/mechcomp/venv/bin/python. The script itself needed no change; it ran first time on re-invocation.

Consequence: This is F-035 exactly, one layer up, and it is the reason F-035's own consequence was written. HANDOFF.md §1 stated runuser as a fact but left the interpreter to be inferred from the Makefile — learnable by pattern-matching an existing command, not by reading. That fails precisely when an assistant composes a command from scratch, which is what happened both times. Now stated as a fact in §1.

The general form is worth keeping: an operational fact that appears only inside example commands has not been documented. Grep the handoff for facts that exist only as examples; each is a future F-035.


Open, not closed

# Status
F-006 Open — cause unproven. Recurrence should capture dpkg lock state.
F-012 Open — cause unproven. Leading candidate ruled out by inspection. default_server was added as independent hardening and does not close this.
F-019 Corrected 2026-08-16. Reboot persistence proven.
F-021 Corrected 2026-08-16. Reboot persistence proven.
F-022 Open — cause unproven, no correction applied.
F-023 Closed 2026-08-17. Cause proven at mx1; delivery proven twice.
F-024 Closed 2026-08-17. Specification corrected.
F-025 Partially corrected 2026-08-17. Project-local half closed and reboot-proven; estate half open, awaiting a CIVICVS decision on wg-pk.
F-026 Corrected 2026-08-17. Reboot persistence proven.
F-027 Corrected 2026-08-17. Applies retroactively to F-018's proofs.
F-028 Corrected 2026-08-17.
F-029 Corrected 2026-08-18.
F-030 Corrected 2026-08-18.
F-031 Corrected 2026-08-18. Root cause of ct-baseline.sh.
F-032 Closed 2026-08-18. No correction required; encoded in ct-baseline.sh.
F-033 Corrected 2026-08-19. Restore path unreachable under set -e.
F-034 Closed 2026-08-22. Measured: the port is exact, the reference is noisy. 30 of 113 accepted cases, worst relative error 2.24e-05, confined to SECTION_AREA_MM2 and its two derivatives. Nothing correctable in the port; resolved in test_oracle.py by bounding at eight units in the last place of the oracle's six-significant-figure record. Suite green at 468 passed. Specification in docs/ACCEPTANCE.md.
F-035 Corrected 2026-08-19. Use runuser, never su; mechcomp is nologin.
F-036 Corrected 2026-08-22. The interpreter is venv/bin/python, never system python3. Same class as F-035.

Everything else is closed with a proven cause and a proven correction.