From bcec11cf43734d5d3e48fe06b13e29a9e7a0f5b3 Mon Sep 17 00:00:00 2001 From: TheRON Date: Mon, 17 Aug 2026 11:48:32 -0400 Subject: [PATCH] Final configuration --- docs/ENVIRONMENT.md | 76 ++++++++++++++++++++++++++-- docs/FAILURES.md | 112 +++++++++++++++++++++++++++++++++++++++++- docs/ROADMAP.md | 12 ++++- docs/STAGING-STATE.md | 67 ++++++++++++++++++------- 4 files changed, 242 insertions(+), 25 deletions(-) diff --git a/docs/ENVIRONMENT.md b/docs/ENVIRONMENT.md index 5e8eb2d..2bf126a 100644 --- a/docs/ENVIRONMENT.md +++ b/docs/ENVIRONMENT.md @@ -4,9 +4,9 @@ Specification for a Mechanical Compiler instance. | | | |---|---| -| Revision | 5.1 (2026-08-17) | +| Revision | 5.2 (2026-08-17) | | Supersedes | Revisions 1 through 4 | -| Basis | Revision 4, reconciled against the proven staging build, then work order 002 | +| Basis | Revision 4, reconciled against the proven staging build, then work orders 002 and 003 | | Scope | **Host-agnostic.** Applies to any instance. | | Instance state | `STAGING-STATE.md`, and later `PRODUCTION-STATE.md` | | Failure evidence | `FAILURES.md` | @@ -174,10 +174,18 @@ nat POSTROUTING -s -o -j MASQUERADE filter FORWARD + -s -d -p tcp + -m multiport --dports 25,465,587 -j DROP # containers do not send mail -s -d -j ACCEPT -s -d -j DROP ``` +**PROVEN (F-025)** — The SMTP rule is not optional where the host masquerades +containers onto a network carrying a mail relay. Without it the containers +inherit whatever relay authority the host's source address holds. The host's own +alerting is unaffected: it originates mail locally, so it takes OUTPUT and never +enters FORWARD. Prove that rather than assuming it. + **REQ** — Rule order is load-bearing. `RETURN` must precede both masquerades. Rules must be persisted, and the persisted set verified to match the live set. @@ -254,6 +262,12 @@ large data volume in a container snapshot is how a backup store fills. **REQ** — Do not over-commit a thin pool. Thin-pool exhaustion is an ugly failure mode. +**PROVEN (F-026)** — Every container that must survive a host restart carries +`onboot: 1`. This is not the default. It is also invisible to `pct reboot`, +which restarts a guest without exercising host-boot autostart at all — so a +container proven to survive `pct reboot` is **not** thereby proven to survive a +host reboot. + ### 4.3 Base OS **REQ** — Debian 12 bookworm, both containers. It matches YunoHost 12 stable and @@ -638,12 +652,34 @@ coinciding with it. Automation that greps a logfile path finds nothing silently, which is worse than failing. -**REQ** — Disk monitoring must be configured **explicitly per device** where the -controller does not present members to automatic scanning. An active +**PROVEN** — Disk monitoring must be configured **explicitly per device** where +the controller does not present members to automatic scanning. An active `smartmontools.service` monitoring zero devices is not monitoring, and is easy -to mistake for success. Acceptance is a delivered test alert, not an active +to mistake for success. Acceptance is a **delivered** test alert, not an active unit. +Behind an HP Smart Array controller the working pattern is: + +``` +DEFAULT -a -m root -M exec /usr/share/smartmontools/smartd-runner +/dev/sda -d cciss,0 +/dev/sda -d cciss,1 +/dev/sda -d cciss,2 +/dev/sda -d cciss,3 +``` + +`-M test` proves the alert path — one message per monitored device — and must be +**reverted afterwards**. Left in place it produces an alert storm on every +restart, which trains everyone to ignore disk alerts. That is worse than no +monitoring. + +**REQ (F-028)** — Query `smartmontools.service`, not `smartd.service`. The +latter is an alias and owns no journal; `journalctl -u smartd` returns nothing +on a host where monitoring is working perfectly. + +**REQ** — Record each member's **serial number** at configuration time. That is +the only thing that maps a future alert to a physical drive in a caddy. + ### 12.1 Downstream infrastructure is out of bounds **REQ** — An instance may configure its own relay client. It may not alter the @@ -705,6 +741,30 @@ committed; an example with empty values is. 11. AGPL-3.0 §13 requires network users be offered the source. The web tier carries a visible source link to the repository. A licence obligation. +### 14.1 Test validity + +**REQ (F-027)** — A test must be able to distinguish failure of the **tool** +from failure of the **thing under test**. + +```bash +# WRONG — any failure of the wrapper reads as a pass +run_in_guest "$c" 'connect to X' && echo "reachable — WRONG" || echo "blocked — correct" + +# RIGHT — establish that the test could have observed the positive case +if [ "$(guest_state "$c")" != "running" ]; then + echo "CT $c NOT RUNNING — TEST INVALID" +elif run_in_guest "$c" 'connect to X'; then + echo "reachable — WRONG" +else + echo "blocked — correct" +fi +``` + +This applies to **every negative assertion** — "X is unreachable", "Y is +absent", "Z is refused". A test whose failure mode is indistinguishable from +success is worse than no test, because it manufactures confidence. In WO-003 the +wrong form reported a pass against containers that were not running. + --- ## 15. Acceptance gates @@ -728,6 +788,12 @@ scaffolding survives a reboot. **PROVEN (F-021)** — the reboot is not optional. Connectivity can be perfect while boot is degraded. +**PROVEN (F-026)** — it must be a **host** reboot, and every guest must return +automatically and report `running`. A guest reboot does not exercise autostart. + +**PROVEN (F-027)** — every negative assertion in this gate must satisfy §14.1 +before its result means anything. + ### Gate 2 — Operational services Mail delivered **end to end and received twice**, with full headers captured, an diff --git a/docs/FAILURES.md b/docs/FAILURES.md index acb4d4b..63984d2 100644 --- a/docs/FAILURES.md +++ b/docs/FAILURES.md @@ -6,7 +6,7 @@ Compiler environment. | | | |---|---| | Scope | All instances. Staging entries are marked `srv-b`. | -| Updated | 2026-08-17, after work order 002 | +| Updated | 2026-08-17, after work order 003 | | Rule | Append only. Never edit an entry except to add a `Resolution` line. | | Numbering | Sequential, never reused. See §0 on the renumbering. | @@ -526,6 +526,111 @@ would break that. The question is which clients may ask, not where it may send. **Consequence:** A correction that fixes the observed failure may widen an adjacent boundary. Record what a trust change grants, not only what it repairs. +**Project-local correction (2026-08-17, WO-003):** Complete. A persistent +`srv-b` FORWARD rule drops TCP 25, 465 and 587 from `10.20.0.0/24` to +`10.110.0.0/22`. Reachability was proven **before** the change, both containers +are blocked after it, ordinary internet egress is intact, host-originated mail +still delivers end to end, and the rule survived a host reboot. + +**Estate half: still open.** Whether `wg-pk` should narrow `mynetworks` from +`10.110.0.0/22` to the explicit hosts that legitimately originate mail is a +CIVICVS decision. No change was made to `wg-pk`, `mx1` or `kane-il.us`. + +F-025 is therefore **partially corrected**, not closed. The two halves are kept +in one entry rather than split, because the exposure is a single chain and +splitting it would let one half be closed while the other is forgotten. + +--- + +### F-026 — containers were not configured to start with the host +`srv-b`, CT 100, CT 101. Work order 003. + +**Observed:** After the WO-003 persistence reboot — the first true **host** +reboot since the containers were built — both were stopped: + +``` +pct status 100 -> status: stopped +pct status 101 -> status: stopped +``` + +**Cause:** **Proven.** Neither container configuration contained `onboot: 1`, so +Proxmox had no instruction to start either guest. `pve-guests.service` was +enabled, active, and its `startall` task completed successfully — proving the +startup machinery worked and simply had nothing to do. +**Correction:** `pct set 100 -onboot 1` and `pct set 101 -onboot 1`. The +containers were deliberately left stopped so the correction could be proven by +reboot rather than masked by a manual start. After a second host reboot both +returned automatically and host plus both guests reported `running`. +**Consequence:** **Specification defect.** `ENVIRONMENT.md` never required +`onboot: 1`, and every earlier reboot in this project was `pct reboot` — which +restarts a guest without ever exercising host-boot autostart. A container that +survives `pct reboot` is not thereby proven to survive a host reboot. Gate 1 +must assert guest state after a **host** reboot specifically. + +--- + +### F-027 — a verification test conflated tool failure with the tested condition +`srv-b`. Work order 003. + +**Observed:** WO-003 B3 and B4 used this pattern to test SMTP reachability from +a container: + +```bash +pct exec $c -- timeout 5 bash -c '...' \ + && echo "REACHABLE — WRONG" || echo "blocked — correct" +``` + +With the containers stopped (F-026) it produced: + +``` +-- CT 100: container '100' not running! +blocked — correct +``` + +**Cause:** **Proven.** The shell `||` branch catches *any* non-zero exit from +`pct exec`, including failures of `pct exec` itself. The test could not +distinguish "the firewall blocked the connection" from "the command never ran." +**Correction:** Assert guest state before interpreting any network result: + +```bash +if [ "$(pct status "$c" | awk '{print $2}')" != "running" ]; then + echo "CT $c NOT RUNNING — TEST INVALID" +elif pct exec "$c" -- timeout 5 bash -c '...'; then + echo "REACHABLE — WRONG" +else + echo "blocked — correct" +fi +``` + +**Consequence:** **A test whose failure mode is indistinguishable from success +is worse than no test**, because it manufactures confidence. This one printed a +pass while the thing under test did not exist. Only a separate health assertion +caught it, and that was ordering luck rather than design. Every negative +assertion — "X is unreachable", "Y is absent", "Z is refused" — must first +establish that the test could have observed the positive case. + +Applies retroactively: the isolation proofs in F-018 used the same shape. They +happened to be run against containers that were up, and the result was +independently corroborated, but the pattern was unsafe there too. + +--- + +### F-028 — three command defects in work order 003 +`srv-b`. Work order 003. Minor, grouped. + +**Observed and corrected in place by the operator:** + +| Defect | Cause | Correction | +|---|---|---| +| Serial numbers absent from A1 output | `grep -E "Serial Number"` is case-sensitive; SCSI output says `Serial number:` | `grep -iE` with a case-insensitive alternation on product, model, serial and SMART support | +| `journalctl -u smartd -b` returned `-- No entries --` | `smartd.service` is an alias; the real unit is `smartmontools.service` and owns the journal | Query `smartmontools` | +| A3 `sed` would have produced a malformed directive | Prefix substitution left the trailing `-M exec ...` in place and duplicated `-m root` | Replace the whole `DEFAULT` line, in both A3 and A4 | + +**Consequence:** Match on the canonical unit name, not a convenience alias. +Prefer whole-line replacement over prefix patching in configuration edits. Use +case-insensitive matching when parsing tool output whose field names vary by +device class. + --- ## Open, not closed @@ -539,6 +644,9 @@ adjacent boundary. Record what a trust change grants, not only what it repairs. | F-022 | **Open** — cause unproven, no correction applied. | | F-023 | **Closed** 2026-08-17. Cause proven at `mx1`; delivery proven twice. | | F-024 | **Closed** 2026-08-17. Specification corrected. | -| F-025 | **Open** — correction pending operator decision. | +| F-025 | **Partially corrected** 2026-08-17. Project-local half closed and reboot-proven; **estate half open**, awaiting a CIVICVS decision on `wg-pk`. | +| F-026 | **Corrected** 2026-08-17. Reboot persistence proven. | +| F-027 | **Corrected** 2026-08-17. Applies retroactively to F-018's proofs. | +| F-028 | **Corrected** 2026-08-17. | Everything else is closed with a proven cause and a proven correction. diff --git a/docs/ROADMAP.md b/docs/ROADMAP.md index c41604a..e0eccd1 100644 --- a/docs/ROADMAP.md +++ b/docs/ROADMAP.md @@ -128,7 +128,11 @@ a new call into the same machinery, not a new machine. operator decision; application deployment remains blocked on application code. See `STAGING-STATE.md`. - **Mail alerting accepted, 2026-08-17.** Delivered end to end and received - twice. `smartd` is now unblocked. One open exposure is recorded as F-025. + twice. +- **Disk monitoring accepted, 2026-08-17.** Four P410i members explicitly + monitored, four test alerts received, serial numbers recorded, reboot-proven. +- **Container SMTP egress blocked, 2026-08-17.** The project-local half of F-025 + is corrected; the estate question about `wg-pk` client trust remains open. ### Next — the port @@ -254,3 +258,9 @@ on whether it makes distributed manufacturing capacity legible. 10. **Shared infrastructure is not ours to reconfigure.** An instance may configure its own client side; changes further down the chain are escalated with evidence. +11. **A test must distinguish tool failure from the condition it tests.** One + whose failure mode is indistinguishable from success manufactures + confidence (F-027). Every negative assertion first establishes that it could + have observed the positive case. +12. **A component surviving its own restart is not proven to survive the + host's.** Different machinery; assert against the one that matters (F-026). diff --git a/docs/STAGING-STATE.md b/docs/STAGING-STATE.md index 9fa1396..321cb0b 100644 --- a/docs/STAGING-STATE.md +++ b/docs/STAGING-STATE.md @@ -4,7 +4,7 @@ Live state of the Mechanical Compiler staging instance on `srv-b`. | | | |---|---| -| Updated | 2026-08-17, after work order 002 (mail) | +| Updated | 2026-08-17, after work order 003 (monitoring, SMTP egress) | | Instance | Staging / development | | Specification | `ENVIRONMENT.md` revision 5 | | Failure log | `FAILURES.md` | @@ -37,7 +37,7 @@ and must not be represented as either hidden failures or completed work: | Subsystem | Status | |---|---| | Mail alert delivery | **Accepted 2026-08-17.** Delivered end to end, twice, headers captured. | -| `smartd` monitoring | Deferred. Now unblocked — mail works. | +| `smartd` monitoring | **Accepted 2026-08-17.** Four members monitored, four alerts received. | | Backup infrastructure | **Postponed by operator decision** — strategy may change | --- @@ -77,6 +77,7 @@ template local:vztmpl/debian-12-standard_12.12-1_amd64.tar.zst | rootfs | 40 GiB `local-lvm` | 8 GiB `local-lvm` | | `mp0` | 60 GiB → `/var/lib/mechcomp`, `backup=0` | — | | `net0` | `eth0` on `vmbr1`, `10.20.0.10/24`, gw `10.20.0.1` | `eth0` on `vmbr1`, `10.20.0.11/24`, gw `10.20.0.1` | +| `onboot` | `1` | `1` | | Debian | 12 bookworm, fully upgraded | 12 bookworm, fully upgraded | | systemd | running, zero failed units | running, zero failed units | @@ -108,6 +109,8 @@ nat POSTROUTING -s 10.20.0.0/24 -o vmbr0 -j MASQUERADE # containers -> internet filter FORWARD + -s 10.20.0.0/24 -d 10.110.0.0/22 -p tcp + -m multiport --dports 25,465,587 -j DROP # SMTP egress blocked -s 10.20.0.0/24 -d 10.20.0.0/24 -j ACCEPT # container <-> container -s 10.20.0.0/24 -d 10.0.0.0/24 -j DROP # LAN blocked ``` @@ -115,6 +118,39 @@ filter FORWARD Rule order is load-bearing. `RETURN` must precede both service-network masquerades. Live and persisted states match. +### Disk monitoring — accepted 2026-08-17 + +`DEVICESCAN` finds nothing behind the P410i. Members are addressed explicitly: + +``` +/etc/smartd.conf +DEFAULT -a -m root -M exec /usr/share/smartmontools/smartd-runner +/dev/sda -d cciss,0 +/dev/sda -d cciss,1 +/dev/sda -d cciss,2 +/dev/sda -d cciss,3 + +before: Monitoring 0 ATA/SATA, 0 SCSI/SAS and 0 NVMe devices +after: Monitoring 0 ATA/SATA, 4 SCSI/SAS and 0 NVMe devices +unit: smartmontools.service (smartd.service is an alias with no journal) +alerts: 4 test alerts generated and received; headers captured + -M test reverted; permanent config restored and reboot-proven +``` + +**Physical identity of the four members.** This is how a future SMART alert is +matched to a caddy: + +| Member | Model | Serial | +|---|---|---| +| `cciss,0` | `EG0146FAWHU` | `6SD0LETK0000B05009W1` | +| `cciss,1` | `EG0146FAWHU` | `3SD3K5P00000905094UN` | +| `cciss,2` | `EG0146FAWHU` | `3SD3GHJ600009047RRRG` | +| `cciss,3` | `EG0146FAWHU` | `3SD3GVP800009045X79B` | + +**Not yet captured:** power-on hours and grown-defect counts. On 2010-vintage +drives holding the only instance, those numbers say how much runway remains. +`smartctl -d cciss,N -A /dev/sda` for each member. + ### TLS ``` @@ -232,6 +268,9 @@ Each line was demonstrated by command output, not inferred. - [x] NAT and FORWARD rules present live **and** persisted, in correct order - [x] Host: `running`, zero failed units - [x] **Mail delivered end to end from `root` on `srv-b`, twice, headers captured** +- [x] **`smartd` monitoring 4 devices; 4 test alerts delivered and received** +- [x] **Container SMTP egress blocked; internet egress and host mail intact** +- [x] **Both containers autostart after a host reboot** (`onboot: 1`) ### CT 100 @@ -278,19 +317,12 @@ Nothing. The infrastructure boundary is accepted. Recorded here so they are visually distinct from failures and from forgotten work. None is a defect. -- [ ] **SMTP egress block for the container network** (F-025). The containers - have no reason to originate mail. A `FORWARD` DROP on ports 25/465/587 - from the service network to `10.110.0.0/22` is local, precise, and touches - no estate policy. `srv-b`'s own alerts are unaffected — they originate - locally and take OUTPUT, never FORWARD. - [ ] **`wg-pk` `mynetworks` scope** (F-025). Estate decision: whether every tunnel peer should originate mail as `kane-il.us` infrastructure, or only the hosts that legitimately do. Not this project's to change unilaterally. -- [ ] **`smartd` explicit four-member configuration.** The package is installed - and the service is active, but `DEVICESCAN` currently monitors **zero - devices**; the P410i members are visible only through explicit - `-d cciss,N`. Active is not the same as monitoring. **Now unblocked** — - mail is accepted, so `-M test` gives a real end-to-end acceptance. +- [ ] **SMART attribute baseline.** Power-on hours and grown-defect counts for + the four members, so drive age is a known quantity rather than an + assumption. `smartctl -d cciss,N -A /dev/sda`. Read-only, one command. - [ ] **Backup infrastructure, entirely.** Postponed 2026-08-16 because the strategy may change: `vzdump` job, archive sizing, retention, free-space guard, host-side pull, `mechcomp-backup`, backup alerting, gold media, @@ -365,12 +397,11 @@ The Shapely port gates all of the above. | # | Question | Blocks | |---|---|---| -| 1 | Should container SMTP egress be blocked at `srv-b`? | F-025, ours to fix | -| 2 | Should `wg-pk` `mynetworks` narrow to explicit hosts? | F-025, estate decision | -| 3 | What is the backup strategy? | all backup work | -| 4 | Where is the 3+ TB USB disk attached? | gold redundancy step | +| 1 | Should `wg-pk` `mynetworks` narrow to explicit hosts? | F-025 estate half | +| 2 | What is the backup strategy? | all backup work | +| 3 | Where is the 3+ TB USB disk attached? | gold redundancy step | -All four need CIVICVS. +All three need CIVICVS. --- @@ -388,3 +419,5 @@ All four need CIVICVS. | `MECHCOMP_BIND` semantics | Scalar, service address only. Loopback not bound. | | Why did delivery fail downstream? | `mx1` relay trust. Proven and corrected (F-023). | | Where does Postfix log on this host? | journald. No `rsyslog`, no `/var/log/mail.log` (F-024). | +| Should container SMTP egress be blocked? | Yes. Implemented and reboot-proven (F-025 project-local half). | +| Do the containers autostart? | Yes, `onboot: 1` on both, proven by host reboot (F-026). |