Final configuration

This commit is contained in:
2026-08-17 11:48:32 -04:00
parent 673d5f26ac
commit bcec11cf43
4 changed files with 242 additions and 25 deletions
+71 -5
View File
@@ -4,9 +4,9 @@ Specification for a Mechanical Compiler instance.
| | |
|---|---|
| Revision | 5.1 (2026-08-17) |
| Revision | 5.2 (2026-08-17) |
| Supersedes | Revisions 1 through 4 |
| Basis | Revision 4, reconciled against the proven staging build, then work order 002 |
| Basis | Revision 4, reconciled against the proven staging build, then work orders 002 and 003 |
| Scope | **Host-agnostic.** Applies to any instance. |
| Instance state | `STAGING-STATE.md`, and later `PRODUCTION-STATE.md` |
| Failure evidence | `FAILURES.md` |
@@ -174,10 +174,18 @@ nat POSTROUTING
-s <SVC_NET> -o <MGMT> -j MASQUERADE
filter FORWARD
-s <SVC_NET> -d <WG_NET> -p tcp
-m multiport --dports 25,465,587 -j DROP # containers do not send mail
-s <SVC_NET> -d <SVC_NET> -j ACCEPT
-s <SVC_NET> -d <LAN_NET> -j DROP
```
**PROVEN (F-025)** — The SMTP rule is not optional where the host masquerades
containers onto a network carrying a mail relay. Without it the containers
inherit whatever relay authority the host's source address holds. The host's own
alerting is unaffected: it originates mail locally, so it takes OUTPUT and never
enters FORWARD. Prove that rather than assuming it.
**REQ** — Rule order is load-bearing. `RETURN` must precede both masquerades.
Rules must be persisted, and the persisted set verified to match the live set.
@@ -254,6 +262,12 @@ large data volume in a container snapshot is how a backup store fills.
**REQ** — Do not over-commit a thin pool. Thin-pool exhaustion is an ugly
failure mode.
**PROVEN (F-026)** — Every container that must survive a host restart carries
`onboot: 1`. This is not the default. It is also invisible to `pct reboot`,
which restarts a guest without exercising host-boot autostart at all — so a
container proven to survive `pct reboot` is **not** thereby proven to survive a
host reboot.
### 4.3 Base OS
**REQ** — Debian 12 bookworm, both containers. It matches YunoHost 12 stable and
@@ -638,12 +652,34 @@ coinciding with it.
Automation that greps a logfile path finds nothing silently, which is worse than
failing.
**REQ** — Disk monitoring must be configured **explicitly per device** where the
controller does not present members to automatic scanning. An active
**PROVEN** — Disk monitoring must be configured **explicitly per device** where
the controller does not present members to automatic scanning. An active
`smartmontools.service` monitoring zero devices is not monitoring, and is easy
to mistake for success. Acceptance is a delivered test alert, not an active
to mistake for success. Acceptance is a **delivered** test alert, not an active
unit.
Behind an HP Smart Array controller the working pattern is:
```
DEFAULT -a -m root -M exec /usr/share/smartmontools/smartd-runner
/dev/sda -d cciss,0
/dev/sda -d cciss,1
/dev/sda -d cciss,2
/dev/sda -d cciss,3
```
`-M test` proves the alert path — one message per monitored device — and must be
**reverted afterwards**. Left in place it produces an alert storm on every
restart, which trains everyone to ignore disk alerts. That is worse than no
monitoring.
**REQ (F-028)** — Query `smartmontools.service`, not `smartd.service`. The
latter is an alias and owns no journal; `journalctl -u smartd` returns nothing
on a host where monitoring is working perfectly.
**REQ** — Record each member's **serial number** at configuration time. That is
the only thing that maps a future alert to a physical drive in a caddy.
### 12.1 Downstream infrastructure is out of bounds
**REQ** — An instance may configure its own relay client. It may not alter the
@@ -705,6 +741,30 @@ committed; an example with empty values is.
11. AGPL-3.0 §13 requires network users be offered the source. The web tier
carries a visible source link to the repository. A licence obligation.
### 14.1 Test validity
**REQ (F-027)** — A test must be able to distinguish failure of the **tool**
from failure of the **thing under test**.
```bash
# WRONG — any failure of the wrapper reads as a pass
run_in_guest "$c" 'connect to X' && echo "reachable — WRONG" || echo "blocked — correct"
# RIGHT — establish that the test could have observed the positive case
if [ "$(guest_state "$c")" != "running" ]; then
echo "CT $c NOT RUNNING — TEST INVALID"
elif run_in_guest "$c" 'connect to X'; then
echo "reachable — WRONG"
else
echo "blocked — correct"
fi
```
This applies to **every negative assertion** — "X is unreachable", "Y is
absent", "Z is refused". A test whose failure mode is indistinguishable from
success is worse than no test, because it manufactures confidence. In WO-003 the
wrong form reported a pass against containers that were not running.
---
## 15. Acceptance gates
@@ -728,6 +788,12 @@ scaffolding survives a reboot.
**PROVEN (F-021)** — the reboot is not optional. Connectivity can be perfect
while boot is degraded.
**PROVEN (F-026)** — it must be a **host** reboot, and every guest must return
automatically and report `running`. A guest reboot does not exercise autostart.
**PROVEN (F-027)** — every negative assertion in this gate must satisfy §14.1
before its result means anything.
### Gate 2 — Operational services
Mail delivered **end to end and received twice**, with full headers captured, an
+110 -2
View File
@@ -6,7 +6,7 @@ Compiler environment.
| | |
|---|---|
| Scope | All instances. Staging entries are marked `srv-b`. |
| Updated | 2026-08-17, after work order 002 |
| Updated | 2026-08-17, after work order 003 |
| Rule | Append only. Never edit an entry except to add a `Resolution` line. |
| Numbering | Sequential, never reused. See §0 on the renumbering. |
@@ -526,6 +526,111 @@ would break that. The question is which clients may ask, not where it may send.
**Consequence:** A correction that fixes the observed failure may widen an
adjacent boundary. Record what a trust change grants, not only what it repairs.
**Project-local correction (2026-08-17, WO-003):** Complete. A persistent
`srv-b` FORWARD rule drops TCP 25, 465 and 587 from `10.20.0.0/24` to
`10.110.0.0/22`. Reachability was proven **before** the change, both containers
are blocked after it, ordinary internet egress is intact, host-originated mail
still delivers end to end, and the rule survived a host reboot.
**Estate half: still open.** Whether `wg-pk` should narrow `mynetworks` from
`10.110.0.0/22` to the explicit hosts that legitimately originate mail is a
CIVICVS decision. No change was made to `wg-pk`, `mx1` or `kane-il.us`.
F-025 is therefore **partially corrected**, not closed. The two halves are kept
in one entry rather than split, because the exposure is a single chain and
splitting it would let one half be closed while the other is forgotten.
---
### F-026 — containers were not configured to start with the host
`srv-b`, CT 100, CT 101. Work order 003.
**Observed:** After the WO-003 persistence reboot — the first true **host**
reboot since the containers were built — both were stopped:
```
pct status 100 -> status: stopped
pct status 101 -> status: stopped
```
**Cause:** **Proven.** Neither container configuration contained `onboot: 1`, so
Proxmox had no instruction to start either guest. `pve-guests.service` was
enabled, active, and its `startall` task completed successfully — proving the
startup machinery worked and simply had nothing to do.
**Correction:** `pct set 100 -onboot 1` and `pct set 101 -onboot 1`. The
containers were deliberately left stopped so the correction could be proven by
reboot rather than masked by a manual start. After a second host reboot both
returned automatically and host plus both guests reported `running`.
**Consequence:** **Specification defect.** `ENVIRONMENT.md` never required
`onboot: 1`, and every earlier reboot in this project was `pct reboot` — which
restarts a guest without ever exercising host-boot autostart. A container that
survives `pct reboot` is not thereby proven to survive a host reboot. Gate 1
must assert guest state after a **host** reboot specifically.
---
### F-027 — a verification test conflated tool failure with the tested condition
`srv-b`. Work order 003.
**Observed:** WO-003 B3 and B4 used this pattern to test SMTP reachability from
a container:
```bash
pct exec $c -- timeout 5 bash -c '...' \
&& echo "REACHABLE — WRONG" || echo "blocked — correct"
```
With the containers stopped (F-026) it produced:
```
-- CT 100: container '100' not running!
blocked — correct
```
**Cause:** **Proven.** The shell `||` branch catches *any* non-zero exit from
`pct exec`, including failures of `pct exec` itself. The test could not
distinguish "the firewall blocked the connection" from "the command never ran."
**Correction:** Assert guest state before interpreting any network result:
```bash
if [ "$(pct status "$c" | awk '{print $2}')" != "running" ]; then
echo "CT $c NOT RUNNING — TEST INVALID"
elif pct exec "$c" -- timeout 5 bash -c '...'; then
echo "REACHABLE — WRONG"
else
echo "blocked — correct"
fi
```
**Consequence:** **A test whose failure mode is indistinguishable from success
is worse than no test**, because it manufactures confidence. This one printed a
pass while the thing under test did not exist. Only a separate health assertion
caught it, and that was ordering luck rather than design. Every negative
assertion — "X is unreachable", "Y is absent", "Z is refused" — must first
establish that the test could have observed the positive case.
Applies retroactively: the isolation proofs in F-018 used the same shape. They
happened to be run against containers that were up, and the result was
independently corroborated, but the pattern was unsafe there too.
---
### F-028 — three command defects in work order 003
`srv-b`. Work order 003. Minor, grouped.
**Observed and corrected in place by the operator:**
| Defect | Cause | Correction |
|---|---|---|
| Serial numbers absent from A1 output | `grep -E "Serial Number"` is case-sensitive; SCSI output says `Serial number:` | `grep -iE` with a case-insensitive alternation on product, model, serial and SMART support |
| `journalctl -u smartd -b` returned `-- No entries --` | `smartd.service` is an alias; the real unit is `smartmontools.service` and owns the journal | Query `smartmontools` |
| A3 `sed` would have produced a malformed directive | Prefix substitution left the trailing `-M exec ...` in place and duplicated `-m root` | Replace the whole `DEFAULT` line, in both A3 and A4 |
**Consequence:** Match on the canonical unit name, not a convenience alias.
Prefer whole-line replacement over prefix patching in configuration edits. Use
case-insensitive matching when parsing tool output whose field names vary by
device class.
---
## Open, not closed
@@ -539,6 +644,9 @@ adjacent boundary. Record what a trust change grants, not only what it repairs.
| F-022 | **Open** — cause unproven, no correction applied. |
| F-023 | **Closed** 2026-08-17. Cause proven at `mx1`; delivery proven twice. |
| F-024 | **Closed** 2026-08-17. Specification corrected. |
| F-025 | **Open** — correction pending operator decision. |
| F-025 | **Partially corrected** 2026-08-17. Project-local half closed and reboot-proven; **estate half open**, awaiting a CIVICVS decision on `wg-pk`. |
| F-026 | **Corrected** 2026-08-17. Reboot persistence proven. |
| F-027 | **Corrected** 2026-08-17. Applies retroactively to F-018's proofs. |
| F-028 | **Corrected** 2026-08-17. |
Everything else is closed with a proven cause and a proven correction.
+11 -1
View File
@@ -128,7 +128,11 @@ a new call into the same machinery, not a new machine.
operator decision; application deployment remains blocked on application
code. See `STAGING-STATE.md`.
- **Mail alerting accepted, 2026-08-17.** Delivered end to end and received
twice. `smartd` is now unblocked. One open exposure is recorded as F-025.
twice.
- **Disk monitoring accepted, 2026-08-17.** Four P410i members explicitly
monitored, four test alerts received, serial numbers recorded, reboot-proven.
- **Container SMTP egress blocked, 2026-08-17.** The project-local half of F-025
is corrected; the estate question about `wg-pk` client trust remains open.
### Next — the port
@@ -254,3 +258,9 @@ on whether it makes distributed manufacturing capacity legible.
10. **Shared infrastructure is not ours to reconfigure.** An instance may
configure its own client side; changes further down the chain are escalated
with evidence.
11. **A test must distinguish tool failure from the condition it tests.** One
whose failure mode is indistinguishable from success manufactures
confidence (F-027). Every negative assertion first establishes that it could
have observed the positive case.
12. **A component surviving its own restart is not proven to survive the
host's.** Different machinery; assert against the one that matters (F-026).
+50 -17
View File
@@ -4,7 +4,7 @@ Live state of the Mechanical Compiler staging instance on `srv-b`.
| | |
|---|---|
| Updated | 2026-08-17, after work order 002 (mail) |
| Updated | 2026-08-17, after work order 003 (monitoring, SMTP egress) |
| Instance | Staging / development |
| Specification | `ENVIRONMENT.md` revision 5 |
| Failure log | `FAILURES.md` |
@@ -37,7 +37,7 @@ and must not be represented as either hidden failures or completed work:
| Subsystem | Status |
|---|---|
| Mail alert delivery | **Accepted 2026-08-17.** Delivered end to end, twice, headers captured. |
| `smartd` monitoring | Deferred. Now unblocked — mail works. |
| `smartd` monitoring | **Accepted 2026-08-17.** Four members monitored, four alerts received. |
| Backup infrastructure | **Postponed by operator decision** — strategy may change |
---
@@ -77,6 +77,7 @@ template local:vztmpl/debian-12-standard_12.12-1_amd64.tar.zst
| rootfs | 40 GiB `local-lvm` | 8 GiB `local-lvm` |
| `mp0` | 60 GiB → `/var/lib/mechcomp`, `backup=0` | — |
| `net0` | `eth0` on `vmbr1`, `10.20.0.10/24`, gw `10.20.0.1` | `eth0` on `vmbr1`, `10.20.0.11/24`, gw `10.20.0.1` |
| `onboot` | `1` | `1` |
| Debian | 12 bookworm, fully upgraded | 12 bookworm, fully upgraded |
| systemd | running, zero failed units | running, zero failed units |
@@ -108,6 +109,8 @@ nat POSTROUTING
-s 10.20.0.0/24 -o vmbr0 -j MASQUERADE # containers -> internet
filter FORWARD
-s 10.20.0.0/24 -d 10.110.0.0/22 -p tcp
-m multiport --dports 25,465,587 -j DROP # SMTP egress blocked
-s 10.20.0.0/24 -d 10.20.0.0/24 -j ACCEPT # container <-> container
-s 10.20.0.0/24 -d 10.0.0.0/24 -j DROP # LAN blocked
```
@@ -115,6 +118,39 @@ filter FORWARD
Rule order is load-bearing. `RETURN` must precede both service-network
masquerades. Live and persisted states match.
### Disk monitoring — accepted 2026-08-17
`DEVICESCAN` finds nothing behind the P410i. Members are addressed explicitly:
```
/etc/smartd.conf
DEFAULT -a -m root -M exec /usr/share/smartmontools/smartd-runner
/dev/sda -d cciss,0
/dev/sda -d cciss,1
/dev/sda -d cciss,2
/dev/sda -d cciss,3
before: Monitoring 0 ATA/SATA, 0 SCSI/SAS and 0 NVMe devices
after: Monitoring 0 ATA/SATA, 4 SCSI/SAS and 0 NVMe devices
unit: smartmontools.service (smartd.service is an alias with no journal)
alerts: 4 test alerts generated and received; headers captured
-M test reverted; permanent config restored and reboot-proven
```
**Physical identity of the four members.** This is how a future SMART alert is
matched to a caddy:
| Member | Model | Serial |
|---|---|---|
| `cciss,0` | `EG0146FAWHU` | `6SD0LETK0000B05009W1` |
| `cciss,1` | `EG0146FAWHU` | `3SD3K5P00000905094UN` |
| `cciss,2` | `EG0146FAWHU` | `3SD3GHJ600009047RRRG` |
| `cciss,3` | `EG0146FAWHU` | `3SD3GVP800009045X79B` |
**Not yet captured:** power-on hours and grown-defect counts. On 2010-vintage
drives holding the only instance, those numbers say how much runway remains.
`smartctl -d cciss,N -A /dev/sda` for each member.
### TLS
```
@@ -232,6 +268,9 @@ Each line was demonstrated by command output, not inferred.
- [x] NAT and FORWARD rules present live **and** persisted, in correct order
- [x] Host: `running`, zero failed units
- [x] **Mail delivered end to end from `root` on `srv-b`, twice, headers captured**
- [x] **`smartd` monitoring 4 devices; 4 test alerts delivered and received**
- [x] **Container SMTP egress blocked; internet egress and host mail intact**
- [x] **Both containers autostart after a host reboot** (`onboot: 1`)
### CT 100
@@ -278,19 +317,12 @@ Nothing. The infrastructure boundary is accepted.
Recorded here so they are visually distinct from failures and from forgotten
work. None is a defect.
- [ ] **SMTP egress block for the container network** (F-025). The containers
have no reason to originate mail. A `FORWARD` DROP on ports 25/465/587
from the service network to `10.110.0.0/22` is local, precise, and touches
no estate policy. `srv-b`'s own alerts are unaffected — they originate
locally and take OUTPUT, never FORWARD.
- [ ] **`wg-pk` `mynetworks` scope** (F-025). Estate decision: whether every
tunnel peer should originate mail as `kane-il.us` infrastructure, or only
the hosts that legitimately do. Not this project's to change unilaterally.
- [ ] **`smartd` explicit four-member configuration.** The package is installed
and the service is active, but `DEVICESCAN` currently monitors **zero
devices**; the P410i members are visible only through explicit
`-d cciss,N`. Active is not the same as monitoring. **Now unblocked** —
mail is accepted, so `-M test` gives a real end-to-end acceptance.
- [ ] **SMART attribute baseline.** Power-on hours and grown-defect counts for
the four members, so drive age is a known quantity rather than an
assumption. `smartctl -d cciss,N -A /dev/sda`. Read-only, one command.
- [ ] **Backup infrastructure, entirely.** Postponed 2026-08-16 because the
strategy may change: `vzdump` job, archive sizing, retention, free-space
guard, host-side pull, `mechcomp-backup`, backup alerting, gold media,
@@ -365,12 +397,11 @@ The Shapely port gates all of the above.
| # | Question | Blocks |
|---|---|---|
| 1 | Should container SMTP egress be blocked at `srv-b`? | F-025, ours to fix |
| 2 | Should `wg-pk` `mynetworks` narrow to explicit hosts? | F-025, estate decision |
| 3 | What is the backup strategy? | all backup work |
| 4 | Where is the 3+ TB USB disk attached? | gold redundancy step |
| 1 | Should `wg-pk` `mynetworks` narrow to explicit hosts? | F-025 estate half |
| 2 | What is the backup strategy? | all backup work |
| 3 | Where is the 3+ TB USB disk attached? | gold redundancy step |
All four need CIVICVS.
All three need CIVICVS.
---
@@ -388,3 +419,5 @@ All four need CIVICVS.
| `MECHCOMP_BIND` semantics | Scalar, service address only. Loopback not bound. |
| Why did delivery fail downstream? | `mx1` relay trust. Proven and corrected (F-023). |
| Where does Postfix log on this host? | journald. No `rsyslog`, no `/var/log/mail.log` (F-024). |
| Should container SMTP egress be blocked? | Yes. Implemented and reboot-proven (F-025 project-local half). |
| Do the containers autostart? | Yes, `onboot: 1` on both, proven by host reboot (F-026). |