Final configuration
This commit is contained in:
+71
-5
@@ -4,9 +4,9 @@ Specification for a Mechanical Compiler instance.
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| Revision | 5.1 (2026-08-17) |
|
||||
| Revision | 5.2 (2026-08-17) |
|
||||
| Supersedes | Revisions 1 through 4 |
|
||||
| Basis | Revision 4, reconciled against the proven staging build, then work order 002 |
|
||||
| Basis | Revision 4, reconciled against the proven staging build, then work orders 002 and 003 |
|
||||
| Scope | **Host-agnostic.** Applies to any instance. |
|
||||
| Instance state | `STAGING-STATE.md`, and later `PRODUCTION-STATE.md` |
|
||||
| Failure evidence | `FAILURES.md` |
|
||||
@@ -174,10 +174,18 @@ nat POSTROUTING
|
||||
-s <SVC_NET> -o <MGMT> -j MASQUERADE
|
||||
|
||||
filter FORWARD
|
||||
-s <SVC_NET> -d <WG_NET> -p tcp
|
||||
-m multiport --dports 25,465,587 -j DROP # containers do not send mail
|
||||
-s <SVC_NET> -d <SVC_NET> -j ACCEPT
|
||||
-s <SVC_NET> -d <LAN_NET> -j DROP
|
||||
```
|
||||
|
||||
**PROVEN (F-025)** — The SMTP rule is not optional where the host masquerades
|
||||
containers onto a network carrying a mail relay. Without it the containers
|
||||
inherit whatever relay authority the host's source address holds. The host's own
|
||||
alerting is unaffected: it originates mail locally, so it takes OUTPUT and never
|
||||
enters FORWARD. Prove that rather than assuming it.
|
||||
|
||||
**REQ** — Rule order is load-bearing. `RETURN` must precede both masquerades.
|
||||
Rules must be persisted, and the persisted set verified to match the live set.
|
||||
|
||||
@@ -254,6 +262,12 @@ large data volume in a container snapshot is how a backup store fills.
|
||||
**REQ** — Do not over-commit a thin pool. Thin-pool exhaustion is an ugly
|
||||
failure mode.
|
||||
|
||||
**PROVEN (F-026)** — Every container that must survive a host restart carries
|
||||
`onboot: 1`. This is not the default. It is also invisible to `pct reboot`,
|
||||
which restarts a guest without exercising host-boot autostart at all — so a
|
||||
container proven to survive `pct reboot` is **not** thereby proven to survive a
|
||||
host reboot.
|
||||
|
||||
### 4.3 Base OS
|
||||
|
||||
**REQ** — Debian 12 bookworm, both containers. It matches YunoHost 12 stable and
|
||||
@@ -638,12 +652,34 @@ coinciding with it.
|
||||
Automation that greps a logfile path finds nothing silently, which is worse than
|
||||
failing.
|
||||
|
||||
**REQ** — Disk monitoring must be configured **explicitly per device** where the
|
||||
controller does not present members to automatic scanning. An active
|
||||
**PROVEN** — Disk monitoring must be configured **explicitly per device** where
|
||||
the controller does not present members to automatic scanning. An active
|
||||
`smartmontools.service` monitoring zero devices is not monitoring, and is easy
|
||||
to mistake for success. Acceptance is a delivered test alert, not an active
|
||||
to mistake for success. Acceptance is a **delivered** test alert, not an active
|
||||
unit.
|
||||
|
||||
Behind an HP Smart Array controller the working pattern is:
|
||||
|
||||
```
|
||||
DEFAULT -a -m root -M exec /usr/share/smartmontools/smartd-runner
|
||||
/dev/sda -d cciss,0
|
||||
/dev/sda -d cciss,1
|
||||
/dev/sda -d cciss,2
|
||||
/dev/sda -d cciss,3
|
||||
```
|
||||
|
||||
`-M test` proves the alert path — one message per monitored device — and must be
|
||||
**reverted afterwards**. Left in place it produces an alert storm on every
|
||||
restart, which trains everyone to ignore disk alerts. That is worse than no
|
||||
monitoring.
|
||||
|
||||
**REQ (F-028)** — Query `smartmontools.service`, not `smartd.service`. The
|
||||
latter is an alias and owns no journal; `journalctl -u smartd` returns nothing
|
||||
on a host where monitoring is working perfectly.
|
||||
|
||||
**REQ** — Record each member's **serial number** at configuration time. That is
|
||||
the only thing that maps a future alert to a physical drive in a caddy.
|
||||
|
||||
### 12.1 Downstream infrastructure is out of bounds
|
||||
|
||||
**REQ** — An instance may configure its own relay client. It may not alter the
|
||||
@@ -705,6 +741,30 @@ committed; an example with empty values is.
|
||||
11. AGPL-3.0 §13 requires network users be offered the source. The web tier
|
||||
carries a visible source link to the repository. A licence obligation.
|
||||
|
||||
### 14.1 Test validity
|
||||
|
||||
**REQ (F-027)** — A test must be able to distinguish failure of the **tool**
|
||||
from failure of the **thing under test**.
|
||||
|
||||
```bash
|
||||
# WRONG — any failure of the wrapper reads as a pass
|
||||
run_in_guest "$c" 'connect to X' && echo "reachable — WRONG" || echo "blocked — correct"
|
||||
|
||||
# RIGHT — establish that the test could have observed the positive case
|
||||
if [ "$(guest_state "$c")" != "running" ]; then
|
||||
echo "CT $c NOT RUNNING — TEST INVALID"
|
||||
elif run_in_guest "$c" 'connect to X'; then
|
||||
echo "reachable — WRONG"
|
||||
else
|
||||
echo "blocked — correct"
|
||||
fi
|
||||
```
|
||||
|
||||
This applies to **every negative assertion** — "X is unreachable", "Y is
|
||||
absent", "Z is refused". A test whose failure mode is indistinguishable from
|
||||
success is worse than no test, because it manufactures confidence. In WO-003 the
|
||||
wrong form reported a pass against containers that were not running.
|
||||
|
||||
---
|
||||
|
||||
## 15. Acceptance gates
|
||||
@@ -728,6 +788,12 @@ scaffolding survives a reboot.
|
||||
**PROVEN (F-021)** — the reboot is not optional. Connectivity can be perfect
|
||||
while boot is degraded.
|
||||
|
||||
**PROVEN (F-026)** — it must be a **host** reboot, and every guest must return
|
||||
automatically and report `running`. A guest reboot does not exercise autostart.
|
||||
|
||||
**PROVEN (F-027)** — every negative assertion in this gate must satisfy §14.1
|
||||
before its result means anything.
|
||||
|
||||
### Gate 2 — Operational services
|
||||
|
||||
Mail delivered **end to end and received twice**, with full headers captured, an
|
||||
|
||||
+110
-2
@@ -6,7 +6,7 @@ Compiler environment.
|
||||
| | |
|
||||
|---|---|
|
||||
| Scope | All instances. Staging entries are marked `srv-b`. |
|
||||
| Updated | 2026-08-17, after work order 002 |
|
||||
| Updated | 2026-08-17, after work order 003 |
|
||||
| Rule | Append only. Never edit an entry except to add a `Resolution` line. |
|
||||
| Numbering | Sequential, never reused. See §0 on the renumbering. |
|
||||
|
||||
@@ -526,6 +526,111 @@ would break that. The question is which clients may ask, not where it may send.
|
||||
**Consequence:** A correction that fixes the observed failure may widen an
|
||||
adjacent boundary. Record what a trust change grants, not only what it repairs.
|
||||
|
||||
**Project-local correction (2026-08-17, WO-003):** Complete. A persistent
|
||||
`srv-b` FORWARD rule drops TCP 25, 465 and 587 from `10.20.0.0/24` to
|
||||
`10.110.0.0/22`. Reachability was proven **before** the change, both containers
|
||||
are blocked after it, ordinary internet egress is intact, host-originated mail
|
||||
still delivers end to end, and the rule survived a host reboot.
|
||||
|
||||
**Estate half: still open.** Whether `wg-pk` should narrow `mynetworks` from
|
||||
`10.110.0.0/22` to the explicit hosts that legitimately originate mail is a
|
||||
CIVICVS decision. No change was made to `wg-pk`, `mx1` or `kane-il.us`.
|
||||
|
||||
F-025 is therefore **partially corrected**, not closed. The two halves are kept
|
||||
in one entry rather than split, because the exposure is a single chain and
|
||||
splitting it would let one half be closed while the other is forgotten.
|
||||
|
||||
---
|
||||
|
||||
### F-026 — containers were not configured to start with the host
|
||||
`srv-b`, CT 100, CT 101. Work order 003.
|
||||
|
||||
**Observed:** After the WO-003 persistence reboot — the first true **host**
|
||||
reboot since the containers were built — both were stopped:
|
||||
|
||||
```
|
||||
pct status 100 -> status: stopped
|
||||
pct status 101 -> status: stopped
|
||||
```
|
||||
|
||||
**Cause:** **Proven.** Neither container configuration contained `onboot: 1`, so
|
||||
Proxmox had no instruction to start either guest. `pve-guests.service` was
|
||||
enabled, active, and its `startall` task completed successfully — proving the
|
||||
startup machinery worked and simply had nothing to do.
|
||||
**Correction:** `pct set 100 -onboot 1` and `pct set 101 -onboot 1`. The
|
||||
containers were deliberately left stopped so the correction could be proven by
|
||||
reboot rather than masked by a manual start. After a second host reboot both
|
||||
returned automatically and host plus both guests reported `running`.
|
||||
**Consequence:** **Specification defect.** `ENVIRONMENT.md` never required
|
||||
`onboot: 1`, and every earlier reboot in this project was `pct reboot` — which
|
||||
restarts a guest without ever exercising host-boot autostart. A container that
|
||||
survives `pct reboot` is not thereby proven to survive a host reboot. Gate 1
|
||||
must assert guest state after a **host** reboot specifically.
|
||||
|
||||
---
|
||||
|
||||
### F-027 — a verification test conflated tool failure with the tested condition
|
||||
`srv-b`. Work order 003.
|
||||
|
||||
**Observed:** WO-003 B3 and B4 used this pattern to test SMTP reachability from
|
||||
a container:
|
||||
|
||||
```bash
|
||||
pct exec $c -- timeout 5 bash -c '...' \
|
||||
&& echo "REACHABLE — WRONG" || echo "blocked — correct"
|
||||
```
|
||||
|
||||
With the containers stopped (F-026) it produced:
|
||||
|
||||
```
|
||||
-- CT 100: container '100' not running!
|
||||
blocked — correct
|
||||
```
|
||||
|
||||
**Cause:** **Proven.** The shell `||` branch catches *any* non-zero exit from
|
||||
`pct exec`, including failures of `pct exec` itself. The test could not
|
||||
distinguish "the firewall blocked the connection" from "the command never ran."
|
||||
**Correction:** Assert guest state before interpreting any network result:
|
||||
|
||||
```bash
|
||||
if [ "$(pct status "$c" | awk '{print $2}')" != "running" ]; then
|
||||
echo "CT $c NOT RUNNING — TEST INVALID"
|
||||
elif pct exec "$c" -- timeout 5 bash -c '...'; then
|
||||
echo "REACHABLE — WRONG"
|
||||
else
|
||||
echo "blocked — correct"
|
||||
fi
|
||||
```
|
||||
|
||||
**Consequence:** **A test whose failure mode is indistinguishable from success
|
||||
is worse than no test**, because it manufactures confidence. This one printed a
|
||||
pass while the thing under test did not exist. Only a separate health assertion
|
||||
caught it, and that was ordering luck rather than design. Every negative
|
||||
assertion — "X is unreachable", "Y is absent", "Z is refused" — must first
|
||||
establish that the test could have observed the positive case.
|
||||
|
||||
Applies retroactively: the isolation proofs in F-018 used the same shape. They
|
||||
happened to be run against containers that were up, and the result was
|
||||
independently corroborated, but the pattern was unsafe there too.
|
||||
|
||||
---
|
||||
|
||||
### F-028 — three command defects in work order 003
|
||||
`srv-b`. Work order 003. Minor, grouped.
|
||||
|
||||
**Observed and corrected in place by the operator:**
|
||||
|
||||
| Defect | Cause | Correction |
|
||||
|---|---|---|
|
||||
| Serial numbers absent from A1 output | `grep -E "Serial Number"` is case-sensitive; SCSI output says `Serial number:` | `grep -iE` with a case-insensitive alternation on product, model, serial and SMART support |
|
||||
| `journalctl -u smartd -b` returned `-- No entries --` | `smartd.service` is an alias; the real unit is `smartmontools.service` and owns the journal | Query `smartmontools` |
|
||||
| A3 `sed` would have produced a malformed directive | Prefix substitution left the trailing `-M exec ...` in place and duplicated `-m root` | Replace the whole `DEFAULT` line, in both A3 and A4 |
|
||||
|
||||
**Consequence:** Match on the canonical unit name, not a convenience alias.
|
||||
Prefer whole-line replacement over prefix patching in configuration edits. Use
|
||||
case-insensitive matching when parsing tool output whose field names vary by
|
||||
device class.
|
||||
|
||||
---
|
||||
|
||||
## Open, not closed
|
||||
@@ -539,6 +644,9 @@ adjacent boundary. Record what a trust change grants, not only what it repairs.
|
||||
| F-022 | **Open** — cause unproven, no correction applied. |
|
||||
| F-023 | **Closed** 2026-08-17. Cause proven at `mx1`; delivery proven twice. |
|
||||
| F-024 | **Closed** 2026-08-17. Specification corrected. |
|
||||
| F-025 | **Open** — correction pending operator decision. |
|
||||
| F-025 | **Partially corrected** 2026-08-17. Project-local half closed and reboot-proven; **estate half open**, awaiting a CIVICVS decision on `wg-pk`. |
|
||||
| F-026 | **Corrected** 2026-08-17. Reboot persistence proven. |
|
||||
| F-027 | **Corrected** 2026-08-17. Applies retroactively to F-018's proofs. |
|
||||
| F-028 | **Corrected** 2026-08-17. |
|
||||
|
||||
Everything else is closed with a proven cause and a proven correction.
|
||||
|
||||
+11
-1
@@ -128,7 +128,11 @@ a new call into the same machinery, not a new machine.
|
||||
operator decision; application deployment remains blocked on application
|
||||
code. See `STAGING-STATE.md`.
|
||||
- **Mail alerting accepted, 2026-08-17.** Delivered end to end and received
|
||||
twice. `smartd` is now unblocked. One open exposure is recorded as F-025.
|
||||
twice.
|
||||
- **Disk monitoring accepted, 2026-08-17.** Four P410i members explicitly
|
||||
monitored, four test alerts received, serial numbers recorded, reboot-proven.
|
||||
- **Container SMTP egress blocked, 2026-08-17.** The project-local half of F-025
|
||||
is corrected; the estate question about `wg-pk` client trust remains open.
|
||||
|
||||
### Next — the port
|
||||
|
||||
@@ -254,3 +258,9 @@ on whether it makes distributed manufacturing capacity legible.
|
||||
10. **Shared infrastructure is not ours to reconfigure.** An instance may
|
||||
configure its own client side; changes further down the chain are escalated
|
||||
with evidence.
|
||||
11. **A test must distinguish tool failure from the condition it tests.** One
|
||||
whose failure mode is indistinguishable from success manufactures
|
||||
confidence (F-027). Every negative assertion first establishes that it could
|
||||
have observed the positive case.
|
||||
12. **A component surviving its own restart is not proven to survive the
|
||||
host's.** Different machinery; assert against the one that matters (F-026).
|
||||
|
||||
+50
-17
@@ -4,7 +4,7 @@ Live state of the Mechanical Compiler staging instance on `srv-b`.
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| Updated | 2026-08-17, after work order 002 (mail) |
|
||||
| Updated | 2026-08-17, after work order 003 (monitoring, SMTP egress) |
|
||||
| Instance | Staging / development |
|
||||
| Specification | `ENVIRONMENT.md` revision 5 |
|
||||
| Failure log | `FAILURES.md` |
|
||||
@@ -37,7 +37,7 @@ and must not be represented as either hidden failures or completed work:
|
||||
| Subsystem | Status |
|
||||
|---|---|
|
||||
| Mail alert delivery | **Accepted 2026-08-17.** Delivered end to end, twice, headers captured. |
|
||||
| `smartd` monitoring | Deferred. Now unblocked — mail works. |
|
||||
| `smartd` monitoring | **Accepted 2026-08-17.** Four members monitored, four alerts received. |
|
||||
| Backup infrastructure | **Postponed by operator decision** — strategy may change |
|
||||
|
||||
---
|
||||
@@ -77,6 +77,7 @@ template local:vztmpl/debian-12-standard_12.12-1_amd64.tar.zst
|
||||
| rootfs | 40 GiB `local-lvm` | 8 GiB `local-lvm` |
|
||||
| `mp0` | 60 GiB → `/var/lib/mechcomp`, `backup=0` | — |
|
||||
| `net0` | `eth0` on `vmbr1`, `10.20.0.10/24`, gw `10.20.0.1` | `eth0` on `vmbr1`, `10.20.0.11/24`, gw `10.20.0.1` |
|
||||
| `onboot` | `1` | `1` |
|
||||
| Debian | 12 bookworm, fully upgraded | 12 bookworm, fully upgraded |
|
||||
| systemd | running, zero failed units | running, zero failed units |
|
||||
|
||||
@@ -108,6 +109,8 @@ nat POSTROUTING
|
||||
-s 10.20.0.0/24 -o vmbr0 -j MASQUERADE # containers -> internet
|
||||
|
||||
filter FORWARD
|
||||
-s 10.20.0.0/24 -d 10.110.0.0/22 -p tcp
|
||||
-m multiport --dports 25,465,587 -j DROP # SMTP egress blocked
|
||||
-s 10.20.0.0/24 -d 10.20.0.0/24 -j ACCEPT # container <-> container
|
||||
-s 10.20.0.0/24 -d 10.0.0.0/24 -j DROP # LAN blocked
|
||||
```
|
||||
@@ -115,6 +118,39 @@ filter FORWARD
|
||||
Rule order is load-bearing. `RETURN` must precede both service-network
|
||||
masquerades. Live and persisted states match.
|
||||
|
||||
### Disk monitoring — accepted 2026-08-17
|
||||
|
||||
`DEVICESCAN` finds nothing behind the P410i. Members are addressed explicitly:
|
||||
|
||||
```
|
||||
/etc/smartd.conf
|
||||
DEFAULT -a -m root -M exec /usr/share/smartmontools/smartd-runner
|
||||
/dev/sda -d cciss,0
|
||||
/dev/sda -d cciss,1
|
||||
/dev/sda -d cciss,2
|
||||
/dev/sda -d cciss,3
|
||||
|
||||
before: Monitoring 0 ATA/SATA, 0 SCSI/SAS and 0 NVMe devices
|
||||
after: Monitoring 0 ATA/SATA, 4 SCSI/SAS and 0 NVMe devices
|
||||
unit: smartmontools.service (smartd.service is an alias with no journal)
|
||||
alerts: 4 test alerts generated and received; headers captured
|
||||
-M test reverted; permanent config restored and reboot-proven
|
||||
```
|
||||
|
||||
**Physical identity of the four members.** This is how a future SMART alert is
|
||||
matched to a caddy:
|
||||
|
||||
| Member | Model | Serial |
|
||||
|---|---|---|
|
||||
| `cciss,0` | `EG0146FAWHU` | `6SD0LETK0000B05009W1` |
|
||||
| `cciss,1` | `EG0146FAWHU` | `3SD3K5P00000905094UN` |
|
||||
| `cciss,2` | `EG0146FAWHU` | `3SD3GHJ600009047RRRG` |
|
||||
| `cciss,3` | `EG0146FAWHU` | `3SD3GVP800009045X79B` |
|
||||
|
||||
**Not yet captured:** power-on hours and grown-defect counts. On 2010-vintage
|
||||
drives holding the only instance, those numbers say how much runway remains.
|
||||
`smartctl -d cciss,N -A /dev/sda` for each member.
|
||||
|
||||
### TLS
|
||||
|
||||
```
|
||||
@@ -232,6 +268,9 @@ Each line was demonstrated by command output, not inferred.
|
||||
- [x] NAT and FORWARD rules present live **and** persisted, in correct order
|
||||
- [x] Host: `running`, zero failed units
|
||||
- [x] **Mail delivered end to end from `root` on `srv-b`, twice, headers captured**
|
||||
- [x] **`smartd` monitoring 4 devices; 4 test alerts delivered and received**
|
||||
- [x] **Container SMTP egress blocked; internet egress and host mail intact**
|
||||
- [x] **Both containers autostart after a host reboot** (`onboot: 1`)
|
||||
|
||||
### CT 100
|
||||
|
||||
@@ -278,19 +317,12 @@ Nothing. The infrastructure boundary is accepted.
|
||||
Recorded here so they are visually distinct from failures and from forgotten
|
||||
work. None is a defect.
|
||||
|
||||
- [ ] **SMTP egress block for the container network** (F-025). The containers
|
||||
have no reason to originate mail. A `FORWARD` DROP on ports 25/465/587
|
||||
from the service network to `10.110.0.0/22` is local, precise, and touches
|
||||
no estate policy. `srv-b`'s own alerts are unaffected — they originate
|
||||
locally and take OUTPUT, never FORWARD.
|
||||
- [ ] **`wg-pk` `mynetworks` scope** (F-025). Estate decision: whether every
|
||||
tunnel peer should originate mail as `kane-il.us` infrastructure, or only
|
||||
the hosts that legitimately do. Not this project's to change unilaterally.
|
||||
- [ ] **`smartd` explicit four-member configuration.** The package is installed
|
||||
and the service is active, but `DEVICESCAN` currently monitors **zero
|
||||
devices**; the P410i members are visible only through explicit
|
||||
`-d cciss,N`. Active is not the same as monitoring. **Now unblocked** —
|
||||
mail is accepted, so `-M test` gives a real end-to-end acceptance.
|
||||
- [ ] **SMART attribute baseline.** Power-on hours and grown-defect counts for
|
||||
the four members, so drive age is a known quantity rather than an
|
||||
assumption. `smartctl -d cciss,N -A /dev/sda`. Read-only, one command.
|
||||
- [ ] **Backup infrastructure, entirely.** Postponed 2026-08-16 because the
|
||||
strategy may change: `vzdump` job, archive sizing, retention, free-space
|
||||
guard, host-side pull, `mechcomp-backup`, backup alerting, gold media,
|
||||
@@ -365,12 +397,11 @@ The Shapely port gates all of the above.
|
||||
|
||||
| # | Question | Blocks |
|
||||
|---|---|---|
|
||||
| 1 | Should container SMTP egress be blocked at `srv-b`? | F-025, ours to fix |
|
||||
| 2 | Should `wg-pk` `mynetworks` narrow to explicit hosts? | F-025, estate decision |
|
||||
| 3 | What is the backup strategy? | all backup work |
|
||||
| 4 | Where is the 3+ TB USB disk attached? | gold redundancy step |
|
||||
| 1 | Should `wg-pk` `mynetworks` narrow to explicit hosts? | F-025 estate half |
|
||||
| 2 | What is the backup strategy? | all backup work |
|
||||
| 3 | Where is the 3+ TB USB disk attached? | gold redundancy step |
|
||||
|
||||
All four need CIVICVS.
|
||||
All three need CIVICVS.
|
||||
|
||||
---
|
||||
|
||||
@@ -388,3 +419,5 @@ All four need CIVICVS.
|
||||
| `MECHCOMP_BIND` semantics | Scalar, service address only. Loopback not bound. |
|
||||
| Why did delivery fail downstream? | `mx1` relay trust. Proven and corrected (F-023). |
|
||||
| Where does Postfix log on this host? | journald. No `rsyslog`, no `/var/log/mail.log` (F-024). |
|
||||
| Should container SMTP egress be blocked? | Yes. Implemented and reboot-proven (F-025 project-local half). |
|
||||
| Do the containers autostart? | Yes, `onboot: 1` on both, proven by host reboot (F-026). |
|
||||
|
||||
Reference in New Issue
Block a user