This commit is contained in:
2026-08-16 20:25:18 -04:00
parent eac601ed99
commit e67988eef5
4 changed files with 727 additions and 828 deletions
+427 -715
View File
File diff suppressed because it is too large Load Diff
+116 -4
View File
@@ -6,6 +6,7 @@ Compiler environment.
| | |
|---|---|
| Scope | All instances. Staging entries are marked `srv-b`. |
| Updated | 2026-08-16, after staging acceptance |
| Rule | Append only. Never edit an entry except to add a `Resolution` line. |
| Numbering | Sequential, never reused. See §0 on the renumbering. |
@@ -315,11 +316,19 @@ CT 100.
**Observed:** After the F-017 reboot, nothing listened on `10.20.0.10:8770`.
**Cause:** The placeholder was a bare foreground process with a PID file. No
supervision.
**Correction:** Pending — to be recreated as `mechcomp-placeholder.service`.
**Correction:** Recreated as `mechcomp-placeholder.service`.
**Consequence:** Anything a proof depends on must be supervised, or the proof
expires silently at the next reboot and the following session diagnoses a
proxy fault that does not exist. Applies to temporary scaffolding as much as to
real services.
**Resolution (2026-08-16):** Installed `/usr/local/libexec/mechcomp-placeholder.py`
and `/etc/systemd/system/mechcomp-placeholder.service`, enabled at
`multi-user.target`, running as `mechcomp:mechcomp`, reading
`/etc/mechcomp/mechcomp.env`, binding `10.20.0.10:8770`, with `Restart=on-failure`
and the standard hardening set. CT 100 was rebooted: the unit restarted
automatically, the listener returned on the service address only, nginx proxied
successfully, `X-Forwarded-Proto: https` was re-observed at the backend, and the
container settled to `running` with zero failed units.
---
@@ -340,10 +349,113 @@ topology decision. Recorded because the reasoning matters more than the value:
---
### F-021 — deleting the Proxmox interface left a stale guest `eth1`
CT 100, CT 101. Network isolation.
**Observed:** After the F-017 topology change and reboot, both containers
reported `systemctl is-system-running` → `degraded`, with
`networking.service` and `ifupdown-wait-online.service` failed. Both guests'
`/etc/network/interfaces` still carried an `auto eth1` / static `iface eth1`
stanza although neither had an `eth1` link. The boot journal on both:
```
Cannot find device "eth1"
ifup: failed to bring up eth1
```
**Cause:** **Proven.** `pct set --delete net1` removes the LXC interface but
does not remove the stanza already written into the guest. `ifup -a` therefore
exited non-zero at boot even though `eth0` came up correctly.
`ifupdown-wait-online` failed as a consequence, not independently.
`systemd-networkd` was investigated and explicitly ruled out — its units were
disabled on both containers.
**Correction:** Preserved a pre-correction copy, removed only the stale `eth1`
stanza on each guest, restarted the affected units. Both containers then
rebooted: the stanza did not return, both units succeeded, both settled to
`running` with zero failed units.
**Consequence:** **Removing a container interface is not proof that the guest
converged.** After any topology mutation, automation must inspect the guest
interface file and assert `systemctl is-system-running = running` with zero
failed units *after a reboot*. Connectivity alone is insufficient — the
surviving interface works fine while boot remains degraded, which is precisely
how this went unnoticed through an entire verification pass.
---
### F-022 — transient DNS resolution timeout in CT 101
CT 101. Formal acceptance.
**Observed:** During the first acceptance pass, two unrelated public HTTPS
tests both failed at name resolution:
```
deb.debian.org: curl: (28) Resolving timed out after 5000 ms
gitea.barternetwork.us: curl: (28) Resolving timed out after 5000 ms
```
IP routing to `10.110.0.1`, `10.0.0.12` and CT 100 remained working throughout.
**Cause:** **Unproven.** Diagnostics showed resolver configuration identical to
CT 100 and the host (`nameserver 75.75.75.75`, `hosts: files dns`),
`systemd-resolved` absent, `75.75.75.75` reachable, and `getent ahostsv4`
resolving both names immediately afterward. Repeat HTTPS tests returned 200.
Worth noting as context, not as cause: since F-017, container DNS traverses the
host's masquerade to an external resolver. That dependency is new.
**Correction:** None. No configuration was changed.
**Consequence:** Do not convert a one-shot resolver timeout into a
configuration change without evidence. On recurrence, capture resolver state
and DNS traffic at the moment of failure before touching anything. A candidate
mitigation — a second `nameserver` line, so a single hiccup retries rather than
fails — is recorded but deliberately not applied on one unexplained event.
---
### F-023 — relay accepted the mail; final delivery failed
`srv-b`. Alerting.
**Observed:** The relay was discovered at `10.110.0.1:25` over `wg0`, banner
`wg-pk.diagnostics.kane-il.us`, offering STARTTLS with a self-signed
`CN = wg-pk`. It accepted unauthenticated SMTP from `10.110.0.12` through
`RCPT TO`, before and after STARTTLS. Ports 465 and 587 were unavailable, and
the public address `198.58.111.109` did not expose SMTP on this path.
With `relayhost = [10.110.0.1]:25` and `root: sandor@kane-il.us`, `srv-b`
recorded successful handoff:
```
relay=10.110.0.1[10.110.0.1]:25 dsn=2.0.0
status=sent (250 2.0.0 Ok: queued as 87E446243A)
```
The local queue emptied. The operator then received a **delivery-failure**
message at `sandor@kane-il.us`.
**Cause:** **Unproven**, downstream of the demonstrated handoff. The bounce
notice itself arriving at `sandor@kane-il.us` establishes that the relay can
deliver to that address, which narrows the problem to the failing message
rather than the destination. Leading hypothesis, untested: the envelope sender
is `root@srv-b.dev.infra`, and `dev.infra` does not resolve publicly, so a
downstream MTA rejects on sender-domain verification. Candidate remedies are
`myorigin` or `smtp_generic_maps`.
**Correction:** None. Mail alerting was deferred by operator decision.
**Consequence:** **SMTP 250 from the relay and an empty local queue prove
handoff, not delivery.** Alerting acceptance requires demonstrated end-to-end
receipt. Until then mail, `smartd` alerting, and any mail-dependent backup
alerting are unaccepted. Note separately that `postfix check` reports
divergence between `/var/spool/postfix` copies and their host originals,
including `/etc/hosts` and NSS libraries — a known cause of resolution failure
inside the chroot, and adjacent enough to this failure to be checked first.
---
## Open, not closed
| # | Status |
|---|---|
| F-006 | Cause unproven. Recurrence should capture `dpkg` lock state. |
| F-012 | Cause unproven. Leading candidate ruled out by inspection. |
| F-019 | Correction pending. |
| F-006 | **Open** — cause unproven. Recurrence should capture `dpkg` lock state. |
| F-012 | **Open** — cause unproven. Leading candidate ruled out by inspection. `default_server` was added as independent hardening and does **not** close this. |
| F-019 | **Corrected** 2026-08-16. Reboot persistence proven. |
| F-021 | **Corrected** 2026-08-16. Reboot persistence proven. |
| F-022 | **Open** — cause unproven, no correction applied. |
| F-023 | **Deferred** — downstream mail failure, cause unproven. Hypothesis recorded. |
Everything else is closed with a proven cause and a proven correction.
+13 -2
View File
@@ -4,7 +4,7 @@ What the Mechanical Compiler is for, and the order in which it gets built.
| | |
|---|---|
| Updated | 2026-08-15 |
| Updated | 2026-08-16 |
| Companions | `ENVIRONMENT.md`, `STAGING-STATE.md`, `FAILURES.md` |
---
@@ -121,7 +121,12 @@ a new call into the same machinery, not a new machine.
provably non-chiral; profile parameters are isolated from one another.
- **Fixture oracle frozen.** 123 cases, `ddd0f154…`, pinned to OpenSCAD 2021.01
and BOSL2 `92d697c`. The ten rejected cases are part of the contract.
- **Staging environment.** In progress — see `STAGING-STATE.md`.
- **Staging infrastructure accepted, 2026-08-16.** Host, both containers,
network isolation, bastion access, TLS trust, reverse proxy and the
application filesystem foundation are proven on `srv-b`. Mail alerting and
explicit disk monitoring are deferred; backup strategy is postponed by
operator decision; application deployment remains blocked on application
code. See `STAGING-STATE.md`.
### Next — the port
@@ -235,3 +240,9 @@ on whether it makes distributed manufacturing capacity legible.
6. **Record the failure before correcting it.** Then apply the smallest
corrective change, not the one that also fixes three things you were
worried about.
7. **Prove the negative.** Isolation is asserted by showing the forbidden path
fails, never by showing the interface is gone (F-018). Health is asserted
after a reboot, never before (F-021).
8. **Handoff is not delivery.** An upstream acceptance code proves the message
left, not that it arrived (F-023). The same distinction applies wherever a
subsystem reports success on behalf of something downstream.
+171 -107
View File
@@ -4,7 +4,7 @@ Live state of the Mechanical Compiler staging instance on `srv-b`.
| | |
|---|---|
| Updated | 2026-08-15, after network isolation |
| Updated | 2026-08-16, formal infrastructure acceptance |
| Instance | Staging / development |
| Specification | `ENVIRONMENT.md` revision 5 |
| Failure log | `FAILURES.md` |
@@ -15,24 +15,36 @@ Live state of the Mechanical Compiler staging instance on `srv-b`.
## 0. How to use this file
This is the authoritative record of what is true on `srv-b`. Where it and
`ENVIRONMENT.md` disagree, **this file wins for facts** and the specification
is defective and must be corrected.
`ENVIRONMENT.md` disagree, **this file wins for facts** and the specification is
defective and must be corrected.
Completed and remaining work are in the same document deliberately. They are
two halves of one boundary; separating them guarantees they drift.
Completed and remaining work are in the same document deliberately. They are two
halves of one boundary; separating them guarantees they drift.
Before any command: read this file, confirm the immediately relevant live state
with a read-only command, then issue one command group. If it fails, record it
in `FAILURES.md` before changing anything else.
with a read-only command, then issue one command group. If it fails, record it in
`FAILURES.md` before changing anything else.
Production will get its own `PRODUCTION-STATE.md`. The specification is shared;
the state is not.
Production gets its own `PRODUCTION-STATE.md`. The specification is shared; the
state is not.
### Acceptance boundary, 2026-08-16
**Staging infrastructure is accepted.** That claim is narrower than "the
application is deployed," and deliberately so. Three subsystems sit outside it
and must not be represented as either hidden failures or completed work:
| Subsystem | Status |
|---|---|
| Mail alert delivery | Partially configured, characterised, **not accepted end to end** |
| `smartd` monitoring | Deferred with mail alerting |
| Backup infrastructure | **Postponed by operator decision** — strategy may change |
---
## 1. Instance values
These are the `srv-b` bindings for the parameters in `ENVIRONMENT.md`.
The `srv-b` bindings for the parameters in `ENVIRONMENT.md`.
### Host
@@ -40,13 +52,16 @@ These are the `srv-b` bindings for the parameters in `ENVIRONMENT.md`.
hostname srv-b / srv-b.dev.infra
platform Proxmox VE 8.4.0, Debian 12, kernel 6.8.12-9-pve
hardware HP ProLiant DL360 G7, 2 x Xeon X5650, 24 threads, 31 GiB
storage P410i, 4 x EG0146FAWHU, RAID 1+0, all SMART OK
storage P410i, 4 x EG0146FAWHU, RAID 1+0, all members SMART OK
local directory /var/lib/vz, ~70 GiB free iso,vztmpl,backup
local-lvm LVM-thin pve/data, 166.9 GiB rootdir,images
vmbr0 10.0.0.12/24 on enp3s0f0, gw 10.0.0.1 management, LAN
vmbr1 10.20.0.1/24, bridge-ports none service, portless
wg0 10.110.0.12/32, peer wg-pk.civicus.us:51820, allowed 10.110.0.0/22
resolver 75.75.75.75, search dev.infra
ip_forward 1
timezone America/Chicago, NTP active, clock synchronised
systemd running, zero failed units
spare NICs enp3s0f1, enp4s0f0, enp4s0f1 — unconfigured, deliberately
template local:vztmpl/debian-12-standard_12.12-1_amd64.tar.zst
```
@@ -62,9 +77,11 @@ template local:vztmpl/debian-12-standard_12.12-1_amd64.tar.zst
| rootfs | 40 GiB `local-lvm` | 8 GiB `local-lvm` |
| `mp0` | 60 GiB → `/var/lib/mechcomp`, `backup=0` | — |
| `net0` | `eth0` on `vmbr1`, `10.20.0.10/24`, gw `10.20.0.1` | `eth0` on `vmbr1`, `10.20.0.11/24`, gw `10.20.0.1` |
| Debian | 12.15 | 12.15 |
| Debian | 12 bookworm, fully upgraded | 12 bookworm, fully upgraded |
| systemd | running, zero failed units | running, zero failed units |
**Neither container has a LAN interface.** See F-017.
**Neither container has a LAN interface** (F-017). **Neither has a stale `eth1`
stanza** (F-021) — verified persistent across reboot.
### Names and identity
@@ -73,16 +90,15 @@ mechanical-compiler.dev.infra -> 10.20.0.11 (CT 101, the proxy)
mechcomp.dev.infra -> 10.20.0.10 PVE-generated
mcproxy.dev.infra -> 10.20.0.11 PVE-generated
ADMIN_USER sandor
uid 1000, groups sandor + mechcomp(996)
shell /bin/bash, both containers
authorized key SHA256:2pNffCscUUW5Wbs9uMepLEvMLKPqUV/7Lpk5tw7iWSY
root@srv-b — bastion pattern, see §4
service user mechcomp, uid 999, gid 996
home /var/www/mechcomp, shell /usr/sbin/nologin
ADMIN_USER sandor, uid 1000, groups sandor + mechcomp(996)
shell /bin/bash, both containers
authorized key SHA256:2pNffCscUUW5Wbs9uMepLEvMLKPqUV/7Lpk5tw7iWSY
matches /root/.ssh/id_rsa.pub on srv-b — bastion, see section 4
service user mechcomp, uid 999, gid 996
home /var/www/mechcomp, shell /usr/sbin/nologin
```
### Host NAT, as persisted in `/etc/iptables/rules.v4`
### Host NAT, persisted in `/etc/iptables/rules.v4`
```
nat POSTROUTING
@@ -96,7 +112,8 @@ filter FORWARD
-s 10.20.0.0/24 -d 10.0.0.0/24 -j DROP # LAN blocked
```
Rule order is load-bearing. `RETURN` must precede both masquerades.
Rule order is load-bearing. `RETURN` must precede both service-network
masquerades. Live and persisted states match.
### TLS
@@ -104,15 +121,14 @@ Rule order is load-bearing. `RETURN` must precede both masquerades.
CA CN = Mechanical Compiler Staging CA (locally generated)
leaf CN = mechanical-compiler.dev.infra
SAN = DNS:mechanical-compiler.dev.infra
validity 2026-08-16 -> 2028-11-18 <-- expires, nothing renews it
validity 2026-08-16 -> 2028-11-18 <-- nothing renews this
trusted srv-b, CT 100, CT 101
key mode 0600 on CT 101
key root:root 0600 /etc/ssl/mechcomp/server.key
```
No Kane County Civic Infrastructure CA issuance path exists on `srv-b`: no
trust anchor, no `step`, no `cfssl`, no EasyRSA. Proxmox's own CA was
deliberately not reused. Replacing this leaf later is two file copies and a
reload.
No Kane County Civic Infrastructure CA issuance path exists on `srv-b`: no trust
anchor, no `step`, no `cfssl`, no EasyRSA. Proxmox's own CA was deliberately not
reused. Replacing this leaf later is two file copies and a reload.
### Application environment
@@ -120,7 +136,7 @@ reload.
```
MECHCOMP_ENV=staging
MECHCOMP_BIND=10.20.0.10
MECHCOMP_BIND=10.20.0.10 # scalar; loopback is NOT bound
MECHCOMP_PORT=8770
MECHCOMP_BASE_URL=https://mechanical-compiler.dev.infra
MECHCOMP_DATA_DIR=/var/lib/mechcomp
@@ -132,6 +148,36 @@ MECHCOMP_ARTIFACT_RETENTION_DAYS=30
MECHCOMP_SECRET_KEY=<generated, not recorded>
```
### Mail, as currently configured
```
relay endpoint 10.110.0.1:25 over wg0
banner wg-pk.diagnostics.kane-il.us, STARTTLS offered
STARTTLS cert self-signed, CN = wg-pk
authentication unauthenticated accepted from 10.110.0.12
ports 465 / 587 unavailable; public 198.58.111.109 exposes no SMTP on this path
srv-b relayhost [10.110.0.1]:25
smtp_tls_security_level may
smtp_sasl_auth_enable no
inet_interfaces loopback-only
root alias sandor@kane-il.us
local handoff succeeds, relay returns SMTP 250, queue empties
FINAL DELIVERY NOT ACCEPTED — operator received a delivery-failure message
status deferred, cause unproven downstream (F-023)
```
### Placeholder backend — staging scaffold, not application code
```
/usr/local/libexec/mechcomp-placeholder.py
/etc/systemd/system/mechcomp-placeholder.service
runs as mechcomp:mechcomp, reads mechcomp.env, binds 10.20.0.10:8770
Restart=on-failure, hardening set applied, enabled at multi-user.target
```
Exists solely to prove the proxy chain independently of the application. It is
removed when the real service arrives.
### Rollback copies on disk
```
@@ -143,7 +189,9 @@ srv-b /etc/network/interfaces.before-mechcomp
/root/pct-100.before-svcnet
/root/pct-101.before-svcnet
CT 100 /etc/hosts.before-svcfqdn
/etc/network/interfaces pre-F-021 copy
CT 101 /etc/hosts.before-svcfqdn
/etc/network/interfaces pre-F-021 copy
```
---
@@ -156,122 +204,136 @@ Each line was demonstrated by command output, not inferred.
- [x] `vmbr1` created, active, `10.20.0.1/24`, portless
- [x] Duplicate address detection run before each container creation
- [x] Container → internet reachable (`deb.debian.org`, `gitea.barternetwork.us`, both 200)
- [x] Container → WireGuard network reachable (`10.110.0.12`)
- [x] **Container → LAN unreachable** — the isolation requirement
- [x] Container → `srv-b` reachable at `10.0.0.12` (INPUT path, required)
- [x] Container ↔ container reachable over `vmbr1`
- [x] LAN → Proxmox console unaffected (`10.0.0.12:8006` → 200)
- [x] Network configuration persisted, survives reboot
- [x] Container to internet reachable (`deb.debian.org`, `gitea.barternetwork.us`)
- [x] Container to WireGuard reachable (`10.110.0.1`)
- [x] **Container to LAN gateway `10.0.0.1` = 100% packet loss** — the requirement
- [x] Container to `srv-b` `10.0.0.12` reachable via INPUT path (required, not a leak)
- [x] Container to container reachable over `vmbr1`
- [x] LAN to Proxmox console `10.0.0.12:8006` returns 200, unaffected
- [x] NAT and FORWARD rules present live **and** persisted, in correct order
- [x] Host: `running`, zero failed units
### CT 100
- [x] Debian 12.15, systemd running, zero failed units, no pending upgrades
- [x] `en_US.UTF-8` generated and active
- [x] Debian 12 bookworm, fully upgraded, zero pending
- [x] `en_US.UTF-8` generated and active; `America/Chicago`
- [x] **`running`, zero failed units — after reboot** (F-021 corrected)
- [x] `/var/lib/mechcomp` is a real separate ext4 filesystem
- [x] Service user `mechcomp`, home `/var/www/mechcomp`, writable
- [x] Application directory layout created with correct ownership
- [x] Base packages installed; OpenSCAD / Qt / X11 absent — verified `clean`
- [x] Docker working, `overlay2` / `systemd`, no `fuse-overlayfs` needed
- [x] Service user `mechcomp`, home `/var/www/mechcomp`, correct ownership
- [x] All four data subdirectories present, `mechcomp:mechcomp`, `0750`
- [x] `mechcomp.env` present, `root:mechcomp`, `0640`, secret generated
- [x] Base packages installed; OpenSCAD, Qt and X11 absent
- [x] Docker active, `overlay2` / `systemd`, no fallback needed
- [x] Repository cloned at `e85c4f4e`, verified as the owning user
- [x] Python 3.11.2 venv created, owned by `mechcomp`
- [x] `mechcomp.env` written with generated secret
- [x] `openssh-server` installed, enabled, active
- [x] Application listener bound to service network only — LAN-side bind refused
- [x] `openssh-server` enabled and active
- [x] **`mechcomp-placeholder.service` enabled, active, reboot-persistent** (F-019)
- [x] Listener on `10.20.0.10:8770` only — not `0.0.0.0`
### CT 101
- [x] Debian 12.15, systemd running, zero failed units after `nesting=1`
- [x] `en_US.UTF-8` generated and active
- [x] nginx 1.22.1 installed, `nginx -t` passes
- [x] Debian `default` site removed; only the project vhost is enabled
- [x] Debian 12 bookworm, fully upgraded, zero pending
- [x] `en_US.UTF-8` generated and active; `America/Chicago`
- [x] **`running`, zero failed units — after reboot** (F-021 corrected)
- [x] nginx 1.22.1, `nginx -t` passes, enabled and active
- [x] Only the project vhost enabled; Debian `default` removed
- [x] `listen 443 ssl default_server` on both address families
- [x] HTTP to HTTPS 301 redirect
- [x] Local CA created, leaf issued, trusted on all three hosts
- [x] HTTP → HTTPS redirect, TLS termination, proxy to `10.20.0.10:8770`
- [x] `openssh-server` installed, enabled, active
- [x] **`X-Forwarded-Proto: https` observed at the backend** — needs re-proof, see §3
- [x] **HTTPS end to end from `srv-b`, CT 100 and CT 101 without `-k`**
- [x] **`X-Forwarded-Proto: https` observed at the backend** — re-proved after the
topology change and again after CT 100's reboot
- [x] `openssh-server` enabled and active
---
## 3. Remaining — infrastructure
In dependency order. Application-dependent work is in §5.
### Immediate
- [ ] **`mechcomp-placeholder.service`** — recreate the backend as a supervised
unit. It died on the F-017 reboot and nothing currently listens on 8770.
Everything below that touches the proxy depends on this. See F-019.
- [ ] **Re-prove `X-Forwarded-Proto`** end to end. The earlier proof was taken
when clients were on `10.0.0.x`; header values will now read `10.20.0.x`.
Re-establish rather than assume it survived the topology change.
- [ ] **`default_server`** on the CT 101 vhost, both `listen 443` lines. Two
words. Prevents catch-all behaviour depending on file ordering once a
second server block exists. See F-012.
Nothing. The infrastructure boundary is accepted.
### Mail — blocks alerting
### Deferred by operator decision
- [ ] **Postfix relay.** Currently `inet_interfaces = loopback-only`,
`relayhost` empty, no `root:` alias. Alerts go to a mailbox nobody opens.
- [ ] **`root:` alias** to a real destination.
- [ ] Needs: address of the `wg-pk` relay, and whether it accepts
unauthenticated from `10.110.0.12`. **Open question for CIVICVS.**
Recorded here so they are visually distinct from failures and from forgotten
work. None is a defect.
### Monitoring and backup
- [ ] **Mail end-to-end delivery** (F-023). Handoff to `wg-pk` works; final
delivery does not. Leading untested hypothesis: envelope sender
`root@srv-b.dev.infra` rejected on sender-domain verification, since
`dev.infra` does not resolve publicly. Candidate remedies `myorigin` or
`smtp_generic_maps`. Check the `postfix check` chroot divergence first.
- [ ] **`smartd` explicit four-member configuration.** The package is installed
and the service is active, but `DEVICESCAN` currently monitors **zero
devices**; the P410i members are visible only through explicit
`-d cciss,N`. Active is not the same as monitoring. Deferred with mail.
- [ ] **Backup infrastructure, entirely.** Postponed 2026-08-16 because the
strategy may change: `vzdump` job, archive sizing, retention, free-space
guard, host-side pull, `mechcomp-backup`, backup alerting, gold media,
3+ TB redundancy.
- [ ] **`smartd`** on `/dev/sda -d cciss,0` through `cciss,3`. Depends on mail.
- [ ] **`vzdump` job** — both containers, `local`, snapshot, zstd, 02:30.
`mp0` already carries `backup=0`.
- [ ] **First `vzdump` run — report the archive size.** Retention is a
placeholder `keep-last=3` until this number exists.
- [ ] **Free-space guard** — refuse and alert below 20 GiB on `/var/lib/vz`.
- [ ] `/var/lib/vz/mechcomp-app/` and the host-side pull script.
- [ ] `mechcomp-backup --stdout` in CT 100. Structure can be built and tested
against the current data directories without application output.
**One request standing against the postponement:** a single manual `vzdump` of
both containers as a point-in-time snapshot before further change. It
presupposes nothing about the eventual strategy and yields the archive size that
has been an open question since revision 4. Four 2010-vintage disks currently
hold the only instance with no monitoring, no alerting and no backup.
### Gold media
### Optional, recorded not scheduled
- [ ] 32 GB USB stick: LUKS2, ext4, label `MC-GOLD-01`, `/mnt/gold` `noauto`.
- [ ] `gold-archive.sh` with `SHA256SUMS` verification before unmount.
- [ ] **Open:** where the 3+ TB USB disk is attached now that `annales` is out
of scope. Until known, stick-to-disk copying is a manual step.
- [ ] Second `nameserver` line in container resolver configuration, as a
mitigation for F-022. One line, no daemon. Deliberately not applied on a
single unexplained event.
- [ ] Staging certificate expires 2028-11-18 and nothing renews it.
---
## 4. Access model
Confirmed by CIVICVS: no workstation access from the home LAN is required.
Confirmed: no workstation access from the home LAN is required.
```
internet -> WireGuard -> srv-b -> containers
```
`srv-b` is the bastion. The authorized key is `root@srv-b`, which is
consistent: root on the host can `pct enter` regardless, so SSH to the
containers adds no privilege. It does mean **every path to a container runs
through `srv-b`** — deliberate, not a limitation.
`srv-b` is the bastion, confirmed by the authorized key matching
`/root/.ssh/id_rsa.pub` on the host. Root on the host can `pct enter` regardless,
so SSH adds no privilege — but **every path to a container runs through
`srv-b`**, which is deliberate.
If direct WireGuard-side access to the catalogue is wanted later, that is a
route addition for `10.20.0.0/24` on the hub. Not now.
Direct WireGuard-side access to the catalogue would be a route addition for
`10.20.0.0/24` on the hub. Not now.
No DHCP anywhere. Two containers with fixed addresses on a portless bridge is
the entire address space; a DHCP server would add a daemon, a lease database
and a failure mode in exchange for nothing.
No DHCP anywhere. Two containers with fixed addresses on a portless bridge is the
entire address space.
---
## 5. Blocked on application code
None of this can be honestly completed while the repository is `LICENSE` and
`README.md`. It is not provisioning work and should not be attempted as such.
Repository state:
- [ ] `requirements-base.txt` / `requirements-cad.txt` and dependency install
- [ ] `mechcomp.service` and `mechcomp-worker.service`
```
HEAD e85c4f4e5bab9a4f032230c99aa6784ace4c80e2
top level .git LICENSE README.md venv
git status ?? venv/
absent requirements-base.txt, requirements-cad.txt, service units
```
None of the following can be honestly completed, and none may be fabricated by
provisioning:
- [ ] Dependency install from committed manifests
- [ ] `mechcomp.service`, `mechcomp-worker.service`
- [ ] Reference toolchain image `mechcomp/reference-toolchain:8.0.0`
- [ ] Fixture reproduction against `ddd0f154…`
- [ ] Fixture reproduction against `ddd0f154...`
- [ ] Application-runtime acceptance
- [ ] Replacing the placeholder with the real service
The Shapely port is the architect's work item and gates all of the above.
**Architect decision:** `venv/` belongs in `.gitignore` — it is a legitimate
artifact inside `install_dir` per YunoHost convention, it simply should not be
tracked. Applied with the first application commit.
The Shapely port gates all of the above.
---
@@ -279,11 +341,11 @@ The Shapely port is the architect's work item and gates all of the above.
| # | Question | Blocks |
|---|---|---|
| 1 | `wg-pk` relay address; unauthenticated from `10.110.0.12`? | mail, `smartd`, backup alerts |
| 2 | Where is the 3+ TB USB disk attached? | gold redundancy step |
| 3 | First `vzdump` archive size | final retention value |
| 1 | Why does delivery fail downstream of `wg-pk`? | mail, `smartd`, backup alerting |
| 2 | What is the backup strategy? | all backup work |
| 3 | Where is the 3+ TB USB disk attached? | gold redundancy step |
Question 3 is answered by doing. Questions 1 and 2 need CIVICVS.
Question 1 is answered by diagnosis. Questions 2 and 3 need CIVICVS.
---
@@ -291,9 +353,11 @@ Question 3 is answered by doing. Questions 1 and 2 need CIVICVS.
| Question | Answer |
|---|---|
| Kane County CA issuance path | None on `srv-b`. Local staging CA generated instead. |
| Kane County CA issuance path | None on `srv-b`. Local staging CA generated. |
| Relay address and authentication | `10.110.0.1:25` over `wg0`, unauthenticated from `10.110.0.12` accepted. |
| `openssh-server` present | Yes, both containers. |
| `ADMIN_USER` / key | `sandor`; `root@srv-b` key, bastion pattern. |
| DHCP pool on `10.0.0.0/24` | **Not applicable.** Containers are no longer on that network. |
| `ADMIN_USER` / key | `sandor`; `root@srv-b` key, bastion pattern confirmed. |
| DHCP pool on `10.0.0.0/24` | **Not applicable.** Containers are not on that network. |
| Docker storage driver | `overlay2` / `systemd`. No fallback needed. |
| LAN workstation access | Not required. WireGuard through `srv-b`. |
| `MECHCOMP_BIND` semantics | Scalar, service address only. Loopback not bound. |