diff --git a/docs/ENVIRONMENT.md b/docs/ENVIRONMENT.md index c317520..b5e8165 100644 --- a/docs/ENVIRONMENT.md +++ b/docs/ENVIRONMENT.md @@ -1,64 +1,59 @@ # ENVIRONMENT.md -Provisioning specification for the **Mechanical Compiler** development and -staging environment. +Specification for a Mechanical Compiler instance. | | | |---|---| -| Revision | 4 (2026-08-15) | -| Supersedes | Revisions 1, 2 and 3 | -| Basis | `ENVIRONMENT_INVESTIGATION_REPORT_2026-08-15.md` (host `srv-b`) | -| Audience | The assistant or operator provisioning the containers | -| Author role | Written by the developer who will work inside this environment | -| Status | All decisions taken. Ready to execute. | -| Repository | `https://gitea.barternetwork.us/TheRON/mechanical-compiler` | -| Licence | AGPL-3.0-or-later (committed at `e85c4f4e`) | +| Revision | 5 (2026-08-16) | +| Supersedes | Revisions 1 through 4 | +| Basis | Revision 4, reconciled against the proven staging build on `srv-b` | +| Scope | **Host-agnostic.** Applies to any instance. | +| Instance state | `STAGING-STATE.md`, and later `PRODUCTION-STATE.md` | +| Failure evidence | `FAILURES.md` | +| Project sequence | `ROADMAP.md` | --- ## 0. How to use this document -This specifies a **staging environment** on `srv-b`, plus the promotion path to -a production host that does not yet exist. +This describes what an instance must **be**. It does not describe what any +particular host currently **is** — that lives in the corresponding state file. -Everything is scripted. Nothing is provisioned by hand. +Revision 4 mixed the two, which is why it aged badly the moment a real host +disagreed with it. Everything specific to `srv-b` has been removed. + +### Authority order + +1. **The instance state file** — factual, wins on any question of what is true. +2. **`FAILURES.md`** — evidence from contact with hosts. Read this *before* + writing automation, not after. +3. **This document** — the specification, corrected whenever proven facts + invalidate it. +4. **`ROADMAP.md`** — product sequence. Environment work never invents + application code. ### Markers | Marker | Meaning | |---|---| | **REQ** | Required. The work is blocked or wrong without it. | -| **PREF** | Preferred. Substitute freely, but report what you substituted. | -| **ASSUMED** | Decided by the author without confirmation. Defensible, but a guess. All are listed in §16. | -| **FILL** | A site value that must be supplied. Goes in `deploy/site.env` and nowhere else. | -| **VERIFY** | Must be checked at run time rather than trusted from this document. | -| **DISCOVER** | Genuinely unknown. Answer by doing, then report. | +| **PREF** | Preferred. Substitute freely, but record the substitution. | +| **PROVEN** | Established by a real build. Do not re-derive; see the cited failure. | +| **ASSUMED** | Decided without confirmation. Listed in §16. | +| **DEFERRED** | Deliberately postponed. Not a defect, not a gap. | +| **INSTANCE** | A value supplied per instance, not fixed here. | -### If you are the provisioning assistant +### Provisioning method -Four requests. **FILL** values live in one file and are referenced from there — -no literals scattered through scripts. **VERIFY** steps run inside the scripts, -not as a pre-flight a human might skip. Report **DISCOVER** outcomes verbatim, -including failures. And if a **REQ** item cannot be satisfied, say so rather -than working around it: the stated reason usually matters more than the -mechanism. +**REQ — manual first.** One command group at a time. Failures are recorded +before they are corrected. The smallest corrective experiment is preferred over +the one that also fixes three adjacent worries. -### What changed in revision 4 - -| # | Change | Source | -|---|---|---| -| 1 | Host is `srv-b`, not `annales` | CHG-001 | -| 2 | Staging and production are separate roles, on separate hosts | CHG-002 | -| 3 | `srv-b` is a standalone node; promotion is export/restore, not cluster migration | CHG-003 | -| 4 | Template `12.12-1`, not `12.7-1` | CHG-004 | -| 5 | Three-tier backup architecture; 32 GB removable gold media | CHG-005 | -| 6 | Backup schedule defined explicitly; none exists to inherit | CHG-006 | -| 7 | `CT_WG_ALIAS` removed | CHG-007 | -| 8 | `PROXY_HOST` not inherited; staging gets its own proxy container | CHG-008 | -| 9 | **Two containers**, not one: application and reverse proxy | this revision | -| 10 | **Second bridge** `vmbr1` isolates service traffic from management traffic | this revision | -| 11 | Container sizing raised substantially to use the host | this revision | -| 12 | Staging is internal-only; the public FQDN is reserved for production | this revision | +Automation is **derived from a proven manual procedure**, never written ahead of +one. Revision 4 said the opposite. `FAILURES.md` is why it changed: three of +this document's assumptions were wrong in ways only a real host revealed, and +one of them (F-003) produced a container that booted, reported success, and had +four broken services. --- @@ -67,104 +62,71 @@ mechanism. ### 1.1 The 2D path and the 3D path are separable The catalogue's product is a 2D cross-section rendered to SVG. Shapely plus our -own code covers that completely: region algebra, offsets, distance queries, -output. CadQuery/OCCT is needed only for STEP and mesh export. +own code covers that completely. CadQuery/OCCT is needed only for STEP and mesh +export. -These stay separate — separate requirements files, separate code paths, and a -CI job that runs the whole suite with the CAD dependencies absent. Three -reasons, none of which is packaging politics: +They stay separate — separate requirements files, separate code paths, and a CI +job running the suite with the CAD dependencies absent. Three reasons, none of +which is packaging politics: 1. Part II §16 requires the control-plane / execution-plane separation - regardless of how the app is distributed. + regardless of distribution. 2. A slow OCCT import must never land in the request path for a page that only draws a cross-section. -3. Test-suite speed. Shapely-only tests run in seconds, which is the difference - between running the 123 frozen fixtures on every save and only before commit. +3. Test-suite speed — the difference between running 123 fixtures on every save + and only before commit. -**Retained correction.** Revision 1 also justified this as a defence against -YunoHost's "resource-hungry" criterion. That was overcalibrated — -`fab-manager_ynh` is in the catalog, `paperless-ngx_ynh` declares nine apt -dependencies including postgresql and redis, and the criterion reads -"resource-hungry **compared to their features**." **CadQuery is a -default-installed dependency.** The boundary stays; the apology goes. +**CadQuery is a default-installed dependency.** Revision 1 restricted it as a +defence against YunoHost's "resource-hungry" criterion; that was overcalibrated. +The criterion reads "compared to their features" and is aimed at marginal apps. +The boundary stays; the apology is gone. -### 1.2 YunoHost and Docker are parallel targets, not sequential ones +### 1.2 YunoHost and Docker are parallel targets, not sequential -YunoHost apps install natively — apt, a venv, systemd, nginx — and the project -does not want Docker inside YunoHost. Docker is a separate distribution -channel, not a stepping stone to the catalog. Neither blocks the other. +YunoHost apps install natively — apt, venv, systemd, nginx — and the project +does not want Docker inside YunoHost. Neither blocks the other, and neither +should be built "in order to" reach the other. -### 1.3 Staging mirrors production topology, not production exposure +### 1.3 Every instance mirrors production topology, not production exposure -`srv-b` runs two containers because production will: an application container -that never terminates TLS, and a reverse proxy that does. A single container -serving directly would never exercise the `X-Forwarded-Proto` path — precisely -the class of bug that cost us time on Hubzilla's session logout. +An instance runs **two containers**: an application container that never +terminates TLS, and a reverse proxy that does. A single container serving +directly would never exercise the `X-Forwarded-Proto` path. -Staging is **not** publicly reachable and does **not** use the public FQDN. It -uses the Kane County Civic Infrastructure CA on an internal name. The public -name and Let's Encrypt belong to production, which forces the promotion path to -be exercised rather than assumed. +Staging is not publicly reachable and does not use the public FQDN. That forces +the promotion path to be exercised rather than assumed. -### 1.4 Directory layout mirrors YunoHost conventions from the start +### 1.4 Directory layout mirrors YunoHost conventions -The paths in §7 are what a YunoHost package would provision anyway -(`install_dir`, `data_dir`, a dedicated system user, a dynamically assigned -port). The eventual `manifest.toml` resource block then describes what already -exists. +`install_dir`, `data_dir`, a dedicated system user, a port treated as data. The +eventual `manifest.toml` resource block then describes what already exists. -### 1.5 All configuration comes from the environment, never from the filesystem +### 1.5 Configuration comes from the environment, never the filesystem -No path, port, URL, or secret may be hardcoded or discovered by convention. One -env file per container; the same variables become `ENV` in a Dockerfile and -`ynh_add_config` substitutions in a package. +No path, port, URL or secret is hardcoded or discovered by convention. One env +file; the same variables become `ENV` in a Dockerfile and `ynh_add_config` +substitutions in a package. -The check that this holds: moving the service to a different hostname is a -one-variable change. It currently is — which is why the staging/production -split below costs almost nothing. +The check: moving an instance to a different hostname is a one-variable change. --- -## 2. Host baseline (confirmed, not to be re-verified) +## 2. Host requirements -| Item | Value | -|---|---| -| Hostname / FQDN | `srv-b` / `srv-b.dev.infra` | -| Role | Development and staging. Standalone node, never clustered. | -| Platform | Proxmox VE 8.4.0 on Debian 12, kernel `6.8.12-9-pve` | -| Hardware | HP ProLiant DL360 G7, firmware P68 | -| CPU | 2 × Xeon X5650, 24 logical CPUs | -| RAM / swap | 31 GiB / 8 GiB | -| Storage | HP Smart Array P410i, 4 × EG0146FAWHU, RAID 1+0, all SMART OK | -| `local` | directory, `/var/lib/vz`, ~70 GiB free — iso, vztmpl, **backup** | -| `local-lvm` | LVM-thin `pve/data`, 166.9 GiB, 0% used — rootdir, images | -| LAN | `vmbr0` on `enp3s0f0`, `10.0.0.12/24`, gateway `10.0.0.1` | -| WireGuard | `wg0` `10.110.0.12/32`, peer `wg-pk.civicus.us:51820`, allowed `10.110.0.0/22` | -| Routing | forwarding on, `MASQUERADE 10.0.0.0/24 -o wg0`, both persistent | -| Resolver | `75.75.75.75`, search `dev.infra` | -| Egress | Direct. No proxy, no allowlist. All required hosts reachable. | -| PVE firewall | `disabled/running` | -| Guests | None. Next ID `100`. | -| Template | `local:vztmpl/debian-12-standard_12.12-1_amd64.tar.zst` | +**REQ** — Proxmox VE 8.x, Debian 12 base. -### 2.1 Two hardware facts worth knowing before anyone is surprised +**REQ** — Two bridges (§3). **REQ** — Directory storage for backups and +templates; LVM-thin or equivalent for container volumes. -The X5650 is Westmere: **no AVX**, and single-thread performance is modest. -Geometry solvers are single-threaded, so regenerating the fixture oracle will -take noticeably longer here than on a modern laptop. That is expected, not a -fault. All compiled wheels target an SSE2 baseline, so nothing breaks. +**PREF** — At least 12 CPU threads and 24 GiB RAM if the instance is also a +development host. The geometry solvers are single-threaded, so single-thread +performance governs fixture regeneration while thread count governs test +parallelism. An older CPU without AVX is acceptable; all compiled wheels target +an SSE2 baseline. -Against that, **24 threads is a great deal of parallel test capacity**. The -test suite runs under `pytest-xdist` with `-n auto`, which is where the box -earns its keep. +**REQ** — Time zone and NTP configured; clock synchronised. -### 2.2 The WireGuard path is outbound only - -`MASQUERADE` means containers reach `10.110.0.0/22`, but nothing on the -WireGuard network can initiate a connection back. Irrelevant for staging. -**Relevant at promotion:** if the production reverse proxy lives on the -WireGuard side, this topology will not serve it unchanged. Recorded now so it -is not rediscovered later. +**INSTANCE** — Hostname, addressing, storage pool names, template path. --- @@ -172,183 +134,185 @@ is not rediscovered later. ### 3.1 Two bridges -**REQ** — `vmbr0`, existing, unchanged. `enp3s0f0`, `10.0.0.12/24`. This is the -**management and egress** network: SSH, apt, PyPI, Gitea, Docker Hub. +**REQ** — A **management bridge** with a physical port, carrying the host's own +address. The host uses it for egress and administration. -**REQ** — `vmbr1`, new, **no physical port**. Host address `10.20.0.1/24`. -This is the **service** network: proxy-to-application traffic and nothing else. +**REQ** — A **service bridge with no physical port**, carrying the host at +`.1`. This is the only network the containers sit on. -A portless bridge needs no cable, no switch configuration, and no spare NIC. It -exists so the application container has no listener on the LAN at all: the app -binds `127.0.0.1` and its `vmbr1` address only, and the only route to it is -through the proxy. +A portless bridge needs no cable, no switch configuration and no spare NIC. It +exists so the containers cannot be reached from the physical network at all. + +### 3.2 Containers have exactly one interface + +**PROVEN (F-017)** — Each container has a single interface on the **service +bridge only**, with the host as gateway. No management-network interface. + +Revision 4 gave each container an address on both bridges, on the unstated +assumption that a workstation on the physical LAN would browse the instance. +That assumption was wrong and put both containers on the operator's home +network. + +**PROVEN (F-021)** — After any interface change, the guest's +`/etc/network/interfaces` must be inspected. `pct set --delete netN` removes the +LXC interface but leaves the stanza in the guest, and `ifup -a` then fails at +boot on a device that no longer exists. See §15 gate 1 for the assertion that +catches this. + +### 3.3 Isolation is routing, not interfaces + +**PROVEN (F-018)** — Removing an interface removes an address, not a route. With +forwarding enabled and a masquerade for internet access, containers reach the +physical LAN through the host, translated to its address. + +**REQ** — These rules, in this order: ``` -# /etc/network/interfaces — append -auto vmbr1 -iface vmbr1 inet static - address 10.20.0.1/24 - bridge-ports none - bridge-stp off - bridge-fd 0 +nat POSTROUTING + -s -d -j RETURN # before any masquerade + -s -o -j MASQUERADE + -s -o -j MASQUERADE + +filter FORWARD + -s -d -j ACCEPT + -s -d -j DROP ``` -### 3.2 The other three NICs stay unconfigured +**REQ** — Rule order is load-bearing. `RETURN` must precede both masquerades. +Rules must be persisted, and the persisted set verified to match the live set. -**REQ** — Do not bond, trunk, or bridge `enp3s0f1` or the second dual-port -adapter. Leave them down and cabled to nothing. +**REQ** — The isolation must be asserted **negatively**: the LAN gateway is +unreachable from both containers. Asserting that the interface is gone is not +sufficient — that is exactly the mistake F-018 records. -This is deliberate. An active-backup bond on `vmbr0` would be defensible, but -it buys link redundancy on a staging host that is allowed to be down, at the -cost of a configuration that has to be understood by everyone who touches the -box afterwards. Not worth it here. - -What the spare ports are *for*, when a reason appears: a point-to-point link to -the production host for promotion transfers, and a dedicated management -interface if the LAN ever becomes untrusted. Neither is now. - -### 3.3 Addressing - -**ASSUMED** — Static, outside any DHCP pool. - -| Guest | `vmbr0` (management) | `vmbr1` (service) | -|---|---|---| -| `srv-b` host | `10.0.0.12/24` | `10.20.0.1/24` | -| CT 100 `mechcomp` | `10.0.0.20/24` | `10.20.0.10/24` | -| CT 101 `mcproxy` | `10.0.0.21/24` | `10.20.0.11/24` | - -**VERIFY** — Stage 2 must confirm `10.0.0.20` and `10.0.0.21` are unclaimed -immediately before creating each container, and abort if not: - -```bash -for ip in 10.0.0.20 10.0.0.21; do - if arping -D -I vmbr0 -c 3 -w 3 "$ip" >/dev/null 2>&1; then - echo " $ip free" - else - echo "ABORT: $ip answered ARP — address is in use" >&2; exit 1 - fi -done -``` - -The investigation deliberately did not scan the LAN, and this document does not -either. An ARP probe for two specific addresses at creation time is the correct -scope. `10.0.0.160` was seen as a neighbour and is avoided. - -**DISCOVER** — If the LAN has a DHCP server, report its pool range so these two -addresses can be confirmed outside it. If `arping` reports a conflict, stop and -ask rather than picking the next number. +**Expected and correct:** containers *can* reach the host itself. A container +addressing its own gateway takes the INPUT path and never enters FORWARD. This +is required — the host is their router and their SSH entry point — and is not a +leak. ### 3.4 Names -**REQ** — `dev.infra` has no resolver: the host points at `75.75.75.75`, which -cannot answer for it. Names resolve through `/etc/hosts` only. +**REQ** — Names resolve through `/etc/hosts` on the host and both containers. -Stage 1 writes to `/etc/hosts` on the host, and stages 3 and 4 write the same -block into both containers: +**REQ — do not stand up an internal DNS server.** Two containers and one host do +not justify one, and it would create a second source of truth for names that +production will not share. + +**REQ** — The service FQDN names **the proxy**, which has exactly one address. +Guest hostnames are separate entries. ``` -10.20.0.10 mechanical-compiler.dev.infra mechcomp -10.20.0.11 mcproxy.dev.infra mcproxy + -> + -> + -> ``` -**REQ — do not stand up an internal DNS server.** Two containers and one host -do not justify one, and adding a resolver here would create a second source of -truth for names that production will not share. +**PROVEN (F-020)** — This value tracked topology through three states before +settling. The reasoning matters more than the value. -### 3.5 Firewall +### 3.5 No DHCP -**ASSUMED** — The Proxmox firewall stays disabled, matching the observed -staging baseline. The service network is portless and the application has no -LAN listener, which is the containment that matters here. +**REQ** — Static addressing. Two containers on a portless bridge is the entire +address space; a DHCP server would add a daemon, a lease database and a failure +mode in exchange for nothing. -**REQ** — Production must revisit this. Recorded as a promotion checklist item -in §14, not as a permanent decision. +### 3.6 Firewall + +**ASSUMED** — The Proxmox firewall may remain disabled where the service bridge +is portless and the application has no listener outside it. **REQ** — Production +must revisit this; see §15 gate 5. --- ## 4. Containers -### 4.1 Roles +### 4.1 Roles and features -| | CT 100 | CT 101 | +| | Application | Proxy | |---|---|---| -| Hostname | `mechcomp` | `mcproxy` | -| Role | Application and worker | Reverse proxy, TLS termination | +| Role | application, worker | reverse proxy, TLS | | Unprivileged | yes | yes | -| Features | `nesting=1,keyctl=1` | none | +| Features | `nesting=1,keyctl=1` | `nesting=1` | -**VERIFY** — Confirm both IDs with `pvesh get /cluster/nextid` and `pct list` -immediately before creation. `100` and `101` are candidates from a host that -had no guests, not reservations. +**PROVEN (F-003)** — **Every Debian 12 unprivileged container requires +`nesting=1`**, including one running nothing but nginx. Without it, +`systemd-logind`, `systemd-networkd`, `systemd-timedated` and +`systemd-networkd.socket` fail with `226/NAMESPACE`. Revision 4 said the proxy +needed no features. + +`keyctl=1` is required only where Docker runs. Tested separately rather than +copied across; it is genuinely unnecessary on the proxy. ### 4.2 Sizing -The host has 24 threads and 31 GiB with nothing else on it. Sizing reflects -that rather than a cautious default. +**INSTANCE.** Sizing is a host capacity decision, not a specification value. -| | CT 100 | CT 101 | Host remaining | -|---|---|---|---| -| vCPU | 16 | 2 | 24 total, soft limits | -| RAM | 16 GiB | 1 GiB | ~14 GiB | -| Swap | 4 GiB | 512 MiB | — | -| rootfs | 40 GiB | 8 GiB | on `local-lvm` | -| `mp0` | 60 GiB at `/var/lib/mechcomp` | — | — | +**REQ** — The application container's data volume is a **real mount point** with +`backup=0`. Backup of application data is a separate tier (§10), and including a +large data volume in a container snapshot is how a backup store fills. -Thin provisioning means `mp0` consumes only what is written. Committed total is -108 GiB of 166.9 GiB — generous but not over-committed, which matters because -thin-pool exhaustion is an ugly failure mode. - -**REQ** — `mp0` is a real Proxmox mount point with `backup=0`. See §9.2 for why -the backup flag is off, and §12 for the check that proves the mount is real. +**REQ** — Do not over-commit a thin pool. Thin-pool exhaustion is an ugly +failure mode. ### 4.3 Base OS -**REQ** — Debian 12 bookworm from -`local:vztmpl/debian-12-standard_12.12-1_amd64.tar.zst`, both containers. +**REQ** — Debian 12 bookworm, both containers. It matches YunoHost 12 stable and +ships Python 3.11, which has the most reliable OCP/CadQuery wheel coverage. -Chosen over Debian 13 and Ubuntu because it matches YunoHost 12 stable, ships -Python 3.11 (the best OCP/CadQuery wheel coverage), and is already present on -the host. +**REQ (F-005)** — `apt full-upgrade` immediately after first boot, before +anything else is installed. Templates are not current, and discovering a +52-package backlog midway through provisioning is avoidable. -**VERIFY** — Stage 2 confirms the template is present, and downloads it only if -absent: +**REQ (F-004)** — Locale must be **generated**, not merely configured. Install +`locales`, enable the locale, run `locale-gen`, then `update-locale`. Setting +`LANG` alone leaves the system with only `C` and `POSIX`, silently, until +something depends on collation. -```bash -pveam list local | grep -q "${PVE_TEMPLATE##*/}" || pveam download local "${PVE_TEMPLATE##*/}" -``` - -**REQ** — `America/Chicago`, `en_US.UTF-8` in both containers, matching the host. +**REQ** — Time zone matching the host. --- -## 5. TLS and service names +## 5. Names and TLS -### 5.1 Staging +### 5.1 Three names, never conflated -**REQ** — Internal name `mechanical-compiler.dev.infra`, certificate issued by -the **Kane County Civic Infrastructure CA**. +| Name | Scope | +|---|---| +| Service FQDN | What users type. Differs per instance. | +| Infrastructure token | System user, paths, units, YunoHost app id. **Same everywhere.** | +| Project name | Documents, README, catalog entry | -**REQ** — The CA root is installed into both containers' trust stores and must -also be trusted on the workstation used to browse staging. +YunoHost app ids accept lowercase alphanumerics and underscores only, so a +hyphenated FQDN cannot serve as the token. -**REQ — the public FQDN is not used on staging.** No DNS record for -`mechanical-compiler.manufacturing.kane-il.us` may point at `srv-b`. No Let's -Encrypt certificate is issued for staging: it consumes rate limit against a -name that must be clean when production needs it, and it would put a -development box on the public internet. +### 5.2 Non-production instances -### 5.2 Production, reserved +**REQ** — An internal service FQDN. **REQ** — No public DNS record and no +publicly-trusted certificate: it consumes rate limit against a name production +needs clean, and puts a development host on the internet. + +**PROVEN** — A locally generated CA is sufficient and is the fallback whenever a +site CA has no discoverable issuance path. The point of the TLS step is +exercising termination and header forwarding, not the trust chain. Replacing the +leaf later is two file copies and a reload. + +**REQ** — The CA root is installed into the trust store of the host, both +containers, and any client used to browse the instance. + +**REQ** — Record the leaf's expiry somewhere that gets read. Nothing renews a +hand-made certificate. + +### 5.3 Production | | Value | |---|---| | FQDN | `mechanical-compiler.manufacturing.kane-il.us` | | Zone | Labels within `kane-il.us`. **Not delegated.** No new DS record. | | Certificate | Let's Encrypt, **per-name, no wildcard** | -| Challenge | **ASSUMED** HTTP-01 on the production proxy. Match existing practice if it is DNS-01; do not run both. | +| Challenge | **ASSUMED** HTTP-01. Match existing practice if it is DNS-01; never run both. | -`manufacturing.kane-il.us` is a namespace, not a service. Nothing is served -from it. `kane-il.us` itself stays reserved for cross-cutting infrastructure — -proxy, mail, directory, certificates — and must not acquire project records. +`manufacturing.kane-il.us` is a namespace, not a service. `kane-il.us` stays +reserved for cross-cutting infrastructure and must not acquire project records. The leaf names the service, not the discipline, because the parent already carries the discipline and the namespace must hold siblings: @@ -359,38 +323,34 @@ field-metrology.manufacturing.kane-il.us Utility One, later stock-processing.manufacturing.kane-il.us Utility Two, later ``` -### 5.3 Three names, never conflated - -| Name | Scope | Value | -|---|---|---| -| Service FQDN | What users type | staging and production differ; see above | -| Infrastructure token | System user, paths, units, YunoHost app id | `mechcomp` | -| Project name | Documents, README, catalog entry | Mechanical Compiler | - -The token stays `mechcomp` in both environments. YunoHost app ids accept -lowercase alphanumerics and underscores only, so a hyphenated FQDN could not -serve as one anyway. - --- -## 6. Reverse proxy (CT 101) +## 6. Reverse proxy -**REQ** — NGINX. Terminates TLS on `vmbr0`, proxies to CT 100 over `vmbr1`. +**REQ** — nginx in the proxy container. Terminates TLS, proxies to the +application container's service address. -**REQ** — The `X-Forwarded-Proto` header is mandatory, not optional. Getting -this wrong is what broke Hubzilla sessions. +**REQ** — Debian's `default` site removed. Only the project vhost enabled. + +**REQ** — `X-Forwarded-Proto` is mandatory. Getting this wrong is what broke +Hubzilla sessions. + +**REQ** — `default_server` declared explicitly on both TLS listeners. With one +vhost the implicit default is correct today; the moment a second is added, which +block catches unmatched requests becomes a function of file ordering. This is +independent hardening and **does not close F-012**. ```nginx server { - listen 443 ssl; - listen [::]:443 ssl; - server_name mechanical-compiler.dev.infra; + listen 443 ssl default_server; + listen [::]:443 ssl default_server; + server_name ; - ssl_certificate /etc/ssl/mechcomp/server.crt; # Kane County CA - ssl_certificate_key /etc/ssl/mechcomp/server.key; + ssl_certificate ; + ssl_certificate_key ; location / { - proxy_pass http://10.20.0.10:8770; + proxy_pass http://:; proxy_http_version 1.1; proxy_set_header Host $host; proxy_set_header X-Real-IP $remote_addr; @@ -404,40 +364,43 @@ server { server { listen 80; listen [::]:80; - server_name mechanical-compiler.dev.infra; + server_name ; return 301 https://$host$request_uri; } ``` -**PREF** — `proxy_read_timeout 300s`. Compile jobs are asynchronous by design, -but the initial synchronous render can take tens of seconds on a cold cache on -this CPU, and the 60s default leaves no headroom. +**REQ** — This vhost differs between instances in exactly two places: +`server_name` and the certificate paths. Any further divergence means the +non-production instance has stopped testing production. -**REQ** — This vhost is the production vhost with two values changed: -`server_name` and the certificate paths. Keep it that way. Any divergence -beyond those two lines means staging has stopped testing production. +**PREF** — `proxy_read_timeout 300s`. Compile jobs are asynchronous, but a cold +synchronous render can take tens of seconds and the 60s default leaves no +headroom. --- -## 7. Users, directories, permissions (CT 100) +## 7. Users, directories, permissions -**REQ** — Dedicated system user, no login shell, no home of its own: +**REQ** — Dedicated system user, no login shell: ``` user:group mechcomp:mechcomp shell /usr/sbin/nologin +home ``` +**REQ (F-016 context)** — The service user's home is the install directory, not +a nonexistent `/home/`. Tooling run as that user writes to `$HOME`; +`ProtectHome=true` makes it moot at runtime but not during provisioning. + **REQ** — Layout, matching what a YunoHost package would provision: -| Path | Owner | Mode | Purpose | YunoHost equivalent | -|---|---|---|---|---| -| `/var/www/mechcomp` | `mechcomp:mechcomp` | `0750` | source and venv | `install_dir` | -| `/var/lib/mechcomp` | `mechcomp:mechcomp` | `0750` | data, on `mp0` | `data_dir` | -| `/var/log/mechcomp` | `mechcomp:mechcomp` | `0750` | logs | — | -| `/etc/mechcomp/mechcomp.env` | `root:mechcomp` | `0640` | configuration | `ynh_add_config` target | - -**REQ** — `/var/lib/mechcomp` subdivides as: +| Path | Owner | Mode | YunoHost equivalent | +|---|---|---|---| +| `/var/www/mechcomp` | `mechcomp:mechcomp` | `0750` | `install_dir` | +| `/var/lib/mechcomp` | `mechcomp:mechcomp` | `0750` | `data_dir` | +| `/var/log/mechcomp` | `mechcomp:mechcomp` | `0750` | — | +| `/etc/mechcomp/mechcomp.env` | `root:mechcomp` | `0640` | `ynh_add_config` | ``` /var/lib/mechcomp/ @@ -448,17 +411,27 @@ shell /usr/sbin/nologin ``` The `cache/` guarantee matters: anything not reproducible from `db/` plus -`artifacts/` must not live there. That is what keeps the backup story simple. +`artifacts/` must not live there. -**REQ** — `${ADMIN_USER}` gets an ordinary shell account in the `mechcomp` -group with `${ADMIN_SSH_KEY}` installed, in **both** containers. The service -user is never used interactively. +**REQ (F-007)** — **Never `chown -R` a mount point root.** `data_dir` is a +filesystem root containing `lost+found`, owned by `nobody:nogroup` at `0700`, +created by `mkfs` and not ours to manage. Enumerate the application directories +explicitly. + +**REQ (F-008)** — All repository operations run as the owning service user. Do +not add a root `safe.directory` exception — it would mask every future instance +of the same mistake. + +**REQ** — The admin account exists in both containers with an authorized key. +**REQ (F-014)** — If the admin is placed in the service group, that group must be +created explicitly in the proxy container, where no service user creates it as a +side effect. --- -## 8. Software (CT 100) +## 8. Software -### 8.1 System packages +### 8.1 Application container **REQ** @@ -466,17 +439,16 @@ user is never used interactively. python3 python3-venv python3-dev build-essential pkg-config git curl ca-certificates -libgeos-dev # Shapely bundles GEOS in its wheel; headers make a - # source build possible if a wheel is unavailable +libgeos-dev sqlite3 ``` -**PREF** — `jq ripgrep tmux htop` +**REQ (F-015)** — The proxy container needs its own minimal toolset, at least +`curl` and `ca-certificates`. Do not assume packages installed in one container +exist in the other. **REQ — do not install** `openscad`, `openscad-nightly`, or any Qt or X11 -package. This is not a weight argument: the running application never invokes -OpenSCAD, and §8.2 confines it where it belongs. If OpenSCAD appears in the -container's package list, something has gone wrong architecturally. +package. The running application never invokes OpenSCAD, and §8.2 confines it. ### 8.2 The pinned reference toolchain @@ -484,7 +456,6 @@ container's package list, something has gone wrong architecturally. frozen against, and nothing else: ```dockerfile -# tools/reference-toolchain/Dockerfile FROM debian:12-slim RUN apt-get update && apt-get install -y --no-install-recommends \ openscad git ca-certificates \ @@ -494,60 +465,41 @@ RUN git clone https://github.com/BelfrySCAD/BOSL2.git /BOSL2 \ ``` Tag `mechcomp/reference-toolchain:8.0.0`. Its only job is to run -`make_fixtures.py` and prove the oracle still reproduces. This is why the -Debian revision skew does not matter: the container has no OpenSCAD, so its -version cannot drift. +`make_fixtures.py` and prove the oracle reproduces. The container has no +OpenSCAD, so its version cannot drift. -**Acceptance test** — must reproduce exactly: +**Acceptance** — must reproduce exactly: ``` -sha256 of strap-beam-fixtures-8.0.0.json +sha256 strap-beam-fixtures-8.0.0.json = ddd0f1548379205dd0c652ec07285b0dae331e52ff0a0437005dfc6cddcc2cb2 ``` -If it does not, stop and report. A drifting oracle is worse than no oracle. - -**DISCOVER** — Docker in an unprivileged LXC is usually fine on this kernel -with `overlay2`, but it is not guaranteed. Report: - -```bash -docker info --format '{{.Driver}} / {{.CgroupDriver}}' -docker run --rm debian:12-slim echo ok -``` - -`fuse-overlayfs` is the documented fallback. If both fail, say so — moving the -image build to a small VM on the same host is a minor change and preferable to -fighting the storage driver. +A drifting oracle is worse than no oracle. ### 8.3 Python -**REQ** — Virtualenv at `/var/www/mechcomp/venv`. Debian 12 enforces PEP 668; -do not use `--break-system-packages` to get around it. +**REQ** — A virtualenv inside `install_dir`. Debian 12 enforces PEP 668; do not +use `--break-system-packages`. + +**REQ** — `venv/` is in `.gitignore`. It is a legitimate artifact in +`install_dir` and simply should not be tracked. **REQ** — Pinned, hash-checked dependencies in two files, both installed: | File | Contents | |---|---| -| `requirements-base.txt` | shapely, fastapi, uvicorn, pydantic, sqlalchemy, jinja2, pytest, **pytest-xdist** | +| `requirements-base.txt` | shapely, fastapi, uvicorn, pydantic, sqlalchemy, jinja2, pytest, pytest-xdist | | `requirements-cad.txt` | cadquery / build123d | -Both install here. The split is not about disk; it is the mechanism that keeps -§1.1 honest. CI runs the suite twice — with both files, then with -`requirements-base.txt` alone — and the second run must pass. - -`pytest-xdist` is in the base set specifically so `-n auto` uses all 24 threads. - -**PREF** — `uv` for resolution and locking; `pip-tools` is a fine substitute. -Either way the lock file is committed. +CI runs the suite twice — with both, then with base alone — and the second run +must pass. That is the mechanism keeping §1.1 honest. --- -## 9. Services and backups +## 9. Services -### 9.1 Systemd units (CT 100) - -**REQ** — Two units, split along the control-plane / execution-plane boundary -Part II §16 requires: +**REQ** — Two units, split along the control-plane / execution-plane boundary: | Unit | Role | |---|---| @@ -578,319 +530,142 @@ RestrictSUIDSGID=true WantedBy=multi-user.target ``` -**REQ** — Worker only: - -```ini -MemoryMax=8G -CPUQuota=1200% -TimeoutStopSec=30 -``` - -A runaway geometry job must not take the container down. `MemoryMax` turning a -hang into a clean OOM-kill of one worker is the desired behaviour. +**REQ** — Worker only: a `MemoryMax` bound, a `CPUQuota`, `TimeoutStopSec=30`. +A runaway geometry job must become a clean OOM-kill of one worker, not a dead +container. **REQ** — Logging to journald. No application-managed rotation. **REQ** — Job queue is SQLite-backed and in-process. No Redis, no Celery, no RabbitMQ. A table with a status column and a claim query is correct at this -scale and adds no services to either packaging target. It sits behind an -interface so it can be swapped if load ever justifies it. It will not. +scale and adds no services to either packaging target. -### 9.2 Three backup tiers +**REQ (F-019)** — **Anything a proof depends on must be supervised.** This +includes temporary scaffolding. A placeholder backend used to prove the proxy +chain before the application exists is a systemd unit, obviously named and +obviously temporary — otherwise the proof expires at the next reboot and the +following session diagnoses a proxy fault that does not exist. -| Tier | What | Where | Retention | -|---|---|---|---| -| Routine | `vzdump` of both container **rootfs** | `local` (`/var/lib/vz/dump`) | see below | -| Application | tar stream of the data set | `/var/lib/vz/mechcomp-app/` on `srv-b` | 14 daily | -| Gold | routine + application + manifest | 32 GB removable USB, LUKS | indefinite, offline | - -**REQ — `mp0` is excluded from `vzdump` (`backup=0`).** `/var/lib/vz` sits on -`pve-root`, which is 78 GiB, and a full root filesystem is how Proxmox hosts -become interesting. Excluding the 60 GiB data volume keeps each routine archive -to the rootfs alone — a few GiB compressed — and application data flows through -the tier below, which is where it belongs anyway. - -**REQ** — `vzdump` job: both containers, `snapshot` mode, `zstd`, daily at -02:30, storage `local`. - -**ASSUMED** — Initial retention `keep-last=3`. Deliberately conservative: no -`vzdump` has been taken on this host, so actual size is unmeasured. **DISCOVER** -— after the first successful run, report the archive size and I will set final -retention. With 70 GiB free, three archives of an unknown size is the safe -starting point. - -**REQ** — A `df` guard in the backup hook that refuses to run and alerts if -`/var/lib/vz` falls below 20 GiB free. - -**REQ** — Application backup is a single script in CT 100 emitting a tar stream -on stdout, pulled host-side: - -```bash -pct exec ${CT_ID_APP} -- /usr/local/bin/mechcomp-backup --stdout \ - > /var/lib/vz/mechcomp-app/mechcomp-$(date +%F).tar -``` - -Covering exactly: - -``` -/etc/mechcomp/mechcomp.env -/var/lib/mechcomp/db/ -/var/lib/mechcomp/artifacts/ -/var/lib/mechcomp/fixtures/ -``` - -`cache/` is excluded. `/var/www/mechcomp` is excluded because it is reproducible -from git plus the lock file. This script is, almost verbatim, the future -YunoHost `backup` script. - -**REQ — pull, never push.** The container has no access to any backup -destination. The alternative is tempting and simpler, but it hands a -network-facing container write access to the last line of defence. It also -avoids idmap ownership problems entirely. - -### 9.3 Gold archive - -**REQ** — 32 GB USB stick, **LUKS2 full-device encryption**, ext4 inside, -filesystem label `MC-GOLD-01`, mounted at `/mnt/gold` with `noauto`. - -Encryption is not optional: the archive contains `mechcomp.env`, which contains -`MECHCOMP_SECRET_KEY`, and a stick that travels between machines is exactly the -thing that gets lost. The passphrase lives in CIVICVS's password manager and -never in the repository. Encrypting the device rather than the files means the -later copy to the 3+ TB disk inherits the protection if the LUKS image is -copied as a block image. - -ext4 rather than exFAT: no 4 GiB file cap, ownership preserved, journalled. -**DISCOVER** — if the stick must ever be read from Windows or macOS, say so and -this becomes exFAT with `age`-encrypted archives instead. - -**REQ** — Archive layout, one directory per mint: - -``` -/mnt/gold/-/ -├── vzdump-lxc-100-*.tar.zst -├── vzdump-lxc-101-*.tar.zst -├── mechcomp-app-*.tar -├── site.env (the FILL values, not secrets) -├── MANIFEST.txt git tag, revisions, host, PVE version -└── SHA256SUMS -``` - -**REQ** — `gold-archive.sh` verifies `SHA256SUMS` after writing and before -unmounting. An archive that has not been verified has not been made. - -**ASSUMED** — A gold archive is minted on **every tagged release** and -**immediately before any promotion or host migration**. This ties the tier to -the existing tag discipline rather than to a calendar. - -**DISCOVER** — Where the 3+ TB USB disk is attached now that `annales` is out -of scope. Until that is known, stick-to-disk copying is a documented manual -step, not a scripted one. - -### 9.4 Disk health monitoring - -**REQ** — Configure `smartd` for the four RAID members using the already -installed `smartmontools`. No third-party HPE repository is added, and `ssacli` -is not installed: the investigation showed the kernel and `smartctl -d cciss,N` -already provide what is needed. - -``` -# /etc/smartd.conf -/dev/sda -d cciss,0 -a -m root -M exec /usr/share/smartmontools/smartd-runner -/dev/sda -d cciss,1 -a -m root -M exec /usr/share/smartmontools/smartd-runner -/dev/sda -d cciss,2 -a -m root -M exec /usr/share/smartmontools/smartd-runner -/dev/sda -d cciss,3 -a -m root -M exec /usr/share/smartmontools/smartd-runner -``` - -These are 2010-vintage spinning disks in a RAID 1+0 holding both the staging -environment and its routine backups. Knowing about a failure before the second -one is worth the four lines. - -**DISCOVER** — Whether Proxmox notifications can relay through the existing -mail infrastructure. If not, `smartd` and `vzdump` failures land in local root -mail only, which nobody reads — report this rather than leaving it silently -broken. +**REQ (F-011)** — Do not use `/tmp` for state shared across users. +`fs.protected_regular` makes that fail in ways that look like permission bugs. +Moot under `PrivateTmp=true`. --- -## 10. Configuration contract +## 10. Backup -**REQ** — `/etc/mechcomp/mechcomp.env` in CT 100, and nothing else, configures -the application: +**DEFERRED — 2026-08-16, operator decision.** The backup strategy may change. +This section is retained as a design sketch and is **not authoritative**; no +acceptance gate depends on it until a strategy is chosen. + +The sketch, for whoever revisits it: + +| Tier | What | Retention | +|---|---|---| +| Routine | container snapshot, **rootfs only** | measured, not guessed | +| Application | tar stream of `db/`, `artifacts/`, `fixtures/`, and the env file | short, rolling | +| Gold | routine + application + manifest, on removable encrypted media | indefinite, offline | + +Three constraints that survive any strategy change: + +1. **Data volumes carry `backup=0`.** A container snapshot store on the root + filesystem is how a Proxmox host fills up. +2. **Pull, never push.** The container has no access to any backup destination. + The alternative hands a network-facing container write access to the last + line of defence, and it avoids idmap ownership problems. +3. **Encrypt removable media.** The backup set includes the env file, which + includes the secret key. Media that travels is media that gets lost. + +**Standing recommendation regardless of strategy:** take one manual snapshot of +both containers before any significant change. It presupposes nothing and yields +the archive size that retention decisions need. + +--- + +## 11. Configuration contract + +**REQ** — One env file configures the application: ```bash -MECHCOMP_ENV=staging # staging | production -MECHCOMP_BIND=10.20.0.10 # service network only; never 0.0.0.0 -MECHCOMP_PORT=8770 -MECHCOMP_BASE_URL=https://mechanical-compiler.dev.infra +MECHCOMP_ENV= # staging | production +MECHCOMP_BIND= # single address — see below +MECHCOMP_PORT= +MECHCOMP_BASE_URL= MECHCOMP_DATA_DIR=/var/lib/mechcomp MECHCOMP_LOG_LEVEL=info MECHCOMP_DB_URL=sqlite:////var/lib/mechcomp/db/mechcomp.sqlite3 -MECHCOMP_WORKER_CONCURRENCY=4 -MECHCOMP_CAD_BACKEND=none # none | cadquery — 'none' until the 3D path exists +MECHCOMP_WORKER_CONCURRENCY= +MECHCOMP_CAD_BACKEND=none # none | cadquery MECHCOMP_ARTIFACT_RETENTION_DAYS=30 -MECHCOMP_SECRET_KEY= # generated by stage 3; see §11.4 +MECHCOMP_SECRET_KEY= # generated on first provision, never committed ``` -**REQ** — The application fails loudly at startup if a required variable is -missing, and never falls back to a compiled-in default path. That single rule -is most of what makes both packaging targets work unchanged. +**REQ** — `MECHCOMP_BIND` is a **single address**. Loopback is not additionally +bound. -**REQ** — `MECHCOMP_BIND` is the service-network address, plus `127.0.0.1`. -Binding `0.0.0.0` would defeat the point of the second bridge. +Revision 4 said "the service-network address, plus `127.0.0.1`," which a scalar +cannot express. Resolved toward one address: `uvicorn` takes one `--host`, and +two sockets would mean either `0.0.0.0` — defeating the isolation — or +multi-socket setup for no benefit. Local checks inside the container can use the +service address. The scalar also generalises correctly: under YunoHost, where +nginx and the application share a host, the same variable takes `127.0.0.1`. + +**REQ** — The application fails loudly at startup if a required variable is +missing, and never falls back to a compiled-in default path. + +**REQ** — The secret is generated on first provision and never overwritten by a +re-run. It is never committed and never logged. Promotion changes `MECHCOMP_ENV`, `MECHCOMP_BIND` and `MECHCOMP_BASE_URL`. Nothing else. --- -## 11. Provisioning scripts +## 12. Mail and monitoring -Everything above is a specification for scripts, not a runbook for a human. +**DEFERRED — end-to-end alerting is not accepted.** See F-023. -### 11.1 Repository +**REQ when resumed** — **SMTP 250 from a relay and an empty local queue prove +handoff, not delivery.** Alerting acceptance requires demonstrated end-to-end +receipt. Until then, mail, disk monitoring and any mail-dependent backup +alerting are unaccepted. -**ASSUMED** — A separate repository, `TheRON/mechanical-compiler-infra`. - -Provisioning scripts do not belong in the application repository. A YunoHost -reviewer eventually reads that tree, and Proxmox tooling beside the app is -exactly the "non-standard directory architecture" the catalog flags. The two -also version independently, which they will. - -If you prefer one repository, put the scripts in `deploy/` and add that path to -the packaging exclusion list from the first commit — not after they have history. - -### 11.2 Four stages - -| Script | Runs as | Where | Does | -|---|---|---|---| -| `stage1-host.sh` | root | `srv-b` | `vmbr1`, `/etc/hosts`, `smartd`, vzdump job, `/var/lib/vz/mechcomp-app` | -| `stage2-create-cts.sh` | root | `srv-b` | template check, ARP verify, `pct create` ×2, `mp0` with `backup=0`, first boot | -| `stage3-app.sh` | root | CT 100 | packages, user, layout, venv, Docker, units, config, secret | -| `stage4-proxy.sh` | root | CT 101 | nginx, CA cert install, vhost, `/etc/hosts` | - -**REQ** — The only interface between stages is `deploy/site.env`. Stage 2 -copies it into both containers; stages 3 and 4 source it. No other state crosses -a boundary. - -**REQ** — Stage 4 installs its own vhost. Unlike revision 3, the proxy is ours, -on this host, and there is no separate operator to hand a config to. - -### 11.3 Idempotency - -**REQ** — All four stages are safe to re-run. Create-if-absent, -converge-if-present, never destroy. A re-run against a healthy host is a no-op -that exits zero. - -This follows from "version, tag, release with discipline": a script you cannot -re-run is a script you cannot test, and a script you cannot test is ceremony, -not discipline. - -**REQ** — Every stage starts with `set -euo pipefail` and validates that each -**FILL** value in `site.env` is non-empty before touching anything. - -### 11.4 Secrets - -**REQ** — `MECHCOMP_SECRET_KEY` is never committed and never appears in -`site.env`. Stage 3 generates it on first run and refuses to overwrite: - -```bash -if ! grep -q '^MECHCOMP_SECRET_KEY=.\+' /etc/mechcomp/mechcomp.env 2>/dev/null; then - key=$(openssl rand -base64 32) - # write it; never log it -fi -``` - -**REQ** — `deploy/site.env` is in `.gitignore`. `deploy/site.env.example`, with -empty values, is committed. The LUKS passphrase is in neither. - -### 11.5 Versioning - -**ASSUMED** — The infrastructure repository carries its own semver, independent -of the application's `SB_REVISION`. They drift for unrelated reasons and tying -them would force meaningless bumps on both. - -**REQ** — Tag a release before each provisioning run and record the tag in the -run log and in the gold `MANIFEST.txt`. "Which version of the script built this -container" must be answerable six months from now. +**REQ** — Disk monitoring must be configured **explicitly per device** where the +controller does not present members to automatic scanning. An active +`smartmontools.service` monitoring zero devices is not monitoring, and is easy +to mistake for success. --- -## 12. `deploy/site.env` +## 13. Instance parameters -```bash -# ---- Proxmox host (confirmed 2026-08-15) -------------------------------- -PVE_HOST=srv-b -PVE_STORAGE=local-lvm -PVE_DATA_STORAGE=local-lvm -PVE_BACKUP_STORAGE=local -PVE_BRIDGE_MGMT=vmbr0 -PVE_BRIDGE_SVC=vmbr1 -PVE_TEMPLATE=local:vztmpl/debian-12-standard_12.12-1_amd64.tar.zst +**INSTANCE** — Every value below is supplied per instance and recorded in that +instance's state file, not here. -# ---- Application container --------------------------------------------- -CT_ID_APP=100 # VERIFY before create -CT_HOSTNAME_APP=mechcomp -CT_IP_APP_MGMT=10.0.0.20/24 # VERIFY free by ARP before create -CT_IP_APP_SVC=10.20.0.10/24 -CT_CORES_APP=16 -CT_RAM_APP=16384 -CT_SWAP_APP=4096 -CT_ROOTFS_APP=40 -CT_MP0_APP=60 - -# ---- Proxy container --------------------------------------------------- -CT_ID_PROXY=101 # VERIFY before create -CT_HOSTNAME_PROXY=mcproxy -CT_IP_PROXY_MGMT=10.0.0.21/24 # VERIFY free by ARP before create -CT_IP_PROXY_SVC=10.20.0.11/24 -CT_CORES_PROXY=2 -CT_RAM_PROXY=1024 -CT_SWAP_PROXY=512 -CT_ROOTFS_PROXY=8 - -# ---- Network ----------------------------------------------------------- -CT_GATEWAY=10.0.0.1 -SVC_NET=10.20.0.0/24 -SVC_HOST_IP=10.20.0.1 - -# ---- Service ----------------------------------------------------------- -MECHCOMP_ENV=staging -SERVICE_FQDN=mechanical-compiler.dev.infra -APP_PORT=8770 -TLS_CA=kane-county-civic-infrastructure - -# ---- Identity ---------------------------------------------------------- -ADMIN_USER= # FILL -ADMIN_SSH_KEY= # FILL - -# ---- Backup ------------------------------------------------------------ -VZDUMP_SCHEDULE="02:30" -VZDUMP_KEEP_LAST=3 # revisit after first size measurement -APP_BACKUP_DIR=/var/lib/vz/mechcomp-app -APP_BACKUP_KEEP=14 -GOLD_LABEL=MC-GOLD-01 -GOLD_MOUNT=/mnt/gold - -# ---- Reserved for production, not used on srv-b ------------------------ -# PRODUCTION_FQDN=mechanical-compiler.manufacturing.kane-il.us +``` +host hostname, storage pools, bridges, template +management network host address, gateway +service network network, host address, container addresses +wireguard interface, address, allowed networks +containers ids, hostnames, sizing +service FQDN per instance +TLS CA and leaf source +admin username, authorized key +mail relay endpoint, root alias ``` -`CT_WG_ALIAS` is gone. The host routes and MASQUERADEs `10.0.0.0/24` out `wg0`, -so containers reach the WireGuard network by their ordinary LAN address, and no -per-container alias mechanism exists to configure. +**REQ** — These live in exactly one file per instance, sourced by every +provisioning step. No literals scattered through scripts. The file is not +committed; an example with empty values is. --- -## 13. Constraints the application code will follow - -Recorded so the environment is not built in a way that quietly permits their -violation. +## 14. Constraints the application code will follow 1. No hardcoded absolute paths. Everything derives from `MECHCOMP_DATA_DIR`. 2. No writes outside `MECHCOMP_DATA_DIR`. Enforced by `ProtectSystem=strict`. 3. No dependency on systemd from application code. Signal handling only. -4. The app listens on a TCP port, on the service network and loopback only. +4. The app listens on one TCP address, never `0.0.0.0`. 5. No OpenSCAD, Qt or X11 dependency in the running application. 6. The 3D backend is reached only through an interface in `src/mechcomp/cad/`. No other module imports CadQuery or OCP directly. @@ -902,162 +677,99 @@ violation. 9. Every generated artifact carries the generator revision that produced it. 10. The test suite passes with `requirements-cad.txt` uninstalled. 11. AGPL-3.0 §13 requires network users be offered the source. The web tier - carries a visible source link to the Gitea repository. This is a licence - obligation, not a nicety. + carries a visible source link to the repository. A licence obligation. --- -## 14. Acceptance checklist +## 15. Acceptance gates -Run and report. This is also the first draft of the YunoHost `tests.toml` -criteria. +**REQ** — Five independent gates. Revision 4 ran them as one list, which meant +an honest infrastructure build appeared to fail because application code did not +exist yet. -```bash -# ---- host ---- -pveversion | head -1 -ip -br address show vmbr1 # expect 10.20.0.1/24 -pct list # expect 100 and 101, both running -grep -c mechanical-compiler.dev.infra /etc/hosts -systemctl is-active smartd -cat /etc/pve/jobs.cfg | grep -c vzdump # expect >= 1 -pct config 100 | grep mp0 # must contain backup=0 +### Gate 1 — Infrastructure -# ---- CT 100 ---- -pct exec 100 -- bash -c ' - cat /etc/debian_version - python3 --version - timedatectl show -p Timezone --value - id mechcomp - findmnt /var/lib/mechcomp # must be a mount, not rootfs - stat -c "%U:%G %a" /etc/mechcomp/mechcomp.env - dpkg -l | grep -Ei "openscad|libqt|xserver" || echo "clean" - ip -br address | grep 10.20.0.10 - ss -lntp | grep 8770 # must NOT bind 0.0.0.0 - grep -c "^MECHCOMP_SECRET_KEY=.\+" /etc/mechcomp/mechcomp.env - /var/www/mechcomp/venv/bin/python -c "import shapely; print(shapely.__version__)" - /var/www/mechcomp/venv/bin/python -c "import cadquery; print(cadquery.__version__)" - docker info --format "{{.Driver}}" - docker run --rm mechcomp/reference-toolchain:8.0.0 openscad --version -' +Host and both containers `running` with **zero failed units**, asserted **after a +reboot**. Locale generated. Data volume is a real mount. Ownership correct with +`lost+found` untouched. Forbidden packages absent. Container egress works; +**the LAN gateway is unreachable from both containers**; the host remains +reachable from them; the management console remains reachable from the LAN. NAT +and FORWARD rules present live and persisted, in order. Proxy serves the service +FQDN over TLS without an insecure bypass, redirects HTTP, and +**`X-Forwarded-Proto: https` is observed at the backend**. Any supervised +scaffolding survives a reboot. -# ---- egress from CT 100 ---- -pct exec 100 -- bash -c ' - for h in deb.debian.org pypi.org files.pythonhosted.org gitea.barternetwork.us; do - curl -sS -o /dev/null -w "$h %{http_code}\n" "https://$h" - done - git ls-remote https://gitea.barternetwork.us/TheRON/mechanical-compiler HEAD | head -1 -' +**PROVEN (F-021)** — the reboot is not optional. Connectivity can be perfect +while boot is degraded. -# ---- CT 101 and end to end ---- -pct exec 101 -- nginx -t -pct exec 101 -- curl -sS -o /dev/null -w '%{http_code} %{ssl_verify_result}\n' \ - https://mechanical-compiler.dev.infra/ +### Gate 2 — Deferred operational services -# ---- forwarded-proto reaches the app ---- -pct exec 101 -- curl -sS https://mechanical-compiler.dev.infra/_debug/headers \ - | grep -i x-forwarded-proto # expect https +Mail delivered **end to end and received**. Disk monitoring reporting a non-zero +device count. Not gated on gate 1. -# ---- idempotency: second run is a clean no-op ---- -./stage1-host.sh && ./stage2-create-cts.sh && echo "re-run clean" +### Gate 3 — Application runtime -# ---- backup round trip ---- -vzdump 100 --storage local --mode snapshot --compress zstd -ls -lh /var/lib/vz/dump/ # REPORT the size -pct exec 100 -- /usr/local/bin/mechcomp-backup --stdout | wc -c -``` +Dependencies installed from committed manifests. Real service and worker units +active. Reference toolchain image reproduces `ddd0f154…`. Scaffolding removed. -### Promotion checklist, for later +### Gate 4 — Backup -Not now, but recorded so it is not invented under pressure: production must -revisit the Proxmox firewall (§3.5), issue a Let's Encrypt certificate for the -public FQDN (§5.2), confirm the reverse proxy is not on the WireGuard side -(§2.2), and change exactly three variables in `mechcomp.env` (§10). +Undefined pending a strategy decision. + +### Gate 5 — Production automation + +Idempotent scripts derived from a proven manual procedure. Firewall posture +revisited. Publicly-trusted certificate issued and renewing. --- -## 15. Deliberately out of scope - -- Hubzilla addon development, or any change to existing Hubzilla containers -- Federation, identity, or qualification storage -- PostgreSQL — SQLite is sufficient and will remain so for a long time -- Redis, Celery, or any external queue -- CI runners — the suite runs locally until there is a reason otherwise -- Monitoring beyond journald, smartd, and Proxmox's own metrics -- NIC bonding, VLANs, or any further network configuration (§3.2) -- An internal DNS server (§3.4) -- The YunoHost package and the Docker image themselves - -Each has a natural moment. None of them is now. - ---- - -## 16. Every assumption in one place +## 16. Assumptions | § | Assumption | If wrong | |---|---|---| -| 3.3 | `10.0.0.20` and `.21` are free and outside any DHCP pool | ARP verify aborts stage 2; supply replacements | -| 3.3 | Static addressing, no DHCP reservation needed | Reserve in DHCP instead; addresses unchanged | -| 3.5 | Proxmox firewall stays disabled in staging | Enable per-container; production revisits regardless | -| 4.1 | CT IDs 100 and 101 | Verified at create time; any free pair works | -| 4.2 | 16 vCPU / 16 GiB for the app container | Lower freely; nothing depends on it | -| 5.1 | Kane County CA issues the staging certificate | A self-signed cert works; trust store step changes | -| 5.2 | Production uses HTTP-01 | Switch to DNS-01 if that is existing practice; never both | -| 9.2 | `keep-last=3` initially | Raise once the first archive size is measured | -| 9.3 | LUKS2 on the gold stick, ext4 inside | exFAT plus `age`-encrypted files if cross-OS reading is needed | -| 9.3 | Gold minted per tagged release and before promotion | Any trigger works; this one ties to existing discipline | -| 11.1 | Separate `mechanical-compiler-infra` repository | Use `deploy/` in the app repo, excluded from packaging | -| 11.3 | All stages idempotent | Only matters if you want create-once semantics; I would argue against | -| 11.5 | Infra versions independently of `SB_REVISION` | Tie them together if you prefer one release train | +| 3.6 | Firewall may stay disabled on a non-production instance | Enable per-container; production revisits regardless | +| 5.3 | Production uses HTTP-01 | Switch to DNS-01 if that is existing practice; never both | +| 11 | `MECHCOMP_BIND` is a single address | Revisit only if a real requirement for two sockets appears | +| 15 | Five gates are the right split | Merge or split further as evidence warrants | --- ## 17. Provenance disclosure This project's documents, code and roadmap are LLM-generated under human -direction. That is stated plainly here, in the README, and in any eventual -catalog submission. We do not obscure it. +direction. Stated plainly here, in the README, and in any catalog submission. **The YunoHost policy is a quality bar with a disclosure requirement, not a -ban.** Revision 1 overstated it. The catalog repository rejects generated -packages *that do not follow the `example_ynh` template*, citing verbose code, -hallucinated helpers and non-standard directory architectures — and then -explicitly permits AI use provided the maintainer is transparent and can -explain every line. The failure mode guarded against is sprawl, not provenance. -The defence is discipline: minimal divergence from the template, no invented -helpers, no extra files. +ban.** The catalog rejects generated packages *that do not follow the +`example_ynh` template*, citing verbose code, hallucinated helpers and +non-standard directory architectures — then explicitly permits AI use where the +maintainer is transparent and can explain every line. The failure mode guarded +against is sprawl, not provenance. -**Package provenance is not upstream provenance.** The policy text concerns the -`_ynh` repository — a few hundred lines of shell and one TOML file. Nobody -audits whether an upstream Rails application was written by a human. The -reviewable surface is small. +**Package provenance is not upstream provenance.** The policy concerns the +`_ynh` repository — a few hundred lines of shell and one TOML file. -**Disclosure is right; headlining it is a tactical mistake.** Part II §29 holds -that early use cases should demonstrate the system's distinct value. If -provenance becomes the pitch, the project is judged on that axis rather than on -whether it makes distributed manufacturing capacity legible. State it clearly -in the README; keep the pitch about manufacturing. The argument is won by an -artifact a reviewer finds shorter and more conventional than average, not by -the claim attached to it. +**Disclosure is right; headlining it is a tactical mistake.** If provenance +becomes the pitch, the project is judged on that axis rather than on whether it +makes distributed manufacturing capacity legible. State it in the README; keep +the pitch about manufacturing. --- -## 18. The packaging path, when it comes +## 18. Packaging, when it comes -**DISCOVER** — Unknown territory, to be documented as it is walked rather than -researched ahead: `manifest.toml` v2 with helpers 2.1, the `scripts/` set -(install, remove, upgrade, backup, restore, change_url), `conf/` templates, -`tests.toml`, and `package_check` levels. +**DISCOVER** — Unknown territory, documented as it is walked: `manifest.toml` v2 +with helpers 2.1, the `scripts/` set, `conf/` templates, `tests.toml`, +`package_check` levels. -Two orienting notes. The useful comparison for acceptable weight is -`paperless-ngx_ynh` or `fab-manager_ynh`, not a hello-world package. And per -§17, the package is drafted against `example_ynh` and reviewed by CIVICVS as -his own work — anything in it he would not defend in a review thread does not -ship. +The useful comparison for acceptable weight is `paperless-ngx_ynh` or +`fab-manager_ynh`, not a hello-world. The package is drafted against +`example_ynh` and reviewed by CIVICVS as his own work. --- -## 19. First work item after acceptance +## 19. First application work item -Port `sb-geom` to Shapely and get `pytest -n auto` green against the 123 frozen -cases in `strap-beam-fixtures-8.0.0.json`. The ten rejected cases are part of -the contract: a port that accepts them is wrong. +Port `sb-geom` to Shapely; `pytest -n auto` green against the 123 frozen cases in +`strap-beam-fixtures-8.0.0.json`. The ten rejected cases are part of the +contract: a port that accepts them is wrong. diff --git a/docs/FAILURES.md b/docs/FAILURES.md index e1f8894..4df101b 100644 --- a/docs/FAILURES.md +++ b/docs/FAILURES.md @@ -6,6 +6,7 @@ Compiler environment. | | | |---|---| | Scope | All instances. Staging entries are marked `srv-b`. | +| Updated | 2026-08-16, after staging acceptance | | Rule | Append only. Never edit an entry except to add a `Resolution` line. | | Numbering | Sequential, never reused. See §0 on the renumbering. | @@ -315,11 +316,19 @@ CT 100. **Observed:** After the F-017 reboot, nothing listened on `10.20.0.10:8770`. **Cause:** The placeholder was a bare foreground process with a PID file. No supervision. -**Correction:** Pending — to be recreated as `mechcomp-placeholder.service`. +**Correction:** Recreated as `mechcomp-placeholder.service`. **Consequence:** Anything a proof depends on must be supervised, or the proof expires silently at the next reboot and the following session diagnoses a proxy fault that does not exist. Applies to temporary scaffolding as much as to real services. +**Resolution (2026-08-16):** Installed `/usr/local/libexec/mechcomp-placeholder.py` +and `/etc/systemd/system/mechcomp-placeholder.service`, enabled at +`multi-user.target`, running as `mechcomp:mechcomp`, reading +`/etc/mechcomp/mechcomp.env`, binding `10.20.0.10:8770`, with `Restart=on-failure` +and the standard hardening set. CT 100 was rebooted: the unit restarted +automatically, the listener returned on the service address only, nginx proxied +successfully, `X-Forwarded-Proto: https` was re-observed at the backend, and the +container settled to `running` with zero failed units. --- @@ -340,10 +349,113 @@ topology decision. Recorded because the reasoning matters more than the value: --- +### F-021 — deleting the Proxmox interface left a stale guest `eth1` +CT 100, CT 101. Network isolation. + +**Observed:** After the F-017 topology change and reboot, both containers +reported `systemctl is-system-running` → `degraded`, with +`networking.service` and `ifupdown-wait-online.service` failed. Both guests' +`/etc/network/interfaces` still carried an `auto eth1` / static `iface eth1` +stanza although neither had an `eth1` link. The boot journal on both: + +``` +Cannot find device "eth1" +ifup: failed to bring up eth1 +``` + +**Cause:** **Proven.** `pct set --delete net1` removes the LXC interface but +does not remove the stanza already written into the guest. `ifup -a` therefore +exited non-zero at boot even though `eth0` came up correctly. +`ifupdown-wait-online` failed as a consequence, not independently. +`systemd-networkd` was investigated and explicitly ruled out — its units were +disabled on both containers. +**Correction:** Preserved a pre-correction copy, removed only the stale `eth1` +stanza on each guest, restarted the affected units. Both containers then +rebooted: the stanza did not return, both units succeeded, both settled to +`running` with zero failed units. +**Consequence:** **Removing a container interface is not proof that the guest +converged.** After any topology mutation, automation must inspect the guest +interface file and assert `systemctl is-system-running = running` with zero +failed units *after a reboot*. Connectivity alone is insufficient — the +surviving interface works fine while boot remains degraded, which is precisely +how this went unnoticed through an entire verification pass. + +--- + +### F-022 — transient DNS resolution timeout in CT 101 +CT 101. Formal acceptance. + +**Observed:** During the first acceptance pass, two unrelated public HTTPS +tests both failed at name resolution: + +``` +deb.debian.org: curl: (28) Resolving timed out after 5000 ms +gitea.barternetwork.us: curl: (28) Resolving timed out after 5000 ms +``` + +IP routing to `10.110.0.1`, `10.0.0.12` and CT 100 remained working throughout. +**Cause:** **Unproven.** Diagnostics showed resolver configuration identical to +CT 100 and the host (`nameserver 75.75.75.75`, `hosts: files dns`), +`systemd-resolved` absent, `75.75.75.75` reachable, and `getent ahostsv4` +resolving both names immediately afterward. Repeat HTTPS tests returned 200. + +Worth noting as context, not as cause: since F-017, container DNS traverses the +host's masquerade to an external resolver. That dependency is new. +**Correction:** None. No configuration was changed. +**Consequence:** Do not convert a one-shot resolver timeout into a +configuration change without evidence. On recurrence, capture resolver state +and DNS traffic at the moment of failure before touching anything. A candidate +mitigation — a second `nameserver` line, so a single hiccup retries rather than +fails — is recorded but deliberately not applied on one unexplained event. + +--- + +### F-023 — relay accepted the mail; final delivery failed +`srv-b`. Alerting. + +**Observed:** The relay was discovered at `10.110.0.1:25` over `wg0`, banner +`wg-pk.diagnostics.kane-il.us`, offering STARTTLS with a self-signed +`CN = wg-pk`. It accepted unauthenticated SMTP from `10.110.0.12` through +`RCPT TO`, before and after STARTTLS. Ports 465 and 587 were unavailable, and +the public address `198.58.111.109` did not expose SMTP on this path. + +With `relayhost = [10.110.0.1]:25` and `root: sandor@kane-il.us`, `srv-b` +recorded successful handoff: + +``` +relay=10.110.0.1[10.110.0.1]:25 dsn=2.0.0 +status=sent (250 2.0.0 Ok: queued as 87E446243A) +``` + +The local queue emptied. The operator then received a **delivery-failure** +message at `sandor@kane-il.us`. +**Cause:** **Unproven**, downstream of the demonstrated handoff. The bounce +notice itself arriving at `sandor@kane-il.us` establishes that the relay can +deliver to that address, which narrows the problem to the failing message +rather than the destination. Leading hypothesis, untested: the envelope sender +is `root@srv-b.dev.infra`, and `dev.infra` does not resolve publicly, so a +downstream MTA rejects on sender-domain verification. Candidate remedies are +`myorigin` or `smtp_generic_maps`. +**Correction:** None. Mail alerting was deferred by operator decision. +**Consequence:** **SMTP 250 from the relay and an empty local queue prove +handoff, not delivery.** Alerting acceptance requires demonstrated end-to-end +receipt. Until then mail, `smartd` alerting, and any mail-dependent backup +alerting are unaccepted. Note separately that `postfix check` reports +divergence between `/var/spool/postfix` copies and their host originals, +including `/etc/hosts` and NSS libraries — a known cause of resolution failure +inside the chroot, and adjacent enough to this failure to be checked first. + +--- + ## Open, not closed | # | Status | |---|---| -| F-006 | Cause unproven. Recurrence should capture `dpkg` lock state. | -| F-012 | Cause unproven. Leading candidate ruled out by inspection. | -| F-019 | Correction pending. | +| F-006 | **Open** — cause unproven. Recurrence should capture `dpkg` lock state. | +| F-012 | **Open** — cause unproven. Leading candidate ruled out by inspection. `default_server` was added as independent hardening and does **not** close this. | +| F-019 | **Corrected** 2026-08-16. Reboot persistence proven. | +| F-021 | **Corrected** 2026-08-16. Reboot persistence proven. | +| F-022 | **Open** — cause unproven, no correction applied. | +| F-023 | **Deferred** — downstream mail failure, cause unproven. Hypothesis recorded. | + +Everything else is closed with a proven cause and a proven correction. diff --git a/docs/ROADMAP.md b/docs/ROADMAP.md index 2fded99..2da85b3 100644 --- a/docs/ROADMAP.md +++ b/docs/ROADMAP.md @@ -4,7 +4,7 @@ What the Mechanical Compiler is for, and the order in which it gets built. | | | |---|---| -| Updated | 2026-08-15 | +| Updated | 2026-08-16 | | Companions | `ENVIRONMENT.md`, `STAGING-STATE.md`, `FAILURES.md` | --- @@ -121,7 +121,12 @@ a new call into the same machinery, not a new machine. provably non-chiral; profile parameters are isolated from one another. - **Fixture oracle frozen.** 123 cases, `ddd0f154…`, pinned to OpenSCAD 2021.01 and BOSL2 `92d697c`. The ten rejected cases are part of the contract. -- **Staging environment.** In progress — see `STAGING-STATE.md`. +- **Staging infrastructure accepted, 2026-08-16.** Host, both containers, + network isolation, bastion access, TLS trust, reverse proxy and the + application filesystem foundation are proven on `srv-b`. Mail alerting and + explicit disk monitoring are deferred; backup strategy is postponed by + operator decision; application deployment remains blocked on application + code. See `STAGING-STATE.md`. ### Next — the port @@ -235,3 +240,9 @@ on whether it makes distributed manufacturing capacity legible. 6. **Record the failure before correcting it.** Then apply the smallest corrective change, not the one that also fixes three things you were worried about. +7. **Prove the negative.** Isolation is asserted by showing the forbidden path + fails, never by showing the interface is gone (F-018). Health is asserted + after a reboot, never before (F-021). +8. **Handoff is not delivery.** An upstream acceptance code proves the message + left, not that it arrived (F-023). The same distinction applies wherever a + subsystem reports success on behalf of something downstream. diff --git a/docs/STAGING-STATE.md b/docs/STAGING-STATE.md index 194e95a..448b10c 100644 --- a/docs/STAGING-STATE.md +++ b/docs/STAGING-STATE.md @@ -4,7 +4,7 @@ Live state of the Mechanical Compiler staging instance on `srv-b`. | | | |---|---| -| Updated | 2026-08-15, after network isolation | +| Updated | 2026-08-16, formal infrastructure acceptance | | Instance | Staging / development | | Specification | `ENVIRONMENT.md` revision 5 | | Failure log | `FAILURES.md` | @@ -15,24 +15,36 @@ Live state of the Mechanical Compiler staging instance on `srv-b`. ## 0. How to use this file This is the authoritative record of what is true on `srv-b`. Where it and -`ENVIRONMENT.md` disagree, **this file wins for facts** and the specification -is defective and must be corrected. +`ENVIRONMENT.md` disagree, **this file wins for facts** and the specification is +defective and must be corrected. -Completed and remaining work are in the same document deliberately. They are -two halves of one boundary; separating them guarantees they drift. +Completed and remaining work are in the same document deliberately. They are two +halves of one boundary; separating them guarantees they drift. Before any command: read this file, confirm the immediately relevant live state -with a read-only command, then issue one command group. If it fails, record it -in `FAILURES.md` before changing anything else. +with a read-only command, then issue one command group. If it fails, record it in +`FAILURES.md` before changing anything else. -Production will get its own `PRODUCTION-STATE.md`. The specification is shared; -the state is not. +Production gets its own `PRODUCTION-STATE.md`. The specification is shared; the +state is not. + +### Acceptance boundary, 2026-08-16 + +**Staging infrastructure is accepted.** That claim is narrower than "the +application is deployed," and deliberately so. Three subsystems sit outside it +and must not be represented as either hidden failures or completed work: + +| Subsystem | Status | +|---|---| +| Mail alert delivery | Partially configured, characterised, **not accepted end to end** | +| `smartd` monitoring | Deferred with mail alerting | +| Backup infrastructure | **Postponed by operator decision** — strategy may change | --- ## 1. Instance values -These are the `srv-b` bindings for the parameters in `ENVIRONMENT.md`. +The `srv-b` bindings for the parameters in `ENVIRONMENT.md`. ### Host @@ -40,13 +52,16 @@ These are the `srv-b` bindings for the parameters in `ENVIRONMENT.md`. hostname srv-b / srv-b.dev.infra platform Proxmox VE 8.4.0, Debian 12, kernel 6.8.12-9-pve hardware HP ProLiant DL360 G7, 2 x Xeon X5650, 24 threads, 31 GiB -storage P410i, 4 x EG0146FAWHU, RAID 1+0, all SMART OK +storage P410i, 4 x EG0146FAWHU, RAID 1+0, all members SMART OK local directory /var/lib/vz, ~70 GiB free iso,vztmpl,backup local-lvm LVM-thin pve/data, 166.9 GiB rootdir,images vmbr0 10.0.0.12/24 on enp3s0f0, gw 10.0.0.1 management, LAN vmbr1 10.20.0.1/24, bridge-ports none service, portless wg0 10.110.0.12/32, peer wg-pk.civicus.us:51820, allowed 10.110.0.0/22 resolver 75.75.75.75, search dev.infra +ip_forward 1 +timezone America/Chicago, NTP active, clock synchronised +systemd running, zero failed units spare NICs enp3s0f1, enp4s0f0, enp4s0f1 — unconfigured, deliberately template local:vztmpl/debian-12-standard_12.12-1_amd64.tar.zst ``` @@ -62,9 +77,11 @@ template local:vztmpl/debian-12-standard_12.12-1_amd64.tar.zst | rootfs | 40 GiB `local-lvm` | 8 GiB `local-lvm` | | `mp0` | 60 GiB → `/var/lib/mechcomp`, `backup=0` | — | | `net0` | `eth0` on `vmbr1`, `10.20.0.10/24`, gw `10.20.0.1` | `eth0` on `vmbr1`, `10.20.0.11/24`, gw `10.20.0.1` | -| Debian | 12.15 | 12.15 | +| Debian | 12 bookworm, fully upgraded | 12 bookworm, fully upgraded | +| systemd | running, zero failed units | running, zero failed units | -**Neither container has a LAN interface.** See F-017. +**Neither container has a LAN interface** (F-017). **Neither has a stale `eth1` +stanza** (F-021) — verified persistent across reboot. ### Names and identity @@ -73,16 +90,15 @@ mechanical-compiler.dev.infra -> 10.20.0.11 (CT 101, the proxy) mechcomp.dev.infra -> 10.20.0.10 PVE-generated mcproxy.dev.infra -> 10.20.0.11 PVE-generated -ADMIN_USER sandor - uid 1000, groups sandor + mechcomp(996) - shell /bin/bash, both containers -authorized key SHA256:2pNffCscUUW5Wbs9uMepLEvMLKPqUV/7Lpk5tw7iWSY - root@srv-b — bastion pattern, see §4 -service user mechcomp, uid 999, gid 996 - home /var/www/mechcomp, shell /usr/sbin/nologin +ADMIN_USER sandor, uid 1000, groups sandor + mechcomp(996) + shell /bin/bash, both containers +authorized key SHA256:2pNffCscUUW5Wbs9uMepLEvMLKPqUV/7Lpk5tw7iWSY + matches /root/.ssh/id_rsa.pub on srv-b — bastion, see section 4 +service user mechcomp, uid 999, gid 996 + home /var/www/mechcomp, shell /usr/sbin/nologin ``` -### Host NAT, as persisted in `/etc/iptables/rules.v4` +### Host NAT, persisted in `/etc/iptables/rules.v4` ``` nat POSTROUTING @@ -96,7 +112,8 @@ filter FORWARD -s 10.20.0.0/24 -d 10.0.0.0/24 -j DROP # LAN blocked ``` -Rule order is load-bearing. `RETURN` must precede both masquerades. +Rule order is load-bearing. `RETURN` must precede both service-network +masquerades. Live and persisted states match. ### TLS @@ -104,15 +121,14 @@ Rule order is load-bearing. `RETURN` must precede both masquerades. CA CN = Mechanical Compiler Staging CA (locally generated) leaf CN = mechanical-compiler.dev.infra SAN = DNS:mechanical-compiler.dev.infra -validity 2026-08-16 -> 2028-11-18 <-- expires, nothing renews it +validity 2026-08-16 -> 2028-11-18 <-- nothing renews this trusted srv-b, CT 100, CT 101 -key mode 0600 on CT 101 +key root:root 0600 /etc/ssl/mechcomp/server.key ``` -No Kane County Civic Infrastructure CA issuance path exists on `srv-b`: no -trust anchor, no `step`, no `cfssl`, no EasyRSA. Proxmox's own CA was -deliberately not reused. Replacing this leaf later is two file copies and a -reload. +No Kane County Civic Infrastructure CA issuance path exists on `srv-b`: no trust +anchor, no `step`, no `cfssl`, no EasyRSA. Proxmox's own CA was deliberately not +reused. Replacing this leaf later is two file copies and a reload. ### Application environment @@ -120,7 +136,7 @@ reload. ``` MECHCOMP_ENV=staging -MECHCOMP_BIND=10.20.0.10 +MECHCOMP_BIND=10.20.0.10 # scalar; loopback is NOT bound MECHCOMP_PORT=8770 MECHCOMP_BASE_URL=https://mechanical-compiler.dev.infra MECHCOMP_DATA_DIR=/var/lib/mechcomp @@ -132,6 +148,36 @@ MECHCOMP_ARTIFACT_RETENTION_DAYS=30 MECHCOMP_SECRET_KEY= ``` +### Mail, as currently configured + +``` +relay endpoint 10.110.0.1:25 over wg0 +banner wg-pk.diagnostics.kane-il.us, STARTTLS offered +STARTTLS cert self-signed, CN = wg-pk +authentication unauthenticated accepted from 10.110.0.12 +ports 465 / 587 unavailable; public 198.58.111.109 exposes no SMTP on this path +srv-b relayhost [10.110.0.1]:25 +smtp_tls_security_level may +smtp_sasl_auth_enable no +inet_interfaces loopback-only +root alias sandor@kane-il.us +local handoff succeeds, relay returns SMTP 250, queue empties +FINAL DELIVERY NOT ACCEPTED — operator received a delivery-failure message +status deferred, cause unproven downstream (F-023) +``` + +### Placeholder backend — staging scaffold, not application code + +``` +/usr/local/libexec/mechcomp-placeholder.py +/etc/systemd/system/mechcomp-placeholder.service +runs as mechcomp:mechcomp, reads mechcomp.env, binds 10.20.0.10:8770 +Restart=on-failure, hardening set applied, enabled at multi-user.target +``` + +Exists solely to prove the proxy chain independently of the application. It is +removed when the real service arrives. + ### Rollback copies on disk ``` @@ -143,7 +189,9 @@ srv-b /etc/network/interfaces.before-mechcomp /root/pct-100.before-svcnet /root/pct-101.before-svcnet CT 100 /etc/hosts.before-svcfqdn + /etc/network/interfaces pre-F-021 copy CT 101 /etc/hosts.before-svcfqdn + /etc/network/interfaces pre-F-021 copy ``` --- @@ -156,122 +204,136 @@ Each line was demonstrated by command output, not inferred. - [x] `vmbr1` created, active, `10.20.0.1/24`, portless - [x] Duplicate address detection run before each container creation -- [x] Container → internet reachable (`deb.debian.org`, `gitea.barternetwork.us`, both 200) -- [x] Container → WireGuard network reachable (`10.110.0.12`) -- [x] **Container → LAN unreachable** — the isolation requirement -- [x] Container → `srv-b` reachable at `10.0.0.12` (INPUT path, required) -- [x] Container ↔ container reachable over `vmbr1` -- [x] LAN → Proxmox console unaffected (`10.0.0.12:8006` → 200) -- [x] Network configuration persisted, survives reboot +- [x] Container to internet reachable (`deb.debian.org`, `gitea.barternetwork.us`) +- [x] Container to WireGuard reachable (`10.110.0.1`) +- [x] **Container to LAN gateway `10.0.0.1` = 100% packet loss** — the requirement +- [x] Container to `srv-b` `10.0.0.12` reachable via INPUT path (required, not a leak) +- [x] Container to container reachable over `vmbr1` +- [x] LAN to Proxmox console `10.0.0.12:8006` returns 200, unaffected +- [x] NAT and FORWARD rules present live **and** persisted, in correct order +- [x] Host: `running`, zero failed units ### CT 100 -- [x] Debian 12.15, systemd running, zero failed units, no pending upgrades -- [x] `en_US.UTF-8` generated and active +- [x] Debian 12 bookworm, fully upgraded, zero pending +- [x] `en_US.UTF-8` generated and active; `America/Chicago` +- [x] **`running`, zero failed units — after reboot** (F-021 corrected) - [x] `/var/lib/mechcomp` is a real separate ext4 filesystem -- [x] Service user `mechcomp`, home `/var/www/mechcomp`, writable -- [x] Application directory layout created with correct ownership -- [x] Base packages installed; OpenSCAD / Qt / X11 absent — verified `clean` -- [x] Docker working, `overlay2` / `systemd`, no `fuse-overlayfs` needed +- [x] Service user `mechcomp`, home `/var/www/mechcomp`, correct ownership +- [x] All four data subdirectories present, `mechcomp:mechcomp`, `0750` +- [x] `mechcomp.env` present, `root:mechcomp`, `0640`, secret generated +- [x] Base packages installed; OpenSCAD, Qt and X11 absent +- [x] Docker active, `overlay2` / `systemd`, no fallback needed - [x] Repository cloned at `e85c4f4e`, verified as the owning user - [x] Python 3.11.2 venv created, owned by `mechcomp` -- [x] `mechcomp.env` written with generated secret -- [x] `openssh-server` installed, enabled, active -- [x] Application listener bound to service network only — LAN-side bind refused +- [x] `openssh-server` enabled and active +- [x] **`mechcomp-placeholder.service` enabled, active, reboot-persistent** (F-019) +- [x] Listener on `10.20.0.10:8770` only — not `0.0.0.0` ### CT 101 -- [x] Debian 12.15, systemd running, zero failed units after `nesting=1` -- [x] `en_US.UTF-8` generated and active -- [x] nginx 1.22.1 installed, `nginx -t` passes -- [x] Debian `default` site removed; only the project vhost is enabled +- [x] Debian 12 bookworm, fully upgraded, zero pending +- [x] `en_US.UTF-8` generated and active; `America/Chicago` +- [x] **`running`, zero failed units — after reboot** (F-021 corrected) +- [x] nginx 1.22.1, `nginx -t` passes, enabled and active +- [x] Only the project vhost enabled; Debian `default` removed +- [x] `listen 443 ssl default_server` on both address families +- [x] HTTP to HTTPS 301 redirect - [x] Local CA created, leaf issued, trusted on all three hosts -- [x] HTTP → HTTPS redirect, TLS termination, proxy to `10.20.0.10:8770` -- [x] `openssh-server` installed, enabled, active -- [x] **`X-Forwarded-Proto: https` observed at the backend** — needs re-proof, see §3 +- [x] **HTTPS end to end from `srv-b`, CT 100 and CT 101 without `-k`** +- [x] **`X-Forwarded-Proto: https` observed at the backend** — re-proved after the + topology change and again after CT 100's reboot +- [x] `openssh-server` enabled and active --- ## 3. Remaining — infrastructure -In dependency order. Application-dependent work is in §5. - ### Immediate -- [ ] **`mechcomp-placeholder.service`** — recreate the backend as a supervised - unit. It died on the F-017 reboot and nothing currently listens on 8770. - Everything below that touches the proxy depends on this. See F-019. -- [ ] **Re-prove `X-Forwarded-Proto`** end to end. The earlier proof was taken - when clients were on `10.0.0.x`; header values will now read `10.20.0.x`. - Re-establish rather than assume it survived the topology change. -- [ ] **`default_server`** on the CT 101 vhost, both `listen 443` lines. Two - words. Prevents catch-all behaviour depending on file ordering once a - second server block exists. See F-012. +Nothing. The infrastructure boundary is accepted. -### Mail — blocks alerting +### Deferred by operator decision -- [ ] **Postfix relay.** Currently `inet_interfaces = loopback-only`, - `relayhost` empty, no `root:` alias. Alerts go to a mailbox nobody opens. -- [ ] **`root:` alias** to a real destination. -- [ ] Needs: address of the `wg-pk` relay, and whether it accepts - unauthenticated from `10.110.0.12`. **Open question for CIVICVS.** +Recorded here so they are visually distinct from failures and from forgotten +work. None is a defect. -### Monitoring and backup +- [ ] **Mail end-to-end delivery** (F-023). Handoff to `wg-pk` works; final + delivery does not. Leading untested hypothesis: envelope sender + `root@srv-b.dev.infra` rejected on sender-domain verification, since + `dev.infra` does not resolve publicly. Candidate remedies `myorigin` or + `smtp_generic_maps`. Check the `postfix check` chroot divergence first. +- [ ] **`smartd` explicit four-member configuration.** The package is installed + and the service is active, but `DEVICESCAN` currently monitors **zero + devices**; the P410i members are visible only through explicit + `-d cciss,N`. Active is not the same as monitoring. Deferred with mail. +- [ ] **Backup infrastructure, entirely.** Postponed 2026-08-16 because the + strategy may change: `vzdump` job, archive sizing, retention, free-space + guard, host-side pull, `mechcomp-backup`, backup alerting, gold media, + 3+ TB redundancy. -- [ ] **`smartd`** on `/dev/sda -d cciss,0` through `cciss,3`. Depends on mail. -- [ ] **`vzdump` job** — both containers, `local`, snapshot, zstd, 02:30. - `mp0` already carries `backup=0`. -- [ ] **First `vzdump` run — report the archive size.** Retention is a - placeholder `keep-last=3` until this number exists. -- [ ] **Free-space guard** — refuse and alert below 20 GiB on `/var/lib/vz`. -- [ ] `/var/lib/vz/mechcomp-app/` and the host-side pull script. -- [ ] `mechcomp-backup --stdout` in CT 100. Structure can be built and tested - against the current data directories without application output. +**One request standing against the postponement:** a single manual `vzdump` of +both containers as a point-in-time snapshot before further change. It +presupposes nothing about the eventual strategy and yields the archive size that +has been an open question since revision 4. Four 2010-vintage disks currently +hold the only instance with no monitoring, no alerting and no backup. -### Gold media +### Optional, recorded not scheduled -- [ ] 32 GB USB stick: LUKS2, ext4, label `MC-GOLD-01`, `/mnt/gold` `noauto`. -- [ ] `gold-archive.sh` with `SHA256SUMS` verification before unmount. -- [ ] **Open:** where the 3+ TB USB disk is attached now that `annales` is out - of scope. Until known, stick-to-disk copying is a manual step. +- [ ] Second `nameserver` line in container resolver configuration, as a + mitigation for F-022. One line, no daemon. Deliberately not applied on a + single unexplained event. +- [ ] Staging certificate expires 2028-11-18 and nothing renews it. --- ## 4. Access model -Confirmed by CIVICVS: no workstation access from the home LAN is required. +Confirmed: no workstation access from the home LAN is required. ``` internet -> WireGuard -> srv-b -> containers ``` -`srv-b` is the bastion. The authorized key is `root@srv-b`, which is -consistent: root on the host can `pct enter` regardless, so SSH to the -containers adds no privilege. It does mean **every path to a container runs -through `srv-b`** — deliberate, not a limitation. +`srv-b` is the bastion, confirmed by the authorized key matching +`/root/.ssh/id_rsa.pub` on the host. Root on the host can `pct enter` regardless, +so SSH adds no privilege — but **every path to a container runs through +`srv-b`**, which is deliberate. -If direct WireGuard-side access to the catalogue is wanted later, that is a -route addition for `10.20.0.0/24` on the hub. Not now. +Direct WireGuard-side access to the catalogue would be a route addition for +`10.20.0.0/24` on the hub. Not now. -No DHCP anywhere. Two containers with fixed addresses on a portless bridge is -the entire address space; a DHCP server would add a daemon, a lease database -and a failure mode in exchange for nothing. +No DHCP anywhere. Two containers with fixed addresses on a portless bridge is the +entire address space. --- ## 5. Blocked on application code -None of this can be honestly completed while the repository is `LICENSE` and -`README.md`. It is not provisioning work and should not be attempted as such. +Repository state: -- [ ] `requirements-base.txt` / `requirements-cad.txt` and dependency install -- [ ] `mechcomp.service` and `mechcomp-worker.service` +``` +HEAD e85c4f4e5bab9a4f032230c99aa6784ace4c80e2 +top level .git LICENSE README.md venv +git status ?? venv/ +absent requirements-base.txt, requirements-cad.txt, service units +``` + +None of the following can be honestly completed, and none may be fabricated by +provisioning: + +- [ ] Dependency install from committed manifests +- [ ] `mechcomp.service`, `mechcomp-worker.service` - [ ] Reference toolchain image `mechcomp/reference-toolchain:8.0.0` -- [ ] Fixture reproduction against `ddd0f154…` +- [ ] Fixture reproduction against `ddd0f154...` - [ ] Application-runtime acceptance - [ ] Replacing the placeholder with the real service -The Shapely port is the architect's work item and gates all of the above. +**Architect decision:** `venv/` belongs in `.gitignore` — it is a legitimate +artifact inside `install_dir` per YunoHost convention, it simply should not be +tracked. Applied with the first application commit. + +The Shapely port gates all of the above. --- @@ -279,11 +341,11 @@ The Shapely port is the architect's work item and gates all of the above. | # | Question | Blocks | |---|---|---| -| 1 | `wg-pk` relay address; unauthenticated from `10.110.0.12`? | mail, `smartd`, backup alerts | -| 2 | Where is the 3+ TB USB disk attached? | gold redundancy step | -| 3 | First `vzdump` archive size | final retention value | +| 1 | Why does delivery fail downstream of `wg-pk`? | mail, `smartd`, backup alerting | +| 2 | What is the backup strategy? | all backup work | +| 3 | Where is the 3+ TB USB disk attached? | gold redundancy step | -Question 3 is answered by doing. Questions 1 and 2 need CIVICVS. +Question 1 is answered by diagnosis. Questions 2 and 3 need CIVICVS. --- @@ -291,9 +353,11 @@ Question 3 is answered by doing. Questions 1 and 2 need CIVICVS. | Question | Answer | |---|---| -| Kane County CA issuance path | None on `srv-b`. Local staging CA generated instead. | +| Kane County CA issuance path | None on `srv-b`. Local staging CA generated. | +| Relay address and authentication | `10.110.0.1:25` over `wg0`, unauthenticated from `10.110.0.12` accepted. | | `openssh-server` present | Yes, both containers. | -| `ADMIN_USER` / key | `sandor`; `root@srv-b` key, bastion pattern. | -| DHCP pool on `10.0.0.0/24` | **Not applicable.** Containers are no longer on that network. | +| `ADMIN_USER` / key | `sandor`; `root@srv-b` key, bastion pattern confirmed. | +| DHCP pool on `10.0.0.0/24` | **Not applicable.** Containers are not on that network. | | Docker storage driver | `overlay2` / `systemd`. No fallback needed. | | LAN workstation access | Not required. WireGuard through `srv-b`. | +| `MECHCOMP_BIND` semantics | Scalar, service address only. Loopback not bound. |