diff --git a/docs/ENVIRONMENT.md b/docs/ENVIRONMENT.md new file mode 100644 index 0000000..c317520 --- /dev/null +++ b/docs/ENVIRONMENT.md @@ -0,0 +1,1063 @@ +# ENVIRONMENT.md + +Provisioning specification for the **Mechanical Compiler** development and +staging environment. + +| | | +|---|---| +| Revision | 4 (2026-08-15) | +| Supersedes | Revisions 1, 2 and 3 | +| Basis | `ENVIRONMENT_INVESTIGATION_REPORT_2026-08-15.md` (host `srv-b`) | +| Audience | The assistant or operator provisioning the containers | +| Author role | Written by the developer who will work inside this environment | +| Status | All decisions taken. Ready to execute. | +| Repository | `https://gitea.barternetwork.us/TheRON/mechanical-compiler` | +| Licence | AGPL-3.0-or-later (committed at `e85c4f4e`) | + +--- + +## 0. How to use this document + +This specifies a **staging environment** on `srv-b`, plus the promotion path to +a production host that does not yet exist. + +Everything is scripted. Nothing is provisioned by hand. + +### Markers + +| Marker | Meaning | +|---|---| +| **REQ** | Required. The work is blocked or wrong without it. | +| **PREF** | Preferred. Substitute freely, but report what you substituted. | +| **ASSUMED** | Decided by the author without confirmation. Defensible, but a guess. All are listed in §16. | +| **FILL** | A site value that must be supplied. Goes in `deploy/site.env` and nowhere else. | +| **VERIFY** | Must be checked at run time rather than trusted from this document. | +| **DISCOVER** | Genuinely unknown. Answer by doing, then report. | + +### If you are the provisioning assistant + +Four requests. **FILL** values live in one file and are referenced from there — +no literals scattered through scripts. **VERIFY** steps run inside the scripts, +not as a pre-flight a human might skip. Report **DISCOVER** outcomes verbatim, +including failures. And if a **REQ** item cannot be satisfied, say so rather +than working around it: the stated reason usually matters more than the +mechanism. + +### What changed in revision 4 + +| # | Change | Source | +|---|---|---| +| 1 | Host is `srv-b`, not `annales` | CHG-001 | +| 2 | Staging and production are separate roles, on separate hosts | CHG-002 | +| 3 | `srv-b` is a standalone node; promotion is export/restore, not cluster migration | CHG-003 | +| 4 | Template `12.12-1`, not `12.7-1` | CHG-004 | +| 5 | Three-tier backup architecture; 32 GB removable gold media | CHG-005 | +| 6 | Backup schedule defined explicitly; none exists to inherit | CHG-006 | +| 7 | `CT_WG_ALIAS` removed | CHG-007 | +| 8 | `PROXY_HOST` not inherited; staging gets its own proxy container | CHG-008 | +| 9 | **Two containers**, not one: application and reverse proxy | this revision | +| 10 | **Second bridge** `vmbr1` isolates service traffic from management traffic | this revision | +| 11 | Container sizing raised substantially to use the host | this revision | +| 12 | Staging is internal-only; the public FQDN is reserved for production | this revision | + +--- + +## 1. Five decisions that shape everything below + +### 1.1 The 2D path and the 3D path are separable + +The catalogue's product is a 2D cross-section rendered to SVG. Shapely plus our +own code covers that completely: region algebra, offsets, distance queries, +output. CadQuery/OCCT is needed only for STEP and mesh export. + +These stay separate — separate requirements files, separate code paths, and a +CI job that runs the whole suite with the CAD dependencies absent. Three +reasons, none of which is packaging politics: + +1. Part II §16 requires the control-plane / execution-plane separation + regardless of how the app is distributed. +2. A slow OCCT import must never land in the request path for a page that only + draws a cross-section. +3. Test-suite speed. Shapely-only tests run in seconds, which is the difference + between running the 123 frozen fixtures on every save and only before commit. + +**Retained correction.** Revision 1 also justified this as a defence against +YunoHost's "resource-hungry" criterion. That was overcalibrated — +`fab-manager_ynh` is in the catalog, `paperless-ngx_ynh` declares nine apt +dependencies including postgresql and redis, and the criterion reads +"resource-hungry **compared to their features**." **CadQuery is a +default-installed dependency.** The boundary stays; the apology goes. + +### 1.2 YunoHost and Docker are parallel targets, not sequential ones + +YunoHost apps install natively — apt, a venv, systemd, nginx — and the project +does not want Docker inside YunoHost. Docker is a separate distribution +channel, not a stepping stone to the catalog. Neither blocks the other. + +### 1.3 Staging mirrors production topology, not production exposure + +`srv-b` runs two containers because production will: an application container +that never terminates TLS, and a reverse proxy that does. A single container +serving directly would never exercise the `X-Forwarded-Proto` path — precisely +the class of bug that cost us time on Hubzilla's session logout. + +Staging is **not** publicly reachable and does **not** use the public FQDN. It +uses the Kane County Civic Infrastructure CA on an internal name. The public +name and Let's Encrypt belong to production, which forces the promotion path to +be exercised rather than assumed. + +### 1.4 Directory layout mirrors YunoHost conventions from the start + +The paths in §7 are what a YunoHost package would provision anyway +(`install_dir`, `data_dir`, a dedicated system user, a dynamically assigned +port). The eventual `manifest.toml` resource block then describes what already +exists. + +### 1.5 All configuration comes from the environment, never from the filesystem + +No path, port, URL, or secret may be hardcoded or discovered by convention. One +env file per container; the same variables become `ENV` in a Dockerfile and +`ynh_add_config` substitutions in a package. + +The check that this holds: moving the service to a different hostname is a +one-variable change. It currently is — which is why the staging/production +split below costs almost nothing. + +--- + +## 2. Host baseline (confirmed, not to be re-verified) + +| Item | Value | +|---|---| +| Hostname / FQDN | `srv-b` / `srv-b.dev.infra` | +| Role | Development and staging. Standalone node, never clustered. | +| Platform | Proxmox VE 8.4.0 on Debian 12, kernel `6.8.12-9-pve` | +| Hardware | HP ProLiant DL360 G7, firmware P68 | +| CPU | 2 × Xeon X5650, 24 logical CPUs | +| RAM / swap | 31 GiB / 8 GiB | +| Storage | HP Smart Array P410i, 4 × EG0146FAWHU, RAID 1+0, all SMART OK | +| `local` | directory, `/var/lib/vz`, ~70 GiB free — iso, vztmpl, **backup** | +| `local-lvm` | LVM-thin `pve/data`, 166.9 GiB, 0% used — rootdir, images | +| LAN | `vmbr0` on `enp3s0f0`, `10.0.0.12/24`, gateway `10.0.0.1` | +| WireGuard | `wg0` `10.110.0.12/32`, peer `wg-pk.civicus.us:51820`, allowed `10.110.0.0/22` | +| Routing | forwarding on, `MASQUERADE 10.0.0.0/24 -o wg0`, both persistent | +| Resolver | `75.75.75.75`, search `dev.infra` | +| Egress | Direct. No proxy, no allowlist. All required hosts reachable. | +| PVE firewall | `disabled/running` | +| Guests | None. Next ID `100`. | +| Template | `local:vztmpl/debian-12-standard_12.12-1_amd64.tar.zst` | + +### 2.1 Two hardware facts worth knowing before anyone is surprised + +The X5650 is Westmere: **no AVX**, and single-thread performance is modest. +Geometry solvers are single-threaded, so regenerating the fixture oracle will +take noticeably longer here than on a modern laptop. That is expected, not a +fault. All compiled wheels target an SSE2 baseline, so nothing breaks. + +Against that, **24 threads is a great deal of parallel test capacity**. The +test suite runs under `pytest-xdist` with `-n auto`, which is where the box +earns its keep. + +### 2.2 The WireGuard path is outbound only + +`MASQUERADE` means containers reach `10.110.0.0/22`, but nothing on the +WireGuard network can initiate a connection back. Irrelevant for staging. +**Relevant at promotion:** if the production reverse proxy lives on the +WireGuard side, this topology will not serve it unchanged. Recorded now so it +is not rediscovered later. + +--- + +## 3. Network + +### 3.1 Two bridges + +**REQ** — `vmbr0`, existing, unchanged. `enp3s0f0`, `10.0.0.12/24`. This is the +**management and egress** network: SSH, apt, PyPI, Gitea, Docker Hub. + +**REQ** — `vmbr1`, new, **no physical port**. Host address `10.20.0.1/24`. +This is the **service** network: proxy-to-application traffic and nothing else. + +A portless bridge needs no cable, no switch configuration, and no spare NIC. It +exists so the application container has no listener on the LAN at all: the app +binds `127.0.0.1` and its `vmbr1` address only, and the only route to it is +through the proxy. + +``` +# /etc/network/interfaces — append +auto vmbr1 +iface vmbr1 inet static + address 10.20.0.1/24 + bridge-ports none + bridge-stp off + bridge-fd 0 +``` + +### 3.2 The other three NICs stay unconfigured + +**REQ** — Do not bond, trunk, or bridge `enp3s0f1` or the second dual-port +adapter. Leave them down and cabled to nothing. + +This is deliberate. An active-backup bond on `vmbr0` would be defensible, but +it buys link redundancy on a staging host that is allowed to be down, at the +cost of a configuration that has to be understood by everyone who touches the +box afterwards. Not worth it here. + +What the spare ports are *for*, when a reason appears: a point-to-point link to +the production host for promotion transfers, and a dedicated management +interface if the LAN ever becomes untrusted. Neither is now. + +### 3.3 Addressing + +**ASSUMED** — Static, outside any DHCP pool. + +| Guest | `vmbr0` (management) | `vmbr1` (service) | +|---|---|---| +| `srv-b` host | `10.0.0.12/24` | `10.20.0.1/24` | +| CT 100 `mechcomp` | `10.0.0.20/24` | `10.20.0.10/24` | +| CT 101 `mcproxy` | `10.0.0.21/24` | `10.20.0.11/24` | + +**VERIFY** — Stage 2 must confirm `10.0.0.20` and `10.0.0.21` are unclaimed +immediately before creating each container, and abort if not: + +```bash +for ip in 10.0.0.20 10.0.0.21; do + if arping -D -I vmbr0 -c 3 -w 3 "$ip" >/dev/null 2>&1; then + echo " $ip free" + else + echo "ABORT: $ip answered ARP — address is in use" >&2; exit 1 + fi +done +``` + +The investigation deliberately did not scan the LAN, and this document does not +either. An ARP probe for two specific addresses at creation time is the correct +scope. `10.0.0.160` was seen as a neighbour and is avoided. + +**DISCOVER** — If the LAN has a DHCP server, report its pool range so these two +addresses can be confirmed outside it. If `arping` reports a conflict, stop and +ask rather than picking the next number. + +### 3.4 Names + +**REQ** — `dev.infra` has no resolver: the host points at `75.75.75.75`, which +cannot answer for it. Names resolve through `/etc/hosts` only. + +Stage 1 writes to `/etc/hosts` on the host, and stages 3 and 4 write the same +block into both containers: + +``` +10.20.0.10 mechanical-compiler.dev.infra mechcomp +10.20.0.11 mcproxy.dev.infra mcproxy +``` + +**REQ — do not stand up an internal DNS server.** Two containers and one host +do not justify one, and adding a resolver here would create a second source of +truth for names that production will not share. + +### 3.5 Firewall + +**ASSUMED** — The Proxmox firewall stays disabled, matching the observed +staging baseline. The service network is portless and the application has no +LAN listener, which is the containment that matters here. + +**REQ** — Production must revisit this. Recorded as a promotion checklist item +in §14, not as a permanent decision. + +--- + +## 4. Containers + +### 4.1 Roles + +| | CT 100 | CT 101 | +|---|---|---| +| Hostname | `mechcomp` | `mcproxy` | +| Role | Application and worker | Reverse proxy, TLS termination | +| Unprivileged | yes | yes | +| Features | `nesting=1,keyctl=1` | none | + +**VERIFY** — Confirm both IDs with `pvesh get /cluster/nextid` and `pct list` +immediately before creation. `100` and `101` are candidates from a host that +had no guests, not reservations. + +### 4.2 Sizing + +The host has 24 threads and 31 GiB with nothing else on it. Sizing reflects +that rather than a cautious default. + +| | CT 100 | CT 101 | Host remaining | +|---|---|---|---| +| vCPU | 16 | 2 | 24 total, soft limits | +| RAM | 16 GiB | 1 GiB | ~14 GiB | +| Swap | 4 GiB | 512 MiB | — | +| rootfs | 40 GiB | 8 GiB | on `local-lvm` | +| `mp0` | 60 GiB at `/var/lib/mechcomp` | — | — | + +Thin provisioning means `mp0` consumes only what is written. Committed total is +108 GiB of 166.9 GiB — generous but not over-committed, which matters because +thin-pool exhaustion is an ugly failure mode. + +**REQ** — `mp0` is a real Proxmox mount point with `backup=0`. See §9.2 for why +the backup flag is off, and §12 for the check that proves the mount is real. + +### 4.3 Base OS + +**REQ** — Debian 12 bookworm from +`local:vztmpl/debian-12-standard_12.12-1_amd64.tar.zst`, both containers. + +Chosen over Debian 13 and Ubuntu because it matches YunoHost 12 stable, ships +Python 3.11 (the best OCP/CadQuery wheel coverage), and is already present on +the host. + +**VERIFY** — Stage 2 confirms the template is present, and downloads it only if +absent: + +```bash +pveam list local | grep -q "${PVE_TEMPLATE##*/}" || pveam download local "${PVE_TEMPLATE##*/}" +``` + +**REQ** — `America/Chicago`, `en_US.UTF-8` in both containers, matching the host. + +--- + +## 5. TLS and service names + +### 5.1 Staging + +**REQ** — Internal name `mechanical-compiler.dev.infra`, certificate issued by +the **Kane County Civic Infrastructure CA**. + +**REQ** — The CA root is installed into both containers' trust stores and must +also be trusted on the workstation used to browse staging. + +**REQ — the public FQDN is not used on staging.** No DNS record for +`mechanical-compiler.manufacturing.kane-il.us` may point at `srv-b`. No Let's +Encrypt certificate is issued for staging: it consumes rate limit against a +name that must be clean when production needs it, and it would put a +development box on the public internet. + +### 5.2 Production, reserved + +| | Value | +|---|---| +| FQDN | `mechanical-compiler.manufacturing.kane-il.us` | +| Zone | Labels within `kane-il.us`. **Not delegated.** No new DS record. | +| Certificate | Let's Encrypt, **per-name, no wildcard** | +| Challenge | **ASSUMED** HTTP-01 on the production proxy. Match existing practice if it is DNS-01; do not run both. | + +`manufacturing.kane-il.us` is a namespace, not a service. Nothing is served +from it. `kane-il.us` itself stays reserved for cross-cutting infrastructure — +proxy, mail, directory, certificates — and must not acquire project records. + +The leaf names the service, not the discipline, because the parent already +carries the discipline and the namespace must hold siblings: + +``` +mechanical-compiler.manufacturing.kane-il.us this application +field-metrology.manufacturing.kane-il.us Utility One, later +stock-processing.manufacturing.kane-il.us Utility Two, later +``` + +### 5.3 Three names, never conflated + +| Name | Scope | Value | +|---|---|---| +| Service FQDN | What users type | staging and production differ; see above | +| Infrastructure token | System user, paths, units, YunoHost app id | `mechcomp` | +| Project name | Documents, README, catalog entry | Mechanical Compiler | + +The token stays `mechcomp` in both environments. YunoHost app ids accept +lowercase alphanumerics and underscores only, so a hyphenated FQDN could not +serve as one anyway. + +--- + +## 6. Reverse proxy (CT 101) + +**REQ** — NGINX. Terminates TLS on `vmbr0`, proxies to CT 100 over `vmbr1`. + +**REQ** — The `X-Forwarded-Proto` header is mandatory, not optional. Getting +this wrong is what broke Hubzilla sessions. + +```nginx +server { + listen 443 ssl; + listen [::]:443 ssl; + server_name mechanical-compiler.dev.infra; + + ssl_certificate /etc/ssl/mechcomp/server.crt; # Kane County CA + ssl_certificate_key /etc/ssl/mechcomp/server.key; + + location / { + proxy_pass http://10.20.0.10:8770; + proxy_http_version 1.1; + proxy_set_header Host $host; + proxy_set_header X-Real-IP $remote_addr; + proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for; + proxy_set_header X-Forwarded-Proto $scheme; + proxy_read_timeout 300s; + client_max_body_size 8m; + } +} + +server { + listen 80; + listen [::]:80; + server_name mechanical-compiler.dev.infra; + return 301 https://$host$request_uri; +} +``` + +**PREF** — `proxy_read_timeout 300s`. Compile jobs are asynchronous by design, +but the initial synchronous render can take tens of seconds on a cold cache on +this CPU, and the 60s default leaves no headroom. + +**REQ** — This vhost is the production vhost with two values changed: +`server_name` and the certificate paths. Keep it that way. Any divergence +beyond those two lines means staging has stopped testing production. + +--- + +## 7. Users, directories, permissions (CT 100) + +**REQ** — Dedicated system user, no login shell, no home of its own: + +``` +user:group mechcomp:mechcomp +shell /usr/sbin/nologin +``` + +**REQ** — Layout, matching what a YunoHost package would provision: + +| Path | Owner | Mode | Purpose | YunoHost equivalent | +|---|---|---|---|---| +| `/var/www/mechcomp` | `mechcomp:mechcomp` | `0750` | source and venv | `install_dir` | +| `/var/lib/mechcomp` | `mechcomp:mechcomp` | `0750` | data, on `mp0` | `data_dir` | +| `/var/log/mechcomp` | `mechcomp:mechcomp` | `0750` | logs | — | +| `/etc/mechcomp/mechcomp.env` | `root:mechcomp` | `0640` | configuration | `ynh_add_config` target | + +**REQ** — `/var/lib/mechcomp` subdivides as: + +``` +/var/lib/mechcomp/ +├── db/ SQLite database +├── artifacts/ generated SVG, STL, STEP, manifests +├── fixtures/ frozen acceptance oracles, read-mostly +└── cache/ safe to delete at any time +``` + +The `cache/` guarantee matters: anything not reproducible from `db/` plus +`artifacts/` must not live there. That is what keeps the backup story simple. + +**REQ** — `${ADMIN_USER}` gets an ordinary shell account in the `mechcomp` +group with `${ADMIN_SSH_KEY}` installed, in **both** containers. The service +user is never used interactively. + +--- + +## 8. Software (CT 100) + +### 8.1 System packages + +**REQ** + +``` +python3 python3-venv python3-dev +build-essential pkg-config +git curl ca-certificates +libgeos-dev # Shapely bundles GEOS in its wheel; headers make a + # source build possible if a wheel is unavailable +sqlite3 +``` + +**PREF** — `jq ripgrep tmux htop` + +**REQ — do not install** `openscad`, `openscad-nightly`, or any Qt or X11 +package. This is not a weight argument: the running application never invokes +OpenSCAD, and §8.2 confines it where it belongs. If OpenSCAD appears in the +container's package list, something has gone wrong architecturally. + +### 8.2 The pinned reference toolchain + +**REQ** — One Docker image with the exact toolchain the fixture oracle was +frozen against, and nothing else: + +```dockerfile +# tools/reference-toolchain/Dockerfile +FROM debian:12-slim +RUN apt-get update && apt-get install -y --no-install-recommends \ + openscad git ca-certificates \ + && rm -rf /var/lib/apt/lists/* +RUN git clone https://github.com/BelfrySCAD/BOSL2.git /BOSL2 \ + && git -C /BOSL2 checkout 92d697c2856de2fed93a33e858068589cefc2898 +``` + +Tag `mechcomp/reference-toolchain:8.0.0`. Its only job is to run +`make_fixtures.py` and prove the oracle still reproduces. This is why the +Debian revision skew does not matter: the container has no OpenSCAD, so its +version cannot drift. + +**Acceptance test** — must reproduce exactly: + +``` +sha256 of strap-beam-fixtures-8.0.0.json + = ddd0f1548379205dd0c652ec07285b0dae331e52ff0a0437005dfc6cddcc2cb2 +``` + +If it does not, stop and report. A drifting oracle is worse than no oracle. + +**DISCOVER** — Docker in an unprivileged LXC is usually fine on this kernel +with `overlay2`, but it is not guaranteed. Report: + +```bash +docker info --format '{{.Driver}} / {{.CgroupDriver}}' +docker run --rm debian:12-slim echo ok +``` + +`fuse-overlayfs` is the documented fallback. If both fail, say so — moving the +image build to a small VM on the same host is a minor change and preferable to +fighting the storage driver. + +### 8.3 Python + +**REQ** — Virtualenv at `/var/www/mechcomp/venv`. Debian 12 enforces PEP 668; +do not use `--break-system-packages` to get around it. + +**REQ** — Pinned, hash-checked dependencies in two files, both installed: + +| File | Contents | +|---|---| +| `requirements-base.txt` | shapely, fastapi, uvicorn, pydantic, sqlalchemy, jinja2, pytest, **pytest-xdist** | +| `requirements-cad.txt` | cadquery / build123d | + +Both install here. The split is not about disk; it is the mechanism that keeps +§1.1 honest. CI runs the suite twice — with both files, then with +`requirements-base.txt` alone — and the second run must pass. + +`pytest-xdist` is in the base set specifically so `-n auto` uses all 24 threads. + +**PREF** — `uv` for resolution and locking; `pip-tools` is a fine substitute. +Either way the lock file is committed. + +--- + +## 9. Services and backups + +### 9.1 Systemd units (CT 100) + +**REQ** — Two units, split along the control-plane / execution-plane boundary +Part II §16 requires: + +| Unit | Role | +|---|---| +| `mechcomp.service` | FastAPI/uvicorn. Serves the catalogue, accepts jobs, returns cached artifacts. **Never runs geometry.** | +| `mechcomp-worker.service` | Consumes the queue, runs generators, writes artifacts. | + +**REQ** — Both: + +```ini +[Service] +User=mechcomp +Group=mechcomp +EnvironmentFile=/etc/mechcomp/mechcomp.env +WorkingDirectory=/var/www/mechcomp +Restart=on-failure +RestartSec=5 + +NoNewPrivileges=true +PrivateTmp=true +ProtectSystem=strict +ProtectHome=true +ReadWritePaths=/var/lib/mechcomp /var/log/mechcomp +ProtectKernelTunables=true +ProtectControlGroups=true +RestrictSUIDSGID=true + +[Install] +WantedBy=multi-user.target +``` + +**REQ** — Worker only: + +```ini +MemoryMax=8G +CPUQuota=1200% +TimeoutStopSec=30 +``` + +A runaway geometry job must not take the container down. `MemoryMax` turning a +hang into a clean OOM-kill of one worker is the desired behaviour. + +**REQ** — Logging to journald. No application-managed rotation. + +**REQ** — Job queue is SQLite-backed and in-process. No Redis, no Celery, no +RabbitMQ. A table with a status column and a claim query is correct at this +scale and adds no services to either packaging target. It sits behind an +interface so it can be swapped if load ever justifies it. It will not. + +### 9.2 Three backup tiers + +| Tier | What | Where | Retention | +|---|---|---|---| +| Routine | `vzdump` of both container **rootfs** | `local` (`/var/lib/vz/dump`) | see below | +| Application | tar stream of the data set | `/var/lib/vz/mechcomp-app/` on `srv-b` | 14 daily | +| Gold | routine + application + manifest | 32 GB removable USB, LUKS | indefinite, offline | + +**REQ — `mp0` is excluded from `vzdump` (`backup=0`).** `/var/lib/vz` sits on +`pve-root`, which is 78 GiB, and a full root filesystem is how Proxmox hosts +become interesting. Excluding the 60 GiB data volume keeps each routine archive +to the rootfs alone — a few GiB compressed — and application data flows through +the tier below, which is where it belongs anyway. + +**REQ** — `vzdump` job: both containers, `snapshot` mode, `zstd`, daily at +02:30, storage `local`. + +**ASSUMED** — Initial retention `keep-last=3`. Deliberately conservative: no +`vzdump` has been taken on this host, so actual size is unmeasured. **DISCOVER** +— after the first successful run, report the archive size and I will set final +retention. With 70 GiB free, three archives of an unknown size is the safe +starting point. + +**REQ** — A `df` guard in the backup hook that refuses to run and alerts if +`/var/lib/vz` falls below 20 GiB free. + +**REQ** — Application backup is a single script in CT 100 emitting a tar stream +on stdout, pulled host-side: + +```bash +pct exec ${CT_ID_APP} -- /usr/local/bin/mechcomp-backup --stdout \ + > /var/lib/vz/mechcomp-app/mechcomp-$(date +%F).tar +``` + +Covering exactly: + +``` +/etc/mechcomp/mechcomp.env +/var/lib/mechcomp/db/ +/var/lib/mechcomp/artifacts/ +/var/lib/mechcomp/fixtures/ +``` + +`cache/` is excluded. `/var/www/mechcomp` is excluded because it is reproducible +from git plus the lock file. This script is, almost verbatim, the future +YunoHost `backup` script. + +**REQ — pull, never push.** The container has no access to any backup +destination. The alternative is tempting and simpler, but it hands a +network-facing container write access to the last line of defence. It also +avoids idmap ownership problems entirely. + +### 9.3 Gold archive + +**REQ** — 32 GB USB stick, **LUKS2 full-device encryption**, ext4 inside, +filesystem label `MC-GOLD-01`, mounted at `/mnt/gold` with `noauto`. + +Encryption is not optional: the archive contains `mechcomp.env`, which contains +`MECHCOMP_SECRET_KEY`, and a stick that travels between machines is exactly the +thing that gets lost. The passphrase lives in CIVICVS's password manager and +never in the repository. Encrypting the device rather than the files means the +later copy to the 3+ TB disk inherits the protection if the LUKS image is +copied as a block image. + +ext4 rather than exFAT: no 4 GiB file cap, ownership preserved, journalled. +**DISCOVER** — if the stick must ever be read from Windows or macOS, say so and +this becomes exFAT with `age`-encrypted archives instead. + +**REQ** — Archive layout, one directory per mint: + +``` +/mnt/gold/-/ +├── vzdump-lxc-100-*.tar.zst +├── vzdump-lxc-101-*.tar.zst +├── mechcomp-app-*.tar +├── site.env (the FILL values, not secrets) +├── MANIFEST.txt git tag, revisions, host, PVE version +└── SHA256SUMS +``` + +**REQ** — `gold-archive.sh` verifies `SHA256SUMS` after writing and before +unmounting. An archive that has not been verified has not been made. + +**ASSUMED** — A gold archive is minted on **every tagged release** and +**immediately before any promotion or host migration**. This ties the tier to +the existing tag discipline rather than to a calendar. + +**DISCOVER** — Where the 3+ TB USB disk is attached now that `annales` is out +of scope. Until that is known, stick-to-disk copying is a documented manual +step, not a scripted one. + +### 9.4 Disk health monitoring + +**REQ** — Configure `smartd` for the four RAID members using the already +installed `smartmontools`. No third-party HPE repository is added, and `ssacli` +is not installed: the investigation showed the kernel and `smartctl -d cciss,N` +already provide what is needed. + +``` +# /etc/smartd.conf +/dev/sda -d cciss,0 -a -m root -M exec /usr/share/smartmontools/smartd-runner +/dev/sda -d cciss,1 -a -m root -M exec /usr/share/smartmontools/smartd-runner +/dev/sda -d cciss,2 -a -m root -M exec /usr/share/smartmontools/smartd-runner +/dev/sda -d cciss,3 -a -m root -M exec /usr/share/smartmontools/smartd-runner +``` + +These are 2010-vintage spinning disks in a RAID 1+0 holding both the staging +environment and its routine backups. Knowing about a failure before the second +one is worth the four lines. + +**DISCOVER** — Whether Proxmox notifications can relay through the existing +mail infrastructure. If not, `smartd` and `vzdump` failures land in local root +mail only, which nobody reads — report this rather than leaving it silently +broken. + +--- + +## 10. Configuration contract + +**REQ** — `/etc/mechcomp/mechcomp.env` in CT 100, and nothing else, configures +the application: + +```bash +MECHCOMP_ENV=staging # staging | production +MECHCOMP_BIND=10.20.0.10 # service network only; never 0.0.0.0 +MECHCOMP_PORT=8770 +MECHCOMP_BASE_URL=https://mechanical-compiler.dev.infra +MECHCOMP_DATA_DIR=/var/lib/mechcomp +MECHCOMP_LOG_LEVEL=info +MECHCOMP_DB_URL=sqlite:////var/lib/mechcomp/db/mechcomp.sqlite3 +MECHCOMP_WORKER_CONCURRENCY=4 +MECHCOMP_CAD_BACKEND=none # none | cadquery — 'none' until the 3D path exists +MECHCOMP_ARTIFACT_RETENTION_DAYS=30 +MECHCOMP_SECRET_KEY= # generated by stage 3; see §11.4 +``` + +**REQ** — The application fails loudly at startup if a required variable is +missing, and never falls back to a compiled-in default path. That single rule +is most of what makes both packaging targets work unchanged. + +**REQ** — `MECHCOMP_BIND` is the service-network address, plus `127.0.0.1`. +Binding `0.0.0.0` would defeat the point of the second bridge. + +Promotion changes `MECHCOMP_ENV`, `MECHCOMP_BIND` and `MECHCOMP_BASE_URL`. +Nothing else. + +--- + +## 11. Provisioning scripts + +Everything above is a specification for scripts, not a runbook for a human. + +### 11.1 Repository + +**ASSUMED** — A separate repository, `TheRON/mechanical-compiler-infra`. + +Provisioning scripts do not belong in the application repository. A YunoHost +reviewer eventually reads that tree, and Proxmox tooling beside the app is +exactly the "non-standard directory architecture" the catalog flags. The two +also version independently, which they will. + +If you prefer one repository, put the scripts in `deploy/` and add that path to +the packaging exclusion list from the first commit — not after they have history. + +### 11.2 Four stages + +| Script | Runs as | Where | Does | +|---|---|---|---| +| `stage1-host.sh` | root | `srv-b` | `vmbr1`, `/etc/hosts`, `smartd`, vzdump job, `/var/lib/vz/mechcomp-app` | +| `stage2-create-cts.sh` | root | `srv-b` | template check, ARP verify, `pct create` ×2, `mp0` with `backup=0`, first boot | +| `stage3-app.sh` | root | CT 100 | packages, user, layout, venv, Docker, units, config, secret | +| `stage4-proxy.sh` | root | CT 101 | nginx, CA cert install, vhost, `/etc/hosts` | + +**REQ** — The only interface between stages is `deploy/site.env`. Stage 2 +copies it into both containers; stages 3 and 4 source it. No other state crosses +a boundary. + +**REQ** — Stage 4 installs its own vhost. Unlike revision 3, the proxy is ours, +on this host, and there is no separate operator to hand a config to. + +### 11.3 Idempotency + +**REQ** — All four stages are safe to re-run. Create-if-absent, +converge-if-present, never destroy. A re-run against a healthy host is a no-op +that exits zero. + +This follows from "version, tag, release with discipline": a script you cannot +re-run is a script you cannot test, and a script you cannot test is ceremony, +not discipline. + +**REQ** — Every stage starts with `set -euo pipefail` and validates that each +**FILL** value in `site.env` is non-empty before touching anything. + +### 11.4 Secrets + +**REQ** — `MECHCOMP_SECRET_KEY` is never committed and never appears in +`site.env`. Stage 3 generates it on first run and refuses to overwrite: + +```bash +if ! grep -q '^MECHCOMP_SECRET_KEY=.\+' /etc/mechcomp/mechcomp.env 2>/dev/null; then + key=$(openssl rand -base64 32) + # write it; never log it +fi +``` + +**REQ** — `deploy/site.env` is in `.gitignore`. `deploy/site.env.example`, with +empty values, is committed. The LUKS passphrase is in neither. + +### 11.5 Versioning + +**ASSUMED** — The infrastructure repository carries its own semver, independent +of the application's `SB_REVISION`. They drift for unrelated reasons and tying +them would force meaningless bumps on both. + +**REQ** — Tag a release before each provisioning run and record the tag in the +run log and in the gold `MANIFEST.txt`. "Which version of the script built this +container" must be answerable six months from now. + +--- + +## 12. `deploy/site.env` + +```bash +# ---- Proxmox host (confirmed 2026-08-15) -------------------------------- +PVE_HOST=srv-b +PVE_STORAGE=local-lvm +PVE_DATA_STORAGE=local-lvm +PVE_BACKUP_STORAGE=local +PVE_BRIDGE_MGMT=vmbr0 +PVE_BRIDGE_SVC=vmbr1 +PVE_TEMPLATE=local:vztmpl/debian-12-standard_12.12-1_amd64.tar.zst + +# ---- Application container --------------------------------------------- +CT_ID_APP=100 # VERIFY before create +CT_HOSTNAME_APP=mechcomp +CT_IP_APP_MGMT=10.0.0.20/24 # VERIFY free by ARP before create +CT_IP_APP_SVC=10.20.0.10/24 +CT_CORES_APP=16 +CT_RAM_APP=16384 +CT_SWAP_APP=4096 +CT_ROOTFS_APP=40 +CT_MP0_APP=60 + +# ---- Proxy container --------------------------------------------------- +CT_ID_PROXY=101 # VERIFY before create +CT_HOSTNAME_PROXY=mcproxy +CT_IP_PROXY_MGMT=10.0.0.21/24 # VERIFY free by ARP before create +CT_IP_PROXY_SVC=10.20.0.11/24 +CT_CORES_PROXY=2 +CT_RAM_PROXY=1024 +CT_SWAP_PROXY=512 +CT_ROOTFS_PROXY=8 + +# ---- Network ----------------------------------------------------------- +CT_GATEWAY=10.0.0.1 +SVC_NET=10.20.0.0/24 +SVC_HOST_IP=10.20.0.1 + +# ---- Service ----------------------------------------------------------- +MECHCOMP_ENV=staging +SERVICE_FQDN=mechanical-compiler.dev.infra +APP_PORT=8770 +TLS_CA=kane-county-civic-infrastructure + +# ---- Identity ---------------------------------------------------------- +ADMIN_USER= # FILL +ADMIN_SSH_KEY= # FILL + +# ---- Backup ------------------------------------------------------------ +VZDUMP_SCHEDULE="02:30" +VZDUMP_KEEP_LAST=3 # revisit after first size measurement +APP_BACKUP_DIR=/var/lib/vz/mechcomp-app +APP_BACKUP_KEEP=14 +GOLD_LABEL=MC-GOLD-01 +GOLD_MOUNT=/mnt/gold + +# ---- Reserved for production, not used on srv-b ------------------------ +# PRODUCTION_FQDN=mechanical-compiler.manufacturing.kane-il.us +``` + +`CT_WG_ALIAS` is gone. The host routes and MASQUERADEs `10.0.0.0/24` out `wg0`, +so containers reach the WireGuard network by their ordinary LAN address, and no +per-container alias mechanism exists to configure. + +--- + +## 13. Constraints the application code will follow + +Recorded so the environment is not built in a way that quietly permits their +violation. + +1. No hardcoded absolute paths. Everything derives from `MECHCOMP_DATA_DIR`. +2. No writes outside `MECHCOMP_DATA_DIR`. Enforced by `ProtectSystem=strict`. +3. No dependency on systemd from application code. Signal handling only. +4. The app listens on a TCP port, on the service network and loopback only. +5. No OpenSCAD, Qt or X11 dependency in the running application. +6. The 3D backend is reached only through an interface in `src/mechcomp/cad/`. + No other module imports CadQuery or OCP directly. +7. Schema migrations are explicit and forward-only. +8. Artifacts are addressed by content hash of **inputs** — parameter set plus + generator revision — never by hash of output bytes. Mesh output is not + reproducible across toolchain versions; input hashing keeps a qualification + valid across an upgrade that did not change the geometry. +9. Every generated artifact carries the generator revision that produced it. +10. The test suite passes with `requirements-cad.txt` uninstalled. +11. AGPL-3.0 §13 requires network users be offered the source. The web tier + carries a visible source link to the Gitea repository. This is a licence + obligation, not a nicety. + +--- + +## 14. Acceptance checklist + +Run and report. This is also the first draft of the YunoHost `tests.toml` +criteria. + +```bash +# ---- host ---- +pveversion | head -1 +ip -br address show vmbr1 # expect 10.20.0.1/24 +pct list # expect 100 and 101, both running +grep -c mechanical-compiler.dev.infra /etc/hosts +systemctl is-active smartd +cat /etc/pve/jobs.cfg | grep -c vzdump # expect >= 1 +pct config 100 | grep mp0 # must contain backup=0 + +# ---- CT 100 ---- +pct exec 100 -- bash -c ' + cat /etc/debian_version + python3 --version + timedatectl show -p Timezone --value + id mechcomp + findmnt /var/lib/mechcomp # must be a mount, not rootfs + stat -c "%U:%G %a" /etc/mechcomp/mechcomp.env + dpkg -l | grep -Ei "openscad|libqt|xserver" || echo "clean" + ip -br address | grep 10.20.0.10 + ss -lntp | grep 8770 # must NOT bind 0.0.0.0 + grep -c "^MECHCOMP_SECRET_KEY=.\+" /etc/mechcomp/mechcomp.env + /var/www/mechcomp/venv/bin/python -c "import shapely; print(shapely.__version__)" + /var/www/mechcomp/venv/bin/python -c "import cadquery; print(cadquery.__version__)" + docker info --format "{{.Driver}}" + docker run --rm mechcomp/reference-toolchain:8.0.0 openscad --version +' + +# ---- egress from CT 100 ---- +pct exec 100 -- bash -c ' + for h in deb.debian.org pypi.org files.pythonhosted.org gitea.barternetwork.us; do + curl -sS -o /dev/null -w "$h %{http_code}\n" "https://$h" + done + git ls-remote https://gitea.barternetwork.us/TheRON/mechanical-compiler HEAD | head -1 +' + +# ---- CT 101 and end to end ---- +pct exec 101 -- nginx -t +pct exec 101 -- curl -sS -o /dev/null -w '%{http_code} %{ssl_verify_result}\n' \ + https://mechanical-compiler.dev.infra/ + +# ---- forwarded-proto reaches the app ---- +pct exec 101 -- curl -sS https://mechanical-compiler.dev.infra/_debug/headers \ + | grep -i x-forwarded-proto # expect https + +# ---- idempotency: second run is a clean no-op ---- +./stage1-host.sh && ./stage2-create-cts.sh && echo "re-run clean" + +# ---- backup round trip ---- +vzdump 100 --storage local --mode snapshot --compress zstd +ls -lh /var/lib/vz/dump/ # REPORT the size +pct exec 100 -- /usr/local/bin/mechcomp-backup --stdout | wc -c +``` + +### Promotion checklist, for later + +Not now, but recorded so it is not invented under pressure: production must +revisit the Proxmox firewall (§3.5), issue a Let's Encrypt certificate for the +public FQDN (§5.2), confirm the reverse proxy is not on the WireGuard side +(§2.2), and change exactly three variables in `mechcomp.env` (§10). + +--- + +## 15. Deliberately out of scope + +- Hubzilla addon development, or any change to existing Hubzilla containers +- Federation, identity, or qualification storage +- PostgreSQL — SQLite is sufficient and will remain so for a long time +- Redis, Celery, or any external queue +- CI runners — the suite runs locally until there is a reason otherwise +- Monitoring beyond journald, smartd, and Proxmox's own metrics +- NIC bonding, VLANs, or any further network configuration (§3.2) +- An internal DNS server (§3.4) +- The YunoHost package and the Docker image themselves + +Each has a natural moment. None of them is now. + +--- + +## 16. Every assumption in one place + +| § | Assumption | If wrong | +|---|---|---| +| 3.3 | `10.0.0.20` and `.21` are free and outside any DHCP pool | ARP verify aborts stage 2; supply replacements | +| 3.3 | Static addressing, no DHCP reservation needed | Reserve in DHCP instead; addresses unchanged | +| 3.5 | Proxmox firewall stays disabled in staging | Enable per-container; production revisits regardless | +| 4.1 | CT IDs 100 and 101 | Verified at create time; any free pair works | +| 4.2 | 16 vCPU / 16 GiB for the app container | Lower freely; nothing depends on it | +| 5.1 | Kane County CA issues the staging certificate | A self-signed cert works; trust store step changes | +| 5.2 | Production uses HTTP-01 | Switch to DNS-01 if that is existing practice; never both | +| 9.2 | `keep-last=3` initially | Raise once the first archive size is measured | +| 9.3 | LUKS2 on the gold stick, ext4 inside | exFAT plus `age`-encrypted files if cross-OS reading is needed | +| 9.3 | Gold minted per tagged release and before promotion | Any trigger works; this one ties to existing discipline | +| 11.1 | Separate `mechanical-compiler-infra` repository | Use `deploy/` in the app repo, excluded from packaging | +| 11.3 | All stages idempotent | Only matters if you want create-once semantics; I would argue against | +| 11.5 | Infra versions independently of `SB_REVISION` | Tie them together if you prefer one release train | + +--- + +## 17. Provenance disclosure + +This project's documents, code and roadmap are LLM-generated under human +direction. That is stated plainly here, in the README, and in any eventual +catalog submission. We do not obscure it. + +**The YunoHost policy is a quality bar with a disclosure requirement, not a +ban.** Revision 1 overstated it. The catalog repository rejects generated +packages *that do not follow the `example_ynh` template*, citing verbose code, +hallucinated helpers and non-standard directory architectures — and then +explicitly permits AI use provided the maintainer is transparent and can +explain every line. The failure mode guarded against is sprawl, not provenance. +The defence is discipline: minimal divergence from the template, no invented +helpers, no extra files. + +**Package provenance is not upstream provenance.** The policy text concerns the +`_ynh` repository — a few hundred lines of shell and one TOML file. Nobody +audits whether an upstream Rails application was written by a human. The +reviewable surface is small. + +**Disclosure is right; headlining it is a tactical mistake.** Part II §29 holds +that early use cases should demonstrate the system's distinct value. If +provenance becomes the pitch, the project is judged on that axis rather than on +whether it makes distributed manufacturing capacity legible. State it clearly +in the README; keep the pitch about manufacturing. The argument is won by an +artifact a reviewer finds shorter and more conventional than average, not by +the claim attached to it. + +--- + +## 18. The packaging path, when it comes + +**DISCOVER** — Unknown territory, to be documented as it is walked rather than +researched ahead: `manifest.toml` v2 with helpers 2.1, the `scripts/` set +(install, remove, upgrade, backup, restore, change_url), `conf/` templates, +`tests.toml`, and `package_check` levels. + +Two orienting notes. The useful comparison for acceptable weight is +`paperless-ngx_ynh` or `fab-manager_ynh`, not a hello-world package. And per +§17, the package is drafted against `example_ynh` and reviewed by CIVICVS as +his own work — anything in it he would not defend in a review thread does not +ship. + +--- + +## 19. First work item after acceptance + +Port `sb-geom` to Shapely and get `pytest -n auto` green against the 123 frozen +cases in `strap-beam-fixtures-8.0.0.json`. The ten rejected cases are part of +the contract: a port that accepts them is wrong.