Files
mechanical-compiler/docs/ENVIRONMENT.md
T
TheRON 318a2e30dc Provisioning specifications
This specifies a **staging environment** on `srv-b`, plus the promotion path to
a production host that does not yet exist.
2026-08-16 11:59:08 -04:00

41 KiB
Raw Blame History

ENVIRONMENT.md

Provisioning specification for the Mechanical Compiler development and staging environment.

Revision 4 (2026-08-15)
Supersedes Revisions 1, 2 and 3
Basis ENVIRONMENT_INVESTIGATION_REPORT_2026-08-15.md (host srv-b)
Audience The assistant or operator provisioning the containers
Author role Written by the developer who will work inside this environment
Status All decisions taken. Ready to execute.
Repository https://gitea.barternetwork.us/TheRON/mechanical-compiler
Licence AGPL-3.0-or-later (committed at e85c4f4e)

0. How to use this document

This specifies a staging environment on srv-b, plus the promotion path to a production host that does not yet exist.

Everything is scripted. Nothing is provisioned by hand.

Markers

Marker Meaning
REQ Required. The work is blocked or wrong without it.
PREF Preferred. Substitute freely, but report what you substituted.
ASSUMED Decided by the author without confirmation. Defensible, but a guess. All are listed in §16.
FILL A site value that must be supplied. Goes in deploy/site.env and nowhere else.
VERIFY Must be checked at run time rather than trusted from this document.
DISCOVER Genuinely unknown. Answer by doing, then report.

If you are the provisioning assistant

Four requests. FILL values live in one file and are referenced from there — no literals scattered through scripts. VERIFY steps run inside the scripts, not as a pre-flight a human might skip. Report DISCOVER outcomes verbatim, including failures. And if a REQ item cannot be satisfied, say so rather than working around it: the stated reason usually matters more than the mechanism.

What changed in revision 4

# Change Source
1 Host is srv-b, not annales CHG-001
2 Staging and production are separate roles, on separate hosts CHG-002
3 srv-b is a standalone node; promotion is export/restore, not cluster migration CHG-003
4 Template 12.12-1, not 12.7-1 CHG-004
5 Three-tier backup architecture; 32 GB removable gold media CHG-005
6 Backup schedule defined explicitly; none exists to inherit CHG-006
7 CT_WG_ALIAS removed CHG-007
8 PROXY_HOST not inherited; staging gets its own proxy container CHG-008
9 Two containers, not one: application and reverse proxy this revision
10 Second bridge vmbr1 isolates service traffic from management traffic this revision
11 Container sizing raised substantially to use the host this revision
12 Staging is internal-only; the public FQDN is reserved for production this revision

1. Five decisions that shape everything below

1.1 The 2D path and the 3D path are separable

The catalogue's product is a 2D cross-section rendered to SVG. Shapely plus our own code covers that completely: region algebra, offsets, distance queries, output. CadQuery/OCCT is needed only for STEP and mesh export.

These stay separate — separate requirements files, separate code paths, and a CI job that runs the whole suite with the CAD dependencies absent. Three reasons, none of which is packaging politics:

  1. Part II §16 requires the control-plane / execution-plane separation regardless of how the app is distributed.
  2. A slow OCCT import must never land in the request path for a page that only draws a cross-section.
  3. Test-suite speed. Shapely-only tests run in seconds, which is the difference between running the 123 frozen fixtures on every save and only before commit.

Retained correction. Revision 1 also justified this as a defence against YunoHost's "resource-hungry" criterion. That was overcalibrated — fab-manager_ynh is in the catalog, paperless-ngx_ynh declares nine apt dependencies including postgresql and redis, and the criterion reads "resource-hungry compared to their features." CadQuery is a default-installed dependency. The boundary stays; the apology goes.

1.2 YunoHost and Docker are parallel targets, not sequential ones

YunoHost apps install natively — apt, a venv, systemd, nginx — and the project does not want Docker inside YunoHost. Docker is a separate distribution channel, not a stepping stone to the catalog. Neither blocks the other.

1.3 Staging mirrors production topology, not production exposure

srv-b runs two containers because production will: an application container that never terminates TLS, and a reverse proxy that does. A single container serving directly would never exercise the X-Forwarded-Proto path — precisely the class of bug that cost us time on Hubzilla's session logout.

Staging is not publicly reachable and does not use the public FQDN. It uses the Kane County Civic Infrastructure CA on an internal name. The public name and Let's Encrypt belong to production, which forces the promotion path to be exercised rather than assumed.

1.4 Directory layout mirrors YunoHost conventions from the start

The paths in §7 are what a YunoHost package would provision anyway (install_dir, data_dir, a dedicated system user, a dynamically assigned port). The eventual manifest.toml resource block then describes what already exists.

1.5 All configuration comes from the environment, never from the filesystem

No path, port, URL, or secret may be hardcoded or discovered by convention. One env file per container; the same variables become ENV in a Dockerfile and ynh_add_config substitutions in a package.

The check that this holds: moving the service to a different hostname is a one-variable change. It currently is — which is why the staging/production split below costs almost nothing.


2. Host baseline (confirmed, not to be re-verified)

Item Value
Hostname / FQDN srv-b / srv-b.dev.infra
Role Development and staging. Standalone node, never clustered.
Platform Proxmox VE 8.4.0 on Debian 12, kernel 6.8.12-9-pve
Hardware HP ProLiant DL360 G7, firmware P68
CPU 2 × Xeon X5650, 24 logical CPUs
RAM / swap 31 GiB / 8 GiB
Storage HP Smart Array P410i, 4 × EG0146FAWHU, RAID 1+0, all SMART OK
local directory, /var/lib/vz, ~70 GiB free — iso, vztmpl, backup
local-lvm LVM-thin pve/data, 166.9 GiB, 0% used — rootdir, images
LAN vmbr0 on enp3s0f0, 10.0.0.12/24, gateway 10.0.0.1
WireGuard wg0 10.110.0.12/32, peer wg-pk.civicus.us:51820, allowed 10.110.0.0/22
Routing forwarding on, MASQUERADE 10.0.0.0/24 -o wg0, both persistent
Resolver 75.75.75.75, search dev.infra
Egress Direct. No proxy, no allowlist. All required hosts reachable.
PVE firewall disabled/running
Guests None. Next ID 100.
Template local:vztmpl/debian-12-standard_12.12-1_amd64.tar.zst

2.1 Two hardware facts worth knowing before anyone is surprised

The X5650 is Westmere: no AVX, and single-thread performance is modest. Geometry solvers are single-threaded, so regenerating the fixture oracle will take noticeably longer here than on a modern laptop. That is expected, not a fault. All compiled wheels target an SSE2 baseline, so nothing breaks.

Against that, 24 threads is a great deal of parallel test capacity. The test suite runs under pytest-xdist with -n auto, which is where the box earns its keep.

2.2 The WireGuard path is outbound only

MASQUERADE means containers reach 10.110.0.0/22, but nothing on the WireGuard network can initiate a connection back. Irrelevant for staging. Relevant at promotion: if the production reverse proxy lives on the WireGuard side, this topology will not serve it unchanged. Recorded now so it is not rediscovered later.


3. Network

3.1 Two bridges

REQ — vmbr0, existing, unchanged. enp3s0f0, 10.0.0.12/24. This is the management and egress network: SSH, apt, PyPI, Gitea, Docker Hub.

REQ — vmbr1, new, no physical port. Host address 10.20.0.1/24. This is the service network: proxy-to-application traffic and nothing else.

A portless bridge needs no cable, no switch configuration, and no spare NIC. It exists so the application container has no listener on the LAN at all: the app binds 127.0.0.1 and its vmbr1 address only, and the only route to it is through the proxy.

# /etc/network/interfaces — append
auto vmbr1
iface vmbr1 inet static
    address  10.20.0.1/24
    bridge-ports none
    bridge-stp off
    bridge-fd 0

3.2 The other three NICs stay unconfigured

REQ — Do not bond, trunk, or bridge enp3s0f1 or the second dual-port adapter. Leave them down and cabled to nothing.

This is deliberate. An active-backup bond on vmbr0 would be defensible, but it buys link redundancy on a staging host that is allowed to be down, at the cost of a configuration that has to be understood by everyone who touches the box afterwards. Not worth it here.

What the spare ports are for, when a reason appears: a point-to-point link to the production host for promotion transfers, and a dedicated management interface if the LAN ever becomes untrusted. Neither is now.

3.3 Addressing

ASSUMED — Static, outside any DHCP pool.

Guest vmbr0 (management) vmbr1 (service)
srv-b host 10.0.0.12/24 10.20.0.1/24
CT 100 mechcomp 10.0.0.20/24 10.20.0.10/24
CT 101 mcproxy 10.0.0.21/24 10.20.0.11/24

VERIFY — Stage 2 must confirm 10.0.0.20 and 10.0.0.21 are unclaimed immediately before creating each container, and abort if not:

for ip in 10.0.0.20 10.0.0.21; do
    if arping -D -I vmbr0 -c 3 -w 3 "$ip" >/dev/null 2>&1; then
        echo "  $ip free"
    else
        echo "ABORT: $ip answered ARP — address is in use" >&2; exit 1
    fi
done

The investigation deliberately did not scan the LAN, and this document does not either. An ARP probe for two specific addresses at creation time is the correct scope. 10.0.0.160 was seen as a neighbour and is avoided.

DISCOVER — If the LAN has a DHCP server, report its pool range so these two addresses can be confirmed outside it. If arping reports a conflict, stop and ask rather than picking the next number.

3.4 Names

REQ — dev.infra has no resolver: the host points at 75.75.75.75, which cannot answer for it. Names resolve through /etc/hosts only.

Stage 1 writes to /etc/hosts on the host, and stages 3 and 4 write the same block into both containers:

10.20.0.10   mechanical-compiler.dev.infra mechcomp
10.20.0.11   mcproxy.dev.infra mcproxy

REQ — do not stand up an internal DNS server. Two containers and one host do not justify one, and adding a resolver here would create a second source of truth for names that production will not share.

3.5 Firewall

ASSUMED — The Proxmox firewall stays disabled, matching the observed staging baseline. The service network is portless and the application has no LAN listener, which is the containment that matters here.

REQ — Production must revisit this. Recorded as a promotion checklist item in §14, not as a permanent decision.


4. Containers

4.1 Roles

CT 100 CT 101
Hostname mechcomp mcproxy
Role Application and worker Reverse proxy, TLS termination
Unprivileged yes yes
Features nesting=1,keyctl=1 none

VERIFY — Confirm both IDs with pvesh get /cluster/nextid and pct list immediately before creation. 100 and 101 are candidates from a host that had no guests, not reservations.

4.2 Sizing

The host has 24 threads and 31 GiB with nothing else on it. Sizing reflects that rather than a cautious default.

CT 100 CT 101 Host remaining
vCPU 16 2 24 total, soft limits
RAM 16 GiB 1 GiB ~14 GiB
Swap 4 GiB 512 MiB —
rootfs 40 GiB 8 GiB on local-lvm
mp0 60 GiB at /var/lib/mechcomp — —

Thin provisioning means mp0 consumes only what is written. Committed total is 108 GiB of 166.9 GiB — generous but not over-committed, which matters because thin-pool exhaustion is an ugly failure mode.

REQ — mp0 is a real Proxmox mount point with backup=0. See §9.2 for why the backup flag is off, and §12 for the check that proves the mount is real.

4.3 Base OS

REQ — Debian 12 bookworm from local:vztmpl/debian-12-standard_12.12-1_amd64.tar.zst, both containers.

Chosen over Debian 13 and Ubuntu because it matches YunoHost 12 stable, ships Python 3.11 (the best OCP/CadQuery wheel coverage), and is already present on the host.

VERIFY — Stage 2 confirms the template is present, and downloads it only if absent:

pveam list local | grep -q "${PVE_TEMPLATE##*/}" || pveam download local "${PVE_TEMPLATE##*/}"

REQ — America/Chicago, en_US.UTF-8 in both containers, matching the host.


5. TLS and service names

5.1 Staging

REQ — Internal name mechanical-compiler.dev.infra, certificate issued by the Kane County Civic Infrastructure CA.

REQ — The CA root is installed into both containers' trust stores and must also be trusted on the workstation used to browse staging.

REQ — the public FQDN is not used on staging. No DNS record for mechanical-compiler.manufacturing.kane-il.us may point at srv-b. No Let's Encrypt certificate is issued for staging: it consumes rate limit against a name that must be clean when production needs it, and it would put a development box on the public internet.

5.2 Production, reserved

Value
FQDN mechanical-compiler.manufacturing.kane-il.us
Zone Labels within kane-il.us. Not delegated. No new DS record.
Certificate Let's Encrypt, per-name, no wildcard
Challenge ASSUMED HTTP-01 on the production proxy. Match existing practice if it is DNS-01; do not run both.

manufacturing.kane-il.us is a namespace, not a service. Nothing is served from it. kane-il.us itself stays reserved for cross-cutting infrastructure — proxy, mail, directory, certificates — and must not acquire project records.

The leaf names the service, not the discipline, because the parent already carries the discipline and the namespace must hold siblings:

mechanical-compiler.manufacturing.kane-il.us    this application
field-metrology.manufacturing.kane-il.us        Utility One, later
stock-processing.manufacturing.kane-il.us       Utility Two, later

5.3 Three names, never conflated

Name Scope Value
Service FQDN What users type staging and production differ; see above
Infrastructure token System user, paths, units, YunoHost app id mechcomp
Project name Documents, README, catalog entry Mechanical Compiler

The token stays mechcomp in both environments. YunoHost app ids accept lowercase alphanumerics and underscores only, so a hyphenated FQDN could not serve as one anyway.


6. Reverse proxy (CT 101)

REQ — NGINX. Terminates TLS on vmbr0, proxies to CT 100 over vmbr1.

REQ — The X-Forwarded-Proto header is mandatory, not optional. Getting this wrong is what broke Hubzilla sessions.

server {
    listen 443 ssl;
    listen [::]:443 ssl;
    server_name mechanical-compiler.dev.infra;

    ssl_certificate     /etc/ssl/mechcomp/server.crt;   # Kane County CA
    ssl_certificate_key /etc/ssl/mechcomp/server.key;

    location / {
        proxy_pass         http://10.20.0.10:8770;
        proxy_http_version 1.1;
        proxy_set_header   Host              $host;
        proxy_set_header   X-Real-IP         $remote_addr;
        proxy_set_header   X-Forwarded-For   $proxy_add_x_forwarded_for;
        proxy_set_header   X-Forwarded-Proto $scheme;
        proxy_read_timeout 300s;
        client_max_body_size 8m;
    }
}

server {
    listen 80;
    listen [::]:80;
    server_name mechanical-compiler.dev.infra;
    return 301 https://$host$request_uri;
}

PREF — proxy_read_timeout 300s. Compile jobs are asynchronous by design, but the initial synchronous render can take tens of seconds on a cold cache on this CPU, and the 60s default leaves no headroom.

REQ — This vhost is the production vhost with two values changed: server_name and the certificate paths. Keep it that way. Any divergence beyond those two lines means staging has stopped testing production.


7. Users, directories, permissions (CT 100)

REQ — Dedicated system user, no login shell, no home of its own:

user:group   mechcomp:mechcomp
shell        /usr/sbin/nologin

REQ — Layout, matching what a YunoHost package would provision:

Path Owner Mode Purpose YunoHost equivalent
/var/www/mechcomp mechcomp:mechcomp 0750 source and venv install_dir
/var/lib/mechcomp mechcomp:mechcomp 0750 data, on mp0 data_dir
/var/log/mechcomp mechcomp:mechcomp 0750 logs —
/etc/mechcomp/mechcomp.env root:mechcomp 0640 configuration ynh_add_config target

REQ — /var/lib/mechcomp subdivides as:

/var/lib/mechcomp/
├── db/           SQLite database
├── artifacts/    generated SVG, STL, STEP, manifests
├── fixtures/     frozen acceptance oracles, read-mostly
└── cache/        safe to delete at any time

The cache/ guarantee matters: anything not reproducible from db/ plus artifacts/ must not live there. That is what keeps the backup story simple.

REQ — ${ADMIN_USER} gets an ordinary shell account in the mechcomp group with ${ADMIN_SSH_KEY} installed, in both containers. The service user is never used interactively.


8. Software (CT 100)

8.1 System packages

REQ

python3 python3-venv python3-dev
build-essential pkg-config
git curl ca-certificates
libgeos-dev            # Shapely bundles GEOS in its wheel; headers make a
                       # source build possible if a wheel is unavailable
sqlite3

PREF — jq ripgrep tmux htop

REQ — do not install openscad, openscad-nightly, or any Qt or X11 package. This is not a weight argument: the running application never invokes OpenSCAD, and §8.2 confines it where it belongs. If OpenSCAD appears in the container's package list, something has gone wrong architecturally.

8.2 The pinned reference toolchain

REQ — One Docker image with the exact toolchain the fixture oracle was frozen against, and nothing else:

# tools/reference-toolchain/Dockerfile
FROM debian:12-slim
RUN apt-get update && apt-get install -y --no-install-recommends \
        openscad git ca-certificates \
    && rm -rf /var/lib/apt/lists/*
RUN git clone https://github.com/BelfrySCAD/BOSL2.git /BOSL2 \
    && git -C /BOSL2 checkout 92d697c2856de2fed93a33e858068589cefc2898

Tag mechcomp/reference-toolchain:8.0.0. Its only job is to run make_fixtures.py and prove the oracle still reproduces. This is why the Debian revision skew does not matter: the container has no OpenSCAD, so its version cannot drift.

Acceptance test — must reproduce exactly:

sha256 of strap-beam-fixtures-8.0.0.json
  = ddd0f1548379205dd0c652ec07285b0dae331e52ff0a0437005dfc6cddcc2cb2

If it does not, stop and report. A drifting oracle is worse than no oracle.

DISCOVER — Docker in an unprivileged LXC is usually fine on this kernel with overlay2, but it is not guaranteed. Report:

docker info --format '{{.Driver}} / {{.CgroupDriver}}'
docker run --rm debian:12-slim echo ok

fuse-overlayfs is the documented fallback. If both fail, say so — moving the image build to a small VM on the same host is a minor change and preferable to fighting the storage driver.

8.3 Python

REQ — Virtualenv at /var/www/mechcomp/venv. Debian 12 enforces PEP 668; do not use --break-system-packages to get around it.

REQ — Pinned, hash-checked dependencies in two files, both installed:

File Contents
requirements-base.txt shapely, fastapi, uvicorn, pydantic, sqlalchemy, jinja2, pytest, pytest-xdist
requirements-cad.txt cadquery / build123d

Both install here. The split is not about disk; it is the mechanism that keeps §1.1 honest. CI runs the suite twice — with both files, then with requirements-base.txt alone — and the second run must pass.

pytest-xdist is in the base set specifically so -n auto uses all 24 threads.

PREF — uv for resolution and locking; pip-tools is a fine substitute. Either way the lock file is committed.


9. Services and backups

9.1 Systemd units (CT 100)

REQ — Two units, split along the control-plane / execution-plane boundary Part II §16 requires:

Unit Role
mechcomp.service FastAPI/uvicorn. Serves the catalogue, accepts jobs, returns cached artifacts. Never runs geometry.
mechcomp-worker.service Consumes the queue, runs generators, writes artifacts.

REQ — Both:

[Service]
User=mechcomp
Group=mechcomp
EnvironmentFile=/etc/mechcomp/mechcomp.env
WorkingDirectory=/var/www/mechcomp
Restart=on-failure
RestartSec=5

NoNewPrivileges=true
PrivateTmp=true
ProtectSystem=strict
ProtectHome=true
ReadWritePaths=/var/lib/mechcomp /var/log/mechcomp
ProtectKernelTunables=true
ProtectControlGroups=true
RestrictSUIDSGID=true

[Install]
WantedBy=multi-user.target

REQ — Worker only:

MemoryMax=8G
CPUQuota=1200%
TimeoutStopSec=30

A runaway geometry job must not take the container down. MemoryMax turning a hang into a clean OOM-kill of one worker is the desired behaviour.

REQ — Logging to journald. No application-managed rotation.

REQ — Job queue is SQLite-backed and in-process. No Redis, no Celery, no RabbitMQ. A table with a status column and a claim query is correct at this scale and adds no services to either packaging target. It sits behind an interface so it can be swapped if load ever justifies it. It will not.

9.2 Three backup tiers

Tier What Where Retention
Routine vzdump of both container rootfs local (/var/lib/vz/dump) see below
Application tar stream of the data set /var/lib/vz/mechcomp-app/ on srv-b 14 daily
Gold routine + application + manifest 32 GB removable USB, LUKS indefinite, offline

REQ — mp0 is excluded from vzdump (backup=0). /var/lib/vz sits on pve-root, which is 78 GiB, and a full root filesystem is how Proxmox hosts become interesting. Excluding the 60 GiB data volume keeps each routine archive to the rootfs alone — a few GiB compressed — and application data flows through the tier below, which is where it belongs anyway.

REQ — vzdump job: both containers, snapshot mode, zstd, daily at 02:30, storage local.

ASSUMED — Initial retention keep-last=3. Deliberately conservative: no vzdump has been taken on this host, so actual size is unmeasured. DISCOVER — after the first successful run, report the archive size and I will set final retention. With 70 GiB free, three archives of an unknown size is the safe starting point.

REQ — A df guard in the backup hook that refuses to run and alerts if /var/lib/vz falls below 20 GiB free.

REQ — Application backup is a single script in CT 100 emitting a tar stream on stdout, pulled host-side:

pct exec ${CT_ID_APP} -- /usr/local/bin/mechcomp-backup --stdout \
  > /var/lib/vz/mechcomp-app/mechcomp-$(date +%F).tar

Covering exactly:

/etc/mechcomp/mechcomp.env
/var/lib/mechcomp/db/
/var/lib/mechcomp/artifacts/
/var/lib/mechcomp/fixtures/

cache/ is excluded. /var/www/mechcomp is excluded because it is reproducible from git plus the lock file. This script is, almost verbatim, the future YunoHost backup script.

REQ — pull, never push. The container has no access to any backup destination. The alternative is tempting and simpler, but it hands a network-facing container write access to the last line of defence. It also avoids idmap ownership problems entirely.

9.3 Gold archive

REQ — 32 GB USB stick, LUKS2 full-device encryption, ext4 inside, filesystem label MC-GOLD-01, mounted at /mnt/gold with noauto.

Encryption is not optional: the archive contains mechcomp.env, which contains MECHCOMP_SECRET_KEY, and a stick that travels between machines is exactly the thing that gets lost. The passphrase lives in CIVICVS's password manager and never in the repository. Encrypting the device rather than the files means the later copy to the 3+ TB disk inherits the protection if the LUKS image is copied as a block image.

ext4 rather than exFAT: no 4 GiB file cap, ownership preserved, journalled. DISCOVER — if the stick must ever be read from Windows or macOS, say so and this becomes exFAT with age-encrypted archives instead.

REQ — Archive layout, one directory per mint:

/mnt/gold/<UTC-timestamp>-<git-tag>/
├── vzdump-lxc-100-*.tar.zst
├── vzdump-lxc-101-*.tar.zst
├── mechcomp-app-*.tar
├── site.env                    (the FILL values, not secrets)
├── MANIFEST.txt                git tag, revisions, host, PVE version
└── SHA256SUMS

REQ — gold-archive.sh verifies SHA256SUMS after writing and before unmounting. An archive that has not been verified has not been made.

ASSUMED — A gold archive is minted on every tagged release and immediately before any promotion or host migration. This ties the tier to the existing tag discipline rather than to a calendar.

DISCOVER — Where the 3+ TB USB disk is attached now that annales is out of scope. Until that is known, stick-to-disk copying is a documented manual step, not a scripted one.

9.4 Disk health monitoring

REQ — Configure smartd for the four RAID members using the already installed smartmontools. No third-party HPE repository is added, and ssacli is not installed: the investigation showed the kernel and smartctl -d cciss,N already provide what is needed.

# /etc/smartd.conf
/dev/sda -d cciss,0 -a -m root -M exec /usr/share/smartmontools/smartd-runner
/dev/sda -d cciss,1 -a -m root -M exec /usr/share/smartmontools/smartd-runner
/dev/sda -d cciss,2 -a -m root -M exec /usr/share/smartmontools/smartd-runner
/dev/sda -d cciss,3 -a -m root -M exec /usr/share/smartmontools/smartd-runner

These are 2010-vintage spinning disks in a RAID 1+0 holding both the staging environment and its routine backups. Knowing about a failure before the second one is worth the four lines.

DISCOVER — Whether Proxmox notifications can relay through the existing mail infrastructure. If not, smartd and vzdump failures land in local root mail only, which nobody reads — report this rather than leaving it silently broken.


10. Configuration contract

REQ — /etc/mechcomp/mechcomp.env in CT 100, and nothing else, configures the application:

MECHCOMP_ENV=staging               # staging | production
MECHCOMP_BIND=10.20.0.10           # service network only; never 0.0.0.0
MECHCOMP_PORT=8770
MECHCOMP_BASE_URL=https://mechanical-compiler.dev.infra
MECHCOMP_DATA_DIR=/var/lib/mechcomp
MECHCOMP_LOG_LEVEL=info
MECHCOMP_DB_URL=sqlite:////var/lib/mechcomp/db/mechcomp.sqlite3
MECHCOMP_WORKER_CONCURRENCY=4
MECHCOMP_CAD_BACKEND=none          # none | cadquery — 'none' until the 3D path exists
MECHCOMP_ARTIFACT_RETENTION_DAYS=30
MECHCOMP_SECRET_KEY=               # generated by stage 3; see §11.4

REQ — The application fails loudly at startup if a required variable is missing, and never falls back to a compiled-in default path. That single rule is most of what makes both packaging targets work unchanged.

REQ — MECHCOMP_BIND is the service-network address, plus 127.0.0.1. Binding 0.0.0.0 would defeat the point of the second bridge.

Promotion changes MECHCOMP_ENV, MECHCOMP_BIND and MECHCOMP_BASE_URL. Nothing else.


11. Provisioning scripts

Everything above is a specification for scripts, not a runbook for a human.

11.1 Repository

ASSUMED — A separate repository, TheRON/mechanical-compiler-infra.

Provisioning scripts do not belong in the application repository. A YunoHost reviewer eventually reads that tree, and Proxmox tooling beside the app is exactly the "non-standard directory architecture" the catalog flags. The two also version independently, which they will.

If you prefer one repository, put the scripts in deploy/ and add that path to the packaging exclusion list from the first commit — not after they have history.

11.2 Four stages

Script Runs as Where Does
stage1-host.sh root srv-b vmbr1, /etc/hosts, smartd, vzdump job, /var/lib/vz/mechcomp-app
stage2-create-cts.sh root srv-b template check, ARP verify, pct create ×2, mp0 with backup=0, first boot
stage3-app.sh root CT 100 packages, user, layout, venv, Docker, units, config, secret
stage4-proxy.sh root CT 101 nginx, CA cert install, vhost, /etc/hosts

REQ — The only interface between stages is deploy/site.env. Stage 2 copies it into both containers; stages 3 and 4 source it. No other state crosses a boundary.

REQ — Stage 4 installs its own vhost. Unlike revision 3, the proxy is ours, on this host, and there is no separate operator to hand a config to.

11.3 Idempotency

REQ — All four stages are safe to re-run. Create-if-absent, converge-if-present, never destroy. A re-run against a healthy host is a no-op that exits zero.

This follows from "version, tag, release with discipline": a script you cannot re-run is a script you cannot test, and a script you cannot test is ceremony, not discipline.

REQ — Every stage starts with set -euo pipefail and validates that each FILL value in site.env is non-empty before touching anything.

11.4 Secrets

REQ — MECHCOMP_SECRET_KEY is never committed and never appears in site.env. Stage 3 generates it on first run and refuses to overwrite:

if ! grep -q '^MECHCOMP_SECRET_KEY=.\+' /etc/mechcomp/mechcomp.env 2>/dev/null; then
    key=$(openssl rand -base64 32)
    # write it; never log it
fi

REQ — deploy/site.env is in .gitignore. deploy/site.env.example, with empty values, is committed. The LUKS passphrase is in neither.

11.5 Versioning

ASSUMED — The infrastructure repository carries its own semver, independent of the application's SB_REVISION. They drift for unrelated reasons and tying them would force meaningless bumps on both.

REQ — Tag a release before each provisioning run and record the tag in the run log and in the gold MANIFEST.txt. "Which version of the script built this container" must be answerable six months from now.


12. deploy/site.env

# ---- Proxmox host (confirmed 2026-08-15) --------------------------------
PVE_HOST=srv-b
PVE_STORAGE=local-lvm
PVE_DATA_STORAGE=local-lvm
PVE_BACKUP_STORAGE=local
PVE_BRIDGE_MGMT=vmbr0
PVE_BRIDGE_SVC=vmbr1
PVE_TEMPLATE=local:vztmpl/debian-12-standard_12.12-1_amd64.tar.zst

# ---- Application container ---------------------------------------------
CT_ID_APP=100                      # VERIFY before create
CT_HOSTNAME_APP=mechcomp
CT_IP_APP_MGMT=10.0.0.20/24        # VERIFY free by ARP before create
CT_IP_APP_SVC=10.20.0.10/24
CT_CORES_APP=16
CT_RAM_APP=16384
CT_SWAP_APP=4096
CT_ROOTFS_APP=40
CT_MP0_APP=60

# ---- Proxy container ---------------------------------------------------
CT_ID_PROXY=101                    # VERIFY before create
CT_HOSTNAME_PROXY=mcproxy
CT_IP_PROXY_MGMT=10.0.0.21/24      # VERIFY free by ARP before create
CT_IP_PROXY_SVC=10.20.0.11/24
CT_CORES_PROXY=2
CT_RAM_PROXY=1024
CT_SWAP_PROXY=512
CT_ROOTFS_PROXY=8

# ---- Network -----------------------------------------------------------
CT_GATEWAY=10.0.0.1
SVC_NET=10.20.0.0/24
SVC_HOST_IP=10.20.0.1

# ---- Service -----------------------------------------------------------
MECHCOMP_ENV=staging
SERVICE_FQDN=mechanical-compiler.dev.infra
APP_PORT=8770
TLS_CA=kane-county-civic-infrastructure

# ---- Identity ----------------------------------------------------------
ADMIN_USER=                        # FILL
ADMIN_SSH_KEY=                     # FILL

# ---- Backup ------------------------------------------------------------
VZDUMP_SCHEDULE="02:30"
VZDUMP_KEEP_LAST=3                 # revisit after first size measurement
APP_BACKUP_DIR=/var/lib/vz/mechcomp-app
APP_BACKUP_KEEP=14
GOLD_LABEL=MC-GOLD-01
GOLD_MOUNT=/mnt/gold

# ---- Reserved for production, not used on srv-b ------------------------
# PRODUCTION_FQDN=mechanical-compiler.manufacturing.kane-il.us

CT_WG_ALIAS is gone. The host routes and MASQUERADEs 10.0.0.0/24 out wg0, so containers reach the WireGuard network by their ordinary LAN address, and no per-container alias mechanism exists to configure.


13. Constraints the application code will follow

Recorded so the environment is not built in a way that quietly permits their violation.

  1. No hardcoded absolute paths. Everything derives from MECHCOMP_DATA_DIR.
  2. No writes outside MECHCOMP_DATA_DIR. Enforced by ProtectSystem=strict.
  3. No dependency on systemd from application code. Signal handling only.
  4. The app listens on a TCP port, on the service network and loopback only.
  5. No OpenSCAD, Qt or X11 dependency in the running application.
  6. The 3D backend is reached only through an interface in src/mechcomp/cad/. No other module imports CadQuery or OCP directly.
  7. Schema migrations are explicit and forward-only.
  8. Artifacts are addressed by content hash of inputs — parameter set plus generator revision — never by hash of output bytes. Mesh output is not reproducible across toolchain versions; input hashing keeps a qualification valid across an upgrade that did not change the geometry.
  9. Every generated artifact carries the generator revision that produced it.
  10. The test suite passes with requirements-cad.txt uninstalled.
  11. AGPL-3.0 §13 requires network users be offered the source. The web tier carries a visible source link to the Gitea repository. This is a licence obligation, not a nicety.

14. Acceptance checklist

Run and report. This is also the first draft of the YunoHost tests.toml criteria.

# ---- host ----
pveversion | head -1
ip -br address show vmbr1                    # expect 10.20.0.1/24
pct list                                     # expect 100 and 101, both running
grep -c mechanical-compiler.dev.infra /etc/hosts
systemctl is-active smartd
cat /etc/pve/jobs.cfg | grep -c vzdump       # expect >= 1
pct config 100 | grep mp0                    # must contain backup=0

# ---- CT 100 ----
pct exec 100 -- bash -c '
  cat /etc/debian_version
  python3 --version
  timedatectl show -p Timezone --value
  id mechcomp
  findmnt /var/lib/mechcomp                  # must be a mount, not rootfs
  stat -c "%U:%G %a" /etc/mechcomp/mechcomp.env
  dpkg -l | grep -Ei "openscad|libqt|xserver" || echo "clean"
  ip -br address | grep 10.20.0.10
  ss -lntp | grep 8770                       # must NOT bind 0.0.0.0
  grep -c "^MECHCOMP_SECRET_KEY=.\+" /etc/mechcomp/mechcomp.env
  /var/www/mechcomp/venv/bin/python -c "import shapely; print(shapely.__version__)"
  /var/www/mechcomp/venv/bin/python -c "import cadquery; print(cadquery.__version__)"
  docker info --format "{{.Driver}}"
  docker run --rm mechcomp/reference-toolchain:8.0.0 openscad --version
'

# ---- egress from CT 100 ----
pct exec 100 -- bash -c '
  for h in deb.debian.org pypi.org files.pythonhosted.org gitea.barternetwork.us; do
    curl -sS -o /dev/null -w "$h %{http_code}\n" "https://$h"
  done
  git ls-remote https://gitea.barternetwork.us/TheRON/mechanical-compiler HEAD | head -1
'

# ---- CT 101 and end to end ----
pct exec 101 -- nginx -t
pct exec 101 -- curl -sS -o /dev/null -w '%{http_code} %{ssl_verify_result}\n' \
  https://mechanical-compiler.dev.infra/

# ---- forwarded-proto reaches the app ----
pct exec 101 -- curl -sS https://mechanical-compiler.dev.infra/_debug/headers \
  | grep -i x-forwarded-proto                # expect https

# ---- idempotency: second run is a clean no-op ----
./stage1-host.sh && ./stage2-create-cts.sh && echo "re-run clean"

# ---- backup round trip ----
vzdump 100 --storage local --mode snapshot --compress zstd
ls -lh /var/lib/vz/dump/                     # REPORT the size
pct exec 100 -- /usr/local/bin/mechcomp-backup --stdout | wc -c

Promotion checklist, for later

Not now, but recorded so it is not invented under pressure: production must revisit the Proxmox firewall (§3.5), issue a Let's Encrypt certificate for the public FQDN (§5.2), confirm the reverse proxy is not on the WireGuard side (§2.2), and change exactly three variables in mechcomp.env (§10).


15. Deliberately out of scope

  • Hubzilla addon development, or any change to existing Hubzilla containers
  • Federation, identity, or qualification storage
  • PostgreSQL — SQLite is sufficient and will remain so for a long time
  • Redis, Celery, or any external queue
  • CI runners — the suite runs locally until there is a reason otherwise
  • Monitoring beyond journald, smartd, and Proxmox's own metrics
  • NIC bonding, VLANs, or any further network configuration (§3.2)
  • An internal DNS server (§3.4)
  • The YunoHost package and the Docker image themselves

Each has a natural moment. None of them is now.


16. Every assumption in one place

§ Assumption If wrong
3.3 10.0.0.20 and .21 are free and outside any DHCP pool ARP verify aborts stage 2; supply replacements
3.3 Static addressing, no DHCP reservation needed Reserve in DHCP instead; addresses unchanged
3.5 Proxmox firewall stays disabled in staging Enable per-container; production revisits regardless
4.1 CT IDs 100 and 101 Verified at create time; any free pair works
4.2 16 vCPU / 16 GiB for the app container Lower freely; nothing depends on it
5.1 Kane County CA issues the staging certificate A self-signed cert works; trust store step changes
5.2 Production uses HTTP-01 Switch to DNS-01 if that is existing practice; never both
9.2 keep-last=3 initially Raise once the first archive size is measured
9.3 LUKS2 on the gold stick, ext4 inside exFAT plus age-encrypted files if cross-OS reading is needed
9.3 Gold minted per tagged release and before promotion Any trigger works; this one ties to existing discipline
11.1 Separate mechanical-compiler-infra repository Use deploy/ in the app repo, excluded from packaging
11.3 All stages idempotent Only matters if you want create-once semantics; I would argue against
11.5 Infra versions independently of SB_REVISION Tie them together if you prefer one release train

17. Provenance disclosure

This project's documents, code and roadmap are LLM-generated under human direction. That is stated plainly here, in the README, and in any eventual catalog submission. We do not obscure it.

The YunoHost policy is a quality bar with a disclosure requirement, not a ban. Revision 1 overstated it. The catalog repository rejects generated packages that do not follow the example_ynh template, citing verbose code, hallucinated helpers and non-standard directory architectures — and then explicitly permits AI use provided the maintainer is transparent and can explain every line. The failure mode guarded against is sprawl, not provenance. The defence is discipline: minimal divergence from the template, no invented helpers, no extra files.

Package provenance is not upstream provenance. The policy text concerns the _ynh repository — a few hundred lines of shell and one TOML file. Nobody audits whether an upstream Rails application was written by a human. The reviewable surface is small.

Disclosure is right; headlining it is a tactical mistake. Part II §29 holds that early use cases should demonstrate the system's distinct value. If provenance becomes the pitch, the project is judged on that axis rather than on whether it makes distributed manufacturing capacity legible. State it clearly in the README; keep the pitch about manufacturing. The argument is won by an artifact a reviewer finds shorter and more conventional than average, not by the claim attached to it.


18. The packaging path, when it comes

DISCOVER — Unknown territory, to be documented as it is walked rather than researched ahead: manifest.toml v2 with helpers 2.1, the scripts/ set (install, remove, upgrade, backup, restore, change_url), conf/ templates, tests.toml, and package_check levels.

Two orienting notes. The useful comparison for acceptable weight is paperless-ngx_ynh or fab-manager_ynh, not a hello-world package. And per §17, the package is drafted against example_ynh and reviewed by CIVICVS as his own work — anything in it he would not defend in a review thread does not ship.


19. First work item after acceptance

Port sb-geom to Shapely and get pytest -n auto green against the 123 frozen cases in strap-beam-fixtures-8.0.0.json. The ten rejected cases are part of the contract: a port that accepts them is wrong.