This specifies a **staging environment** on `srv-b`, plus the promotion path to a production host that does not yet exist.
41 KiB
ENVIRONMENT.md
Provisioning specification for the Mechanical Compiler development and staging environment.
| Revision | 4 (2026-08-15) |
| Supersedes | Revisions 1, 2 and 3 |
| Basis | ENVIRONMENT_INVESTIGATION_REPORT_2026-08-15.md (host srv-b) |
| Audience | The assistant or operator provisioning the containers |
| Author role | Written by the developer who will work inside this environment |
| Status | All decisions taken. Ready to execute. |
| Repository | https://gitea.barternetwork.us/TheRON/mechanical-compiler |
| Licence | AGPL-3.0-or-later (committed at e85c4f4e) |
0. How to use this document
This specifies a staging environment on srv-b, plus the promotion path to
a production host that does not yet exist.
Everything is scripted. Nothing is provisioned by hand.
Markers
| Marker | Meaning |
|---|---|
| REQ | Required. The work is blocked or wrong without it. |
| PREF | Preferred. Substitute freely, but report what you substituted. |
| ASSUMED | Decided by the author without confirmation. Defensible, but a guess. All are listed in §16. |
| FILL | A site value that must be supplied. Goes in deploy/site.env and nowhere else. |
| VERIFY | Must be checked at run time rather than trusted from this document. |
| DISCOVER | Genuinely unknown. Answer by doing, then report. |
If you are the provisioning assistant
Four requests. FILL values live in one file and are referenced from there — no literals scattered through scripts. VERIFY steps run inside the scripts, not as a pre-flight a human might skip. Report DISCOVER outcomes verbatim, including failures. And if a REQ item cannot be satisfied, say so rather than working around it: the stated reason usually matters more than the mechanism.
What changed in revision 4
| # | Change | Source |
|---|---|---|
| 1 | Host is srv-b, not annales |
CHG-001 |
| 2 | Staging and production are separate roles, on separate hosts | CHG-002 |
| 3 | srv-b is a standalone node; promotion is export/restore, not cluster migration |
CHG-003 |
| 4 | Template 12.12-1, not 12.7-1 |
CHG-004 |
| 5 | Three-tier backup architecture; 32 GB removable gold media | CHG-005 |
| 6 | Backup schedule defined explicitly; none exists to inherit | CHG-006 |
| 7 | CT_WG_ALIAS removed |
CHG-007 |
| 8 | PROXY_HOST not inherited; staging gets its own proxy container |
CHG-008 |
| 9 | Two containers, not one: application and reverse proxy | this revision |
| 10 | Second bridge vmbr1 isolates service traffic from management traffic |
this revision |
| 11 | Container sizing raised substantially to use the host | this revision |
| 12 | Staging is internal-only; the public FQDN is reserved for production | this revision |
1. Five decisions that shape everything below
1.1 The 2D path and the 3D path are separable
The catalogue's product is a 2D cross-section rendered to SVG. Shapely plus our own code covers that completely: region algebra, offsets, distance queries, output. CadQuery/OCCT is needed only for STEP and mesh export.
These stay separate — separate requirements files, separate code paths, and a CI job that runs the whole suite with the CAD dependencies absent. Three reasons, none of which is packaging politics:
- Part II §16 requires the control-plane / execution-plane separation regardless of how the app is distributed.
- A slow OCCT import must never land in the request path for a page that only draws a cross-section.
- Test-suite speed. Shapely-only tests run in seconds, which is the difference between running the 123 frozen fixtures on every save and only before commit.
Retained correction. Revision 1 also justified this as a defence against
YunoHost's "resource-hungry" criterion. That was overcalibrated —
fab-manager_ynh is in the catalog, paperless-ngx_ynh declares nine apt
dependencies including postgresql and redis, and the criterion reads
"resource-hungry compared to their features." CadQuery is a
default-installed dependency. The boundary stays; the apology goes.
1.2 YunoHost and Docker are parallel targets, not sequential ones
YunoHost apps install natively — apt, a venv, systemd, nginx — and the project does not want Docker inside YunoHost. Docker is a separate distribution channel, not a stepping stone to the catalog. Neither blocks the other.
1.3 Staging mirrors production topology, not production exposure
srv-b runs two containers because production will: an application container
that never terminates TLS, and a reverse proxy that does. A single container
serving directly would never exercise the X-Forwarded-Proto path — precisely
the class of bug that cost us time on Hubzilla's session logout.
Staging is not publicly reachable and does not use the public FQDN. It uses the Kane County Civic Infrastructure CA on an internal name. The public name and Let's Encrypt belong to production, which forces the promotion path to be exercised rather than assumed.
1.4 Directory layout mirrors YunoHost conventions from the start
The paths in §7 are what a YunoHost package would provision anyway
(install_dir, data_dir, a dedicated system user, a dynamically assigned
port). The eventual manifest.toml resource block then describes what already
exists.
1.5 All configuration comes from the environment, never from the filesystem
No path, port, URL, or secret may be hardcoded or discovered by convention. One
env file per container; the same variables become ENV in a Dockerfile and
ynh_add_config substitutions in a package.
The check that this holds: moving the service to a different hostname is a one-variable change. It currently is — which is why the staging/production split below costs almost nothing.
2. Host baseline (confirmed, not to be re-verified)
| Item | Value |
|---|---|
| Hostname / FQDN | srv-b / srv-b.dev.infra |
| Role | Development and staging. Standalone node, never clustered. |
| Platform | Proxmox VE 8.4.0 on Debian 12, kernel 6.8.12-9-pve |
| Hardware | HP ProLiant DL360 G7, firmware P68 |
| CPU | 2 × Xeon X5650, 24 logical CPUs |
| RAM / swap | 31 GiB / 8 GiB |
| Storage | HP Smart Array P410i, 4 × EG0146FAWHU, RAID 1+0, all SMART OK |
local |
directory, /var/lib/vz, ~70 GiB free — iso, vztmpl, backup |
local-lvm |
LVM-thin pve/data, 166.9 GiB, 0% used — rootdir, images |
| LAN | vmbr0 on enp3s0f0, 10.0.0.12/24, gateway 10.0.0.1 |
| WireGuard | wg0 10.110.0.12/32, peer wg-pk.civicus.us:51820, allowed 10.110.0.0/22 |
| Routing | forwarding on, MASQUERADE 10.0.0.0/24 -o wg0, both persistent |
| Resolver | 75.75.75.75, search dev.infra |
| Egress | Direct. No proxy, no allowlist. All required hosts reachable. |
| PVE firewall | disabled/running |
| Guests | None. Next ID 100. |
| Template | local:vztmpl/debian-12-standard_12.12-1_amd64.tar.zst |
2.1 Two hardware facts worth knowing before anyone is surprised
The X5650 is Westmere: no AVX, and single-thread performance is modest. Geometry solvers are single-threaded, so regenerating the fixture oracle will take noticeably longer here than on a modern laptop. That is expected, not a fault. All compiled wheels target an SSE2 baseline, so nothing breaks.
Against that, 24 threads is a great deal of parallel test capacity. The
test suite runs under pytest-xdist with -n auto, which is where the box
earns its keep.
2.2 The WireGuard path is outbound only
MASQUERADE means containers reach 10.110.0.0/22, but nothing on the
WireGuard network can initiate a connection back. Irrelevant for staging.
Relevant at promotion: if the production reverse proxy lives on the
WireGuard side, this topology will not serve it unchanged. Recorded now so it
is not rediscovered later.
3. Network
3.1 Two bridges
REQ — vmbr0, existing, unchanged. enp3s0f0, 10.0.0.12/24. This is the
management and egress network: SSH, apt, PyPI, Gitea, Docker Hub.
REQ — vmbr1, new, no physical port. Host address 10.20.0.1/24.
This is the service network: proxy-to-application traffic and nothing else.
A portless bridge needs no cable, no switch configuration, and no spare NIC. It
exists so the application container has no listener on the LAN at all: the app
binds 127.0.0.1 and its vmbr1 address only, and the only route to it is
through the proxy.
# /etc/network/interfaces — append
auto vmbr1
iface vmbr1 inet static
address 10.20.0.1/24
bridge-ports none
bridge-stp off
bridge-fd 0
3.2 The other three NICs stay unconfigured
REQ — Do not bond, trunk, or bridge enp3s0f1 or the second dual-port
adapter. Leave them down and cabled to nothing.
This is deliberate. An active-backup bond on vmbr0 would be defensible, but
it buys link redundancy on a staging host that is allowed to be down, at the
cost of a configuration that has to be understood by everyone who touches the
box afterwards. Not worth it here.
What the spare ports are for, when a reason appears: a point-to-point link to the production host for promotion transfers, and a dedicated management interface if the LAN ever becomes untrusted. Neither is now.
3.3 Addressing
ASSUMED — Static, outside any DHCP pool.
| Guest | vmbr0 (management) |
vmbr1 (service) |
|---|---|---|
srv-b host |
10.0.0.12/24 |
10.20.0.1/24 |
CT 100 mechcomp |
10.0.0.20/24 |
10.20.0.10/24 |
CT 101 mcproxy |
10.0.0.21/24 |
10.20.0.11/24 |
VERIFY — Stage 2 must confirm 10.0.0.20 and 10.0.0.21 are unclaimed
immediately before creating each container, and abort if not:
for ip in 10.0.0.20 10.0.0.21; do
if arping -D -I vmbr0 -c 3 -w 3 "$ip" >/dev/null 2>&1; then
echo " $ip free"
else
echo "ABORT: $ip answered ARP — address is in use" >&2; exit 1
fi
done
The investigation deliberately did not scan the LAN, and this document does not
either. An ARP probe for two specific addresses at creation time is the correct
scope. 10.0.0.160 was seen as a neighbour and is avoided.
DISCOVER — If the LAN has a DHCP server, report its pool range so these two
addresses can be confirmed outside it. If arping reports a conflict, stop and
ask rather than picking the next number.
3.4 Names
REQ — dev.infra has no resolver: the host points at 75.75.75.75, which
cannot answer for it. Names resolve through /etc/hosts only.
Stage 1 writes to /etc/hosts on the host, and stages 3 and 4 write the same
block into both containers:
10.20.0.10 mechanical-compiler.dev.infra mechcomp
10.20.0.11 mcproxy.dev.infra mcproxy
REQ — do not stand up an internal DNS server. Two containers and one host do not justify one, and adding a resolver here would create a second source of truth for names that production will not share.
3.5 Firewall
ASSUMED — The Proxmox firewall stays disabled, matching the observed staging baseline. The service network is portless and the application has no LAN listener, which is the containment that matters here.
REQ — Production must revisit this. Recorded as a promotion checklist item in §14, not as a permanent decision.
4. Containers
4.1 Roles
| CT 100 | CT 101 | |
|---|---|---|
| Hostname | mechcomp |
mcproxy |
| Role | Application and worker | Reverse proxy, TLS termination |
| Unprivileged | yes | yes |
| Features | nesting=1,keyctl=1 |
none |
VERIFY — Confirm both IDs with pvesh get /cluster/nextid and pct list
immediately before creation. 100 and 101 are candidates from a host that
had no guests, not reservations.
4.2 Sizing
The host has 24 threads and 31 GiB with nothing else on it. Sizing reflects that rather than a cautious default.
| CT 100 | CT 101 | Host remaining | |
|---|---|---|---|
| vCPU | 16 | 2 | 24 total, soft limits |
| RAM | 16 GiB | 1 GiB | ~14 GiB |
| Swap | 4 GiB | 512 MiB | — |
| rootfs | 40 GiB | 8 GiB | on local-lvm |
mp0 |
60 GiB at /var/lib/mechcomp |
— | — |
Thin provisioning means mp0 consumes only what is written. Committed total is
108 GiB of 166.9 GiB — generous but not over-committed, which matters because
thin-pool exhaustion is an ugly failure mode.
REQ — mp0 is a real Proxmox mount point with backup=0. See §9.2 for why
the backup flag is off, and §12 for the check that proves the mount is real.
4.3 Base OS
REQ — Debian 12 bookworm from
local:vztmpl/debian-12-standard_12.12-1_amd64.tar.zst, both containers.
Chosen over Debian 13 and Ubuntu because it matches YunoHost 12 stable, ships Python 3.11 (the best OCP/CadQuery wheel coverage), and is already present on the host.
VERIFY — Stage 2 confirms the template is present, and downloads it only if absent:
pveam list local | grep -q "${PVE_TEMPLATE##*/}" || pveam download local "${PVE_TEMPLATE##*/}"
REQ — America/Chicago, en_US.UTF-8 in both containers, matching the host.
5. TLS and service names
5.1 Staging
REQ — Internal name mechanical-compiler.dev.infra, certificate issued by
the Kane County Civic Infrastructure CA.
REQ — The CA root is installed into both containers' trust stores and must also be trusted on the workstation used to browse staging.
REQ — the public FQDN is not used on staging. No DNS record for
mechanical-compiler.manufacturing.kane-il.us may point at srv-b. No Let's
Encrypt certificate is issued for staging: it consumes rate limit against a
name that must be clean when production needs it, and it would put a
development box on the public internet.
5.2 Production, reserved
| Value | |
|---|---|
| FQDN | mechanical-compiler.manufacturing.kane-il.us |
| Zone | Labels within kane-il.us. Not delegated. No new DS record. |
| Certificate | Let's Encrypt, per-name, no wildcard |
| Challenge | ASSUMED HTTP-01 on the production proxy. Match existing practice if it is DNS-01; do not run both. |
manufacturing.kane-il.us is a namespace, not a service. Nothing is served
from it. kane-il.us itself stays reserved for cross-cutting infrastructure —
proxy, mail, directory, certificates — and must not acquire project records.
The leaf names the service, not the discipline, because the parent already carries the discipline and the namespace must hold siblings:
mechanical-compiler.manufacturing.kane-il.us this application
field-metrology.manufacturing.kane-il.us Utility One, later
stock-processing.manufacturing.kane-il.us Utility Two, later
5.3 Three names, never conflated
| Name | Scope | Value |
|---|---|---|
| Service FQDN | What users type | staging and production differ; see above |
| Infrastructure token | System user, paths, units, YunoHost app id | mechcomp |
| Project name | Documents, README, catalog entry | Mechanical Compiler |
The token stays mechcomp in both environments. YunoHost app ids accept
lowercase alphanumerics and underscores only, so a hyphenated FQDN could not
serve as one anyway.
6. Reverse proxy (CT 101)
REQ — NGINX. Terminates TLS on vmbr0, proxies to CT 100 over vmbr1.
REQ — The X-Forwarded-Proto header is mandatory, not optional. Getting
this wrong is what broke Hubzilla sessions.
server {
listen 443 ssl;
listen [::]:443 ssl;
server_name mechanical-compiler.dev.infra;
ssl_certificate /etc/ssl/mechcomp/server.crt; # Kane County CA
ssl_certificate_key /etc/ssl/mechcomp/server.key;
location / {
proxy_pass http://10.20.0.10:8770;
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
proxy_read_timeout 300s;
client_max_body_size 8m;
}
}
server {
listen 80;
listen [::]:80;
server_name mechanical-compiler.dev.infra;
return 301 https://$host$request_uri;
}
PREF — proxy_read_timeout 300s. Compile jobs are asynchronous by design,
but the initial synchronous render can take tens of seconds on a cold cache on
this CPU, and the 60s default leaves no headroom.
REQ — This vhost is the production vhost with two values changed:
server_name and the certificate paths. Keep it that way. Any divergence
beyond those two lines means staging has stopped testing production.
7. Users, directories, permissions (CT 100)
REQ — Dedicated system user, no login shell, no home of its own:
user:group mechcomp:mechcomp
shell /usr/sbin/nologin
REQ — Layout, matching what a YunoHost package would provision:
| Path | Owner | Mode | Purpose | YunoHost equivalent |
|---|---|---|---|---|
/var/www/mechcomp |
mechcomp:mechcomp |
0750 |
source and venv | install_dir |
/var/lib/mechcomp |
mechcomp:mechcomp |
0750 |
data, on mp0 |
data_dir |
/var/log/mechcomp |
mechcomp:mechcomp |
0750 |
logs | — |
/etc/mechcomp/mechcomp.env |
root:mechcomp |
0640 |
configuration | ynh_add_config target |
REQ — /var/lib/mechcomp subdivides as:
/var/lib/mechcomp/
├── db/ SQLite database
├── artifacts/ generated SVG, STL, STEP, manifests
├── fixtures/ frozen acceptance oracles, read-mostly
└── cache/ safe to delete at any time
The cache/ guarantee matters: anything not reproducible from db/ plus
artifacts/ must not live there. That is what keeps the backup story simple.
REQ — ${ADMIN_USER} gets an ordinary shell account in the mechcomp
group with ${ADMIN_SSH_KEY} installed, in both containers. The service
user is never used interactively.
8. Software (CT 100)
8.1 System packages
REQ
python3 python3-venv python3-dev
build-essential pkg-config
git curl ca-certificates
libgeos-dev # Shapely bundles GEOS in its wheel; headers make a
# source build possible if a wheel is unavailable
sqlite3
PREF — jq ripgrep tmux htop
REQ — do not install openscad, openscad-nightly, or any Qt or X11
package. This is not a weight argument: the running application never invokes
OpenSCAD, and §8.2 confines it where it belongs. If OpenSCAD appears in the
container's package list, something has gone wrong architecturally.
8.2 The pinned reference toolchain
REQ — One Docker image with the exact toolchain the fixture oracle was frozen against, and nothing else:
# tools/reference-toolchain/Dockerfile
FROM debian:12-slim
RUN apt-get update && apt-get install -y --no-install-recommends \
openscad git ca-certificates \
&& rm -rf /var/lib/apt/lists/*
RUN git clone https://github.com/BelfrySCAD/BOSL2.git /BOSL2 \
&& git -C /BOSL2 checkout 92d697c2856de2fed93a33e858068589cefc2898
Tag mechcomp/reference-toolchain:8.0.0. Its only job is to run
make_fixtures.py and prove the oracle still reproduces. This is why the
Debian revision skew does not matter: the container has no OpenSCAD, so its
version cannot drift.
Acceptance test — must reproduce exactly:
sha256 of strap-beam-fixtures-8.0.0.json
= ddd0f1548379205dd0c652ec07285b0dae331e52ff0a0437005dfc6cddcc2cb2
If it does not, stop and report. A drifting oracle is worse than no oracle.
DISCOVER — Docker in an unprivileged LXC is usually fine on this kernel
with overlay2, but it is not guaranteed. Report:
docker info --format '{{.Driver}} / {{.CgroupDriver}}'
docker run --rm debian:12-slim echo ok
fuse-overlayfs is the documented fallback. If both fail, say so — moving the
image build to a small VM on the same host is a minor change and preferable to
fighting the storage driver.
8.3 Python
REQ — Virtualenv at /var/www/mechcomp/venv. Debian 12 enforces PEP 668;
do not use --break-system-packages to get around it.
REQ — Pinned, hash-checked dependencies in two files, both installed:
| File | Contents |
|---|---|
requirements-base.txt |
shapely, fastapi, uvicorn, pydantic, sqlalchemy, jinja2, pytest, pytest-xdist |
requirements-cad.txt |
cadquery / build123d |
Both install here. The split is not about disk; it is the mechanism that keeps
§1.1 honest. CI runs the suite twice — with both files, then with
requirements-base.txt alone — and the second run must pass.
pytest-xdist is in the base set specifically so -n auto uses all 24 threads.
PREF — uv for resolution and locking; pip-tools is a fine substitute.
Either way the lock file is committed.
9. Services and backups
9.1 Systemd units (CT 100)
REQ — Two units, split along the control-plane / execution-plane boundary Part II §16 requires:
| Unit | Role |
|---|---|
mechcomp.service |
FastAPI/uvicorn. Serves the catalogue, accepts jobs, returns cached artifacts. Never runs geometry. |
mechcomp-worker.service |
Consumes the queue, runs generators, writes artifacts. |
REQ — Both:
[Service]
User=mechcomp
Group=mechcomp
EnvironmentFile=/etc/mechcomp/mechcomp.env
WorkingDirectory=/var/www/mechcomp
Restart=on-failure
RestartSec=5
NoNewPrivileges=true
PrivateTmp=true
ProtectSystem=strict
ProtectHome=true
ReadWritePaths=/var/lib/mechcomp /var/log/mechcomp
ProtectKernelTunables=true
ProtectControlGroups=true
RestrictSUIDSGID=true
[Install]
WantedBy=multi-user.target
REQ — Worker only:
MemoryMax=8G
CPUQuota=1200%
TimeoutStopSec=30
A runaway geometry job must not take the container down. MemoryMax turning a
hang into a clean OOM-kill of one worker is the desired behaviour.
REQ — Logging to journald. No application-managed rotation.
REQ — Job queue is SQLite-backed and in-process. No Redis, no Celery, no RabbitMQ. A table with a status column and a claim query is correct at this scale and adds no services to either packaging target. It sits behind an interface so it can be swapped if load ever justifies it. It will not.
9.2 Three backup tiers
| Tier | What | Where | Retention |
|---|---|---|---|
| Routine | vzdump of both container rootfs |
local (/var/lib/vz/dump) |
see below |
| Application | tar stream of the data set | /var/lib/vz/mechcomp-app/ on srv-b |
14 daily |
| Gold | routine + application + manifest | 32 GB removable USB, LUKS | indefinite, offline |
REQ — mp0 is excluded from vzdump (backup=0). /var/lib/vz sits on
pve-root, which is 78 GiB, and a full root filesystem is how Proxmox hosts
become interesting. Excluding the 60 GiB data volume keeps each routine archive
to the rootfs alone — a few GiB compressed — and application data flows through
the tier below, which is where it belongs anyway.
REQ — vzdump job: both containers, snapshot mode, zstd, daily at
02:30, storage local.
ASSUMED — Initial retention keep-last=3. Deliberately conservative: no
vzdump has been taken on this host, so actual size is unmeasured. DISCOVER
— after the first successful run, report the archive size and I will set final
retention. With 70 GiB free, three archives of an unknown size is the safe
starting point.
REQ — A df guard in the backup hook that refuses to run and alerts if
/var/lib/vz falls below 20 GiB free.
REQ — Application backup is a single script in CT 100 emitting a tar stream on stdout, pulled host-side:
pct exec ${CT_ID_APP} -- /usr/local/bin/mechcomp-backup --stdout \
> /var/lib/vz/mechcomp-app/mechcomp-$(date +%F).tar
Covering exactly:
/etc/mechcomp/mechcomp.env
/var/lib/mechcomp/db/
/var/lib/mechcomp/artifacts/
/var/lib/mechcomp/fixtures/
cache/ is excluded. /var/www/mechcomp is excluded because it is reproducible
from git plus the lock file. This script is, almost verbatim, the future
YunoHost backup script.
REQ — pull, never push. The container has no access to any backup destination. The alternative is tempting and simpler, but it hands a network-facing container write access to the last line of defence. It also avoids idmap ownership problems entirely.
9.3 Gold archive
REQ — 32 GB USB stick, LUKS2 full-device encryption, ext4 inside,
filesystem label MC-GOLD-01, mounted at /mnt/gold with noauto.
Encryption is not optional: the archive contains mechcomp.env, which contains
MECHCOMP_SECRET_KEY, and a stick that travels between machines is exactly the
thing that gets lost. The passphrase lives in CIVICVS's password manager and
never in the repository. Encrypting the device rather than the files means the
later copy to the 3+ TB disk inherits the protection if the LUKS image is
copied as a block image.
ext4 rather than exFAT: no 4 GiB file cap, ownership preserved, journalled.
DISCOVER — if the stick must ever be read from Windows or macOS, say so and
this becomes exFAT with age-encrypted archives instead.
REQ — Archive layout, one directory per mint:
/mnt/gold/<UTC-timestamp>-<git-tag>/
├── vzdump-lxc-100-*.tar.zst
├── vzdump-lxc-101-*.tar.zst
├── mechcomp-app-*.tar
├── site.env (the FILL values, not secrets)
├── MANIFEST.txt git tag, revisions, host, PVE version
└── SHA256SUMS
REQ — gold-archive.sh verifies SHA256SUMS after writing and before
unmounting. An archive that has not been verified has not been made.
ASSUMED — A gold archive is minted on every tagged release and immediately before any promotion or host migration. This ties the tier to the existing tag discipline rather than to a calendar.
DISCOVER — Where the 3+ TB USB disk is attached now that annales is out
of scope. Until that is known, stick-to-disk copying is a documented manual
step, not a scripted one.
9.4 Disk health monitoring
REQ — Configure smartd for the four RAID members using the already
installed smartmontools. No third-party HPE repository is added, and ssacli
is not installed: the investigation showed the kernel and smartctl -d cciss,N
already provide what is needed.
# /etc/smartd.conf
/dev/sda -d cciss,0 -a -m root -M exec /usr/share/smartmontools/smartd-runner
/dev/sda -d cciss,1 -a -m root -M exec /usr/share/smartmontools/smartd-runner
/dev/sda -d cciss,2 -a -m root -M exec /usr/share/smartmontools/smartd-runner
/dev/sda -d cciss,3 -a -m root -M exec /usr/share/smartmontools/smartd-runner
These are 2010-vintage spinning disks in a RAID 1+0 holding both the staging environment and its routine backups. Knowing about a failure before the second one is worth the four lines.
DISCOVER — Whether Proxmox notifications can relay through the existing
mail infrastructure. If not, smartd and vzdump failures land in local root
mail only, which nobody reads — report this rather than leaving it silently
broken.
10. Configuration contract
REQ — /etc/mechcomp/mechcomp.env in CT 100, and nothing else, configures
the application:
MECHCOMP_ENV=staging # staging | production
MECHCOMP_BIND=10.20.0.10 # service network only; never 0.0.0.0
MECHCOMP_PORT=8770
MECHCOMP_BASE_URL=https://mechanical-compiler.dev.infra
MECHCOMP_DATA_DIR=/var/lib/mechcomp
MECHCOMP_LOG_LEVEL=info
MECHCOMP_DB_URL=sqlite:////var/lib/mechcomp/db/mechcomp.sqlite3
MECHCOMP_WORKER_CONCURRENCY=4
MECHCOMP_CAD_BACKEND=none # none | cadquery — 'none' until the 3D path exists
MECHCOMP_ARTIFACT_RETENTION_DAYS=30
MECHCOMP_SECRET_KEY= # generated by stage 3; see §11.4
REQ — The application fails loudly at startup if a required variable is missing, and never falls back to a compiled-in default path. That single rule is most of what makes both packaging targets work unchanged.
REQ — MECHCOMP_BIND is the service-network address, plus 127.0.0.1.
Binding 0.0.0.0 would defeat the point of the second bridge.
Promotion changes MECHCOMP_ENV, MECHCOMP_BIND and MECHCOMP_BASE_URL.
Nothing else.
11. Provisioning scripts
Everything above is a specification for scripts, not a runbook for a human.
11.1 Repository
ASSUMED — A separate repository, TheRON/mechanical-compiler-infra.
Provisioning scripts do not belong in the application repository. A YunoHost reviewer eventually reads that tree, and Proxmox tooling beside the app is exactly the "non-standard directory architecture" the catalog flags. The two also version independently, which they will.
If you prefer one repository, put the scripts in deploy/ and add that path to
the packaging exclusion list from the first commit — not after they have history.
11.2 Four stages
| Script | Runs as | Where | Does |
|---|---|---|---|
stage1-host.sh |
root | srv-b |
vmbr1, /etc/hosts, smartd, vzdump job, /var/lib/vz/mechcomp-app |
stage2-create-cts.sh |
root | srv-b |
template check, ARP verify, pct create ×2, mp0 with backup=0, first boot |
stage3-app.sh |
root | CT 100 | packages, user, layout, venv, Docker, units, config, secret |
stage4-proxy.sh |
root | CT 101 | nginx, CA cert install, vhost, /etc/hosts |
REQ — The only interface between stages is deploy/site.env. Stage 2
copies it into both containers; stages 3 and 4 source it. No other state crosses
a boundary.
REQ — Stage 4 installs its own vhost. Unlike revision 3, the proxy is ours, on this host, and there is no separate operator to hand a config to.
11.3 Idempotency
REQ — All four stages are safe to re-run. Create-if-absent, converge-if-present, never destroy. A re-run against a healthy host is a no-op that exits zero.
This follows from "version, tag, release with discipline": a script you cannot re-run is a script you cannot test, and a script you cannot test is ceremony, not discipline.
REQ — Every stage starts with set -euo pipefail and validates that each
FILL value in site.env is non-empty before touching anything.
11.4 Secrets
REQ — MECHCOMP_SECRET_KEY is never committed and never appears in
site.env. Stage 3 generates it on first run and refuses to overwrite:
if ! grep -q '^MECHCOMP_SECRET_KEY=.\+' /etc/mechcomp/mechcomp.env 2>/dev/null; then
key=$(openssl rand -base64 32)
# write it; never log it
fi
REQ — deploy/site.env is in .gitignore. deploy/site.env.example, with
empty values, is committed. The LUKS passphrase is in neither.
11.5 Versioning
ASSUMED — The infrastructure repository carries its own semver, independent
of the application's SB_REVISION. They drift for unrelated reasons and tying
them would force meaningless bumps on both.
REQ — Tag a release before each provisioning run and record the tag in the
run log and in the gold MANIFEST.txt. "Which version of the script built this
container" must be answerable six months from now.
12. deploy/site.env
# ---- Proxmox host (confirmed 2026-08-15) --------------------------------
PVE_HOST=srv-b
PVE_STORAGE=local-lvm
PVE_DATA_STORAGE=local-lvm
PVE_BACKUP_STORAGE=local
PVE_BRIDGE_MGMT=vmbr0
PVE_BRIDGE_SVC=vmbr1
PVE_TEMPLATE=local:vztmpl/debian-12-standard_12.12-1_amd64.tar.zst
# ---- Application container ---------------------------------------------
CT_ID_APP=100 # VERIFY before create
CT_HOSTNAME_APP=mechcomp
CT_IP_APP_MGMT=10.0.0.20/24 # VERIFY free by ARP before create
CT_IP_APP_SVC=10.20.0.10/24
CT_CORES_APP=16
CT_RAM_APP=16384
CT_SWAP_APP=4096
CT_ROOTFS_APP=40
CT_MP0_APP=60
# ---- Proxy container ---------------------------------------------------
CT_ID_PROXY=101 # VERIFY before create
CT_HOSTNAME_PROXY=mcproxy
CT_IP_PROXY_MGMT=10.0.0.21/24 # VERIFY free by ARP before create
CT_IP_PROXY_SVC=10.20.0.11/24
CT_CORES_PROXY=2
CT_RAM_PROXY=1024
CT_SWAP_PROXY=512
CT_ROOTFS_PROXY=8
# ---- Network -----------------------------------------------------------
CT_GATEWAY=10.0.0.1
SVC_NET=10.20.0.0/24
SVC_HOST_IP=10.20.0.1
# ---- Service -----------------------------------------------------------
MECHCOMP_ENV=staging
SERVICE_FQDN=mechanical-compiler.dev.infra
APP_PORT=8770
TLS_CA=kane-county-civic-infrastructure
# ---- Identity ----------------------------------------------------------
ADMIN_USER= # FILL
ADMIN_SSH_KEY= # FILL
# ---- Backup ------------------------------------------------------------
VZDUMP_SCHEDULE="02:30"
VZDUMP_KEEP_LAST=3 # revisit after first size measurement
APP_BACKUP_DIR=/var/lib/vz/mechcomp-app
APP_BACKUP_KEEP=14
GOLD_LABEL=MC-GOLD-01
GOLD_MOUNT=/mnt/gold
# ---- Reserved for production, not used on srv-b ------------------------
# PRODUCTION_FQDN=mechanical-compiler.manufacturing.kane-il.us
CT_WG_ALIAS is gone. The host routes and MASQUERADEs 10.0.0.0/24 out wg0,
so containers reach the WireGuard network by their ordinary LAN address, and no
per-container alias mechanism exists to configure.
13. Constraints the application code will follow
Recorded so the environment is not built in a way that quietly permits their violation.
- No hardcoded absolute paths. Everything derives from
MECHCOMP_DATA_DIR. - No writes outside
MECHCOMP_DATA_DIR. Enforced byProtectSystem=strict. - No dependency on systemd from application code. Signal handling only.
- The app listens on a TCP port, on the service network and loopback only.
- No OpenSCAD, Qt or X11 dependency in the running application.
- The 3D backend is reached only through an interface in
src/mechcomp/cad/. No other module imports CadQuery or OCP directly. - Schema migrations are explicit and forward-only.
- Artifacts are addressed by content hash of inputs — parameter set plus generator revision — never by hash of output bytes. Mesh output is not reproducible across toolchain versions; input hashing keeps a qualification valid across an upgrade that did not change the geometry.
- Every generated artifact carries the generator revision that produced it.
- The test suite passes with
requirements-cad.txtuninstalled. - AGPL-3.0 §13 requires network users be offered the source. The web tier carries a visible source link to the Gitea repository. This is a licence obligation, not a nicety.
14. Acceptance checklist
Run and report. This is also the first draft of the YunoHost tests.toml
criteria.
# ---- host ----
pveversion | head -1
ip -br address show vmbr1 # expect 10.20.0.1/24
pct list # expect 100 and 101, both running
grep -c mechanical-compiler.dev.infra /etc/hosts
systemctl is-active smartd
cat /etc/pve/jobs.cfg | grep -c vzdump # expect >= 1
pct config 100 | grep mp0 # must contain backup=0
# ---- CT 100 ----
pct exec 100 -- bash -c '
cat /etc/debian_version
python3 --version
timedatectl show -p Timezone --value
id mechcomp
findmnt /var/lib/mechcomp # must be a mount, not rootfs
stat -c "%U:%G %a" /etc/mechcomp/mechcomp.env
dpkg -l | grep -Ei "openscad|libqt|xserver" || echo "clean"
ip -br address | grep 10.20.0.10
ss -lntp | grep 8770 # must NOT bind 0.0.0.0
grep -c "^MECHCOMP_SECRET_KEY=.\+" /etc/mechcomp/mechcomp.env
/var/www/mechcomp/venv/bin/python -c "import shapely; print(shapely.__version__)"
/var/www/mechcomp/venv/bin/python -c "import cadquery; print(cadquery.__version__)"
docker info --format "{{.Driver}}"
docker run --rm mechcomp/reference-toolchain:8.0.0 openscad --version
'
# ---- egress from CT 100 ----
pct exec 100 -- bash -c '
for h in deb.debian.org pypi.org files.pythonhosted.org gitea.barternetwork.us; do
curl -sS -o /dev/null -w "$h %{http_code}\n" "https://$h"
done
git ls-remote https://gitea.barternetwork.us/TheRON/mechanical-compiler HEAD | head -1
'
# ---- CT 101 and end to end ----
pct exec 101 -- nginx -t
pct exec 101 -- curl -sS -o /dev/null -w '%{http_code} %{ssl_verify_result}\n' \
https://mechanical-compiler.dev.infra/
# ---- forwarded-proto reaches the app ----
pct exec 101 -- curl -sS https://mechanical-compiler.dev.infra/_debug/headers \
| grep -i x-forwarded-proto # expect https
# ---- idempotency: second run is a clean no-op ----
./stage1-host.sh && ./stage2-create-cts.sh && echo "re-run clean"
# ---- backup round trip ----
vzdump 100 --storage local --mode snapshot --compress zstd
ls -lh /var/lib/vz/dump/ # REPORT the size
pct exec 100 -- /usr/local/bin/mechcomp-backup --stdout | wc -c
Promotion checklist, for later
Not now, but recorded so it is not invented under pressure: production must
revisit the Proxmox firewall (§3.5), issue a Let's Encrypt certificate for the
public FQDN (§5.2), confirm the reverse proxy is not on the WireGuard side
(§2.2), and change exactly three variables in mechcomp.env (§10).
15. Deliberately out of scope
- Hubzilla addon development, or any change to existing Hubzilla containers
- Federation, identity, or qualification storage
- PostgreSQL — SQLite is sufficient and will remain so for a long time
- Redis, Celery, or any external queue
- CI runners — the suite runs locally until there is a reason otherwise
- Monitoring beyond journald, smartd, and Proxmox's own metrics
- NIC bonding, VLANs, or any further network configuration (§3.2)
- An internal DNS server (§3.4)
- The YunoHost package and the Docker image themselves
Each has a natural moment. None of them is now.
16. Every assumption in one place
| § | Assumption | If wrong |
|---|---|---|
| 3.3 | 10.0.0.20 and .21 are free and outside any DHCP pool |
ARP verify aborts stage 2; supply replacements |
| 3.3 | Static addressing, no DHCP reservation needed | Reserve in DHCP instead; addresses unchanged |
| 3.5 | Proxmox firewall stays disabled in staging | Enable per-container; production revisits regardless |
| 4.1 | CT IDs 100 and 101 | Verified at create time; any free pair works |
| 4.2 | 16 vCPU / 16 GiB for the app container | Lower freely; nothing depends on it |
| 5.1 | Kane County CA issues the staging certificate | A self-signed cert works; trust store step changes |
| 5.2 | Production uses HTTP-01 | Switch to DNS-01 if that is existing practice; never both |
| 9.2 | keep-last=3 initially |
Raise once the first archive size is measured |
| 9.3 | LUKS2 on the gold stick, ext4 inside | exFAT plus age-encrypted files if cross-OS reading is needed |
| 9.3 | Gold minted per tagged release and before promotion | Any trigger works; this one ties to existing discipline |
| 11.1 | Separate mechanical-compiler-infra repository |
Use deploy/ in the app repo, excluded from packaging |
| 11.3 | All stages idempotent | Only matters if you want create-once semantics; I would argue against |
| 11.5 | Infra versions independently of SB_REVISION |
Tie them together if you prefer one release train |
17. Provenance disclosure
This project's documents, code and roadmap are LLM-generated under human direction. That is stated plainly here, in the README, and in any eventual catalog submission. We do not obscure it.
The YunoHost policy is a quality bar with a disclosure requirement, not a
ban. Revision 1 overstated it. The catalog repository rejects generated
packages that do not follow the example_ynh template, citing verbose code,
hallucinated helpers and non-standard directory architectures — and then
explicitly permits AI use provided the maintainer is transparent and can
explain every line. The failure mode guarded against is sprawl, not provenance.
The defence is discipline: minimal divergence from the template, no invented
helpers, no extra files.
Package provenance is not upstream provenance. The policy text concerns the
_ynh repository — a few hundred lines of shell and one TOML file. Nobody
audits whether an upstream Rails application was written by a human. The
reviewable surface is small.
Disclosure is right; headlining it is a tactical mistake. Part II §29 holds that early use cases should demonstrate the system's distinct value. If provenance becomes the pitch, the project is judged on that axis rather than on whether it makes distributed manufacturing capacity legible. State it clearly in the README; keep the pitch about manufacturing. The argument is won by an artifact a reviewer finds shorter and more conventional than average, not by the claim attached to it.
18. The packaging path, when it comes
DISCOVER — Unknown territory, to be documented as it is walked rather than
researched ahead: manifest.toml v2 with helpers 2.1, the scripts/ set
(install, remove, upgrade, backup, restore, change_url), conf/ templates,
tests.toml, and package_check levels.
Two orienting notes. The useful comparison for acceptable weight is
paperless-ngx_ynh or fab-manager_ynh, not a hello-world package. And per
§17, the package is drafted against example_ynh and reviewed by CIVICVS as
his own work — anything in it he would not defend in a review thread does not
ship.
19. First work item after acceptance
Port sb-geom to Shapely and get pytest -n auto green against the 123 frozen
cases in strap-beam-fixtures-8.0.0.json. The ten rejected cases are part of
the contract: a port that accepts them is wrong.