Files
mechanical-compiler/docs/STAGING-STATE.md
T
2026-08-16 20:25:18 -04:00

14 KiB

STAGING-STATE.md

Live state of the Mechanical Compiler staging instance on srv-b.

Updated 2026-08-16, formal infrastructure acceptance
Instance Staging / development
Specification ENVIRONMENT.md revision 5
Failure log FAILURES.md
Method Manual, one command group at a time

0. How to use this file

This is the authoritative record of what is true on srv-b. Where it and ENVIRONMENT.md disagree, this file wins for facts and the specification is defective and must be corrected.

Completed and remaining work are in the same document deliberately. They are two halves of one boundary; separating them guarantees they drift.

Before any command: read this file, confirm the immediately relevant live state with a read-only command, then issue one command group. If it fails, record it in FAILURES.md before changing anything else.

Production gets its own PRODUCTION-STATE.md. The specification is shared; the state is not.

Acceptance boundary, 2026-08-16

Staging infrastructure is accepted. That claim is narrower than "the application is deployed," and deliberately so. Three subsystems sit outside it and must not be represented as either hidden failures or completed work:

Subsystem Status
Mail alert delivery Partially configured, characterised, not accepted end to end
smartd monitoring Deferred with mail alerting
Backup infrastructure Postponed by operator decision — strategy may change

1. Instance values

The srv-b bindings for the parameters in ENVIRONMENT.md.

Host

hostname          srv-b  /  srv-b.dev.infra
platform          Proxmox VE 8.4.0, Debian 12, kernel 6.8.12-9-pve
hardware          HP ProLiant DL360 G7, 2 x Xeon X5650, 24 threads, 31 GiB
storage           P410i, 4 x EG0146FAWHU, RAID 1+0, all members SMART OK
local             directory /var/lib/vz, ~70 GiB free   iso,vztmpl,backup
local-lvm         LVM-thin pve/data, 166.9 GiB          rootdir,images
vmbr0             10.0.0.12/24 on enp3s0f0, gw 10.0.0.1   management, LAN
vmbr1             10.20.0.1/24, bridge-ports none         service, portless
wg0               10.110.0.12/32, peer wg-pk.civicus.us:51820, allowed 10.110.0.0/22
resolver          75.75.75.75, search dev.infra
ip_forward        1
timezone          America/Chicago, NTP active, clock synchronised
systemd           running, zero failed units
spare NICs        enp3s0f1, enp4s0f0, enp4s0f1 — unconfigured, deliberately
template          local:vztmpl/debian-12-standard_12.12-1_amd64.tar.zst

Containers

CT 100 CT 101
hostname mechcomp mcproxy
role application, worker reverse proxy, TLS
features nesting=1,keyctl=1 nesting=1
cores / RAM / swap 16 / 16384 / 4096 2 / 1024 / 512
rootfs 40 GiB local-lvm 8 GiB local-lvm
mp0 60 GiB → /var/lib/mechcomp, backup=0 —
net0 eth0 on vmbr1, 10.20.0.10/24, gw 10.20.0.1 eth0 on vmbr1, 10.20.0.11/24, gw 10.20.0.1
Debian 12 bookworm, fully upgraded 12 bookworm, fully upgraded
systemd running, zero failed units running, zero failed units

Neither container has a LAN interface (F-017). Neither has a stale eth1 stanza (F-021) — verified persistent across reboot.

Names and identity

mechanical-compiler.dev.infra   ->  10.20.0.11    (CT 101, the proxy)
mechcomp.dev.infra              ->  10.20.0.10    PVE-generated
mcproxy.dev.infra               ->  10.20.0.11    PVE-generated

ADMIN_USER      sandor, uid 1000, groups sandor + mechcomp(996)
                shell /bin/bash, both containers
authorized key  SHA256:2pNffCscUUW5Wbs9uMepLEvMLKPqUV/7Lpk5tw7iWSY
                matches /root/.ssh/id_rsa.pub on srv-b — bastion, see section 4
service user    mechcomp, uid 999, gid 996
                home /var/www/mechcomp, shell /usr/sbin/nologin

Host NAT, persisted in /etc/iptables/rules.v4

nat POSTROUTING
  -s 10.20.0.0/24 -d 10.0.0.0/24            -j RETURN       # LAN not translated
  -s 10.0.0.0/24                  -o wg0    -j MASQUERADE   # pre-existing
  -s 10.20.0.0/24                 -o wg0    -j MASQUERADE   # containers -> WireGuard
  -s 10.20.0.0/24                 -o vmbr0  -j MASQUERADE   # containers -> internet

filter FORWARD
  -s 10.20.0.0/24 -d 10.20.0.0/24           -j ACCEPT       # container <-> container
  -s 10.20.0.0/24 -d 10.0.0.0/24            -j DROP         # LAN blocked

Rule order is load-bearing. RETURN must precede both service-network masquerades. Live and persisted states match.

TLS

CA        CN = Mechanical Compiler Staging CA   (locally generated)
leaf      CN = mechanical-compiler.dev.infra
          SAN = DNS:mechanical-compiler.dev.infra
validity  2026-08-16 -> 2028-11-18            <-- nothing renews this
trusted   srv-b, CT 100, CT 101
key       root:root 0600 /etc/ssl/mechcomp/server.key

No Kane County Civic Infrastructure CA issuance path exists on srv-b: no trust anchor, no step, no cfssl, no EasyRSA. Proxmox's own CA was deliberately not reused. Replacing this leaf later is two file copies and a reload.

Application environment

/etc/mechcomp/mechcomp.env, root:mechcomp, 0640:

MECHCOMP_ENV=staging
MECHCOMP_BIND=10.20.0.10          # scalar; loopback is NOT bound
MECHCOMP_PORT=8770
MECHCOMP_BASE_URL=https://mechanical-compiler.dev.infra
MECHCOMP_DATA_DIR=/var/lib/mechcomp
MECHCOMP_LOG_LEVEL=info
MECHCOMP_DB_URL=sqlite:////var/lib/mechcomp/db/mechcomp.sqlite3
MECHCOMP_WORKER_CONCURRENCY=4
MECHCOMP_CAD_BACKEND=none
MECHCOMP_ARTIFACT_RETENTION_DAYS=30
MECHCOMP_SECRET_KEY=<generated, not recorded>

Mail, as currently configured

relay endpoint    10.110.0.1:25 over wg0
banner            wg-pk.diagnostics.kane-il.us, STARTTLS offered
STARTTLS cert     self-signed, CN = wg-pk
authentication    unauthenticated accepted from 10.110.0.12
ports 465 / 587   unavailable; public 198.58.111.109 exposes no SMTP on this path
srv-b relayhost   [10.110.0.1]:25
smtp_tls_security_level  may
smtp_sasl_auth_enable    no
inet_interfaces   loopback-only
root alias        sandor@kane-il.us
local handoff     succeeds, relay returns SMTP 250, queue empties
FINAL DELIVERY    NOT ACCEPTED — operator received a delivery-failure message
status            deferred, cause unproven downstream (F-023)

Placeholder backend — staging scaffold, not application code

/usr/local/libexec/mechcomp-placeholder.py
/etc/systemd/system/mechcomp-placeholder.service
runs as mechcomp:mechcomp, reads mechcomp.env, binds 10.20.0.10:8770
Restart=on-failure, hardening set applied, enabled at multi-user.target

Exists solely to prove the proxy chain independently of the application. It is removed when the real service arrives.

Rollback copies on disk

srv-b     /etc/network/interfaces.before-mechcomp
          /etc/hosts.before-mechcomp
          /etc/hosts.before-svcfqdn
          /etc/iptables/rules.v4.before-svcnat
          /etc/iptables/rules.v4.before-lanblock
          /root/pct-100.before-svcnet
          /root/pct-101.before-svcnet
CT 100    /etc/hosts.before-svcfqdn
          /etc/network/interfaces pre-F-021 copy
CT 101    /etc/hosts.before-svcfqdn
          /etc/network/interfaces pre-F-021 copy

2. Verified

Each line was demonstrated by command output, not inferred.

Host and network

  • vmbr1 created, active, 10.20.0.1/24, portless
  • Duplicate address detection run before each container creation
  • Container to internet reachable (deb.debian.org, gitea.barternetwork.us)
  • Container to WireGuard reachable (10.110.0.1)
  • Container to LAN gateway 10.0.0.1 = 100% packet loss — the requirement
  • Container to srv-b 10.0.0.12 reachable via INPUT path (required, not a leak)
  • Container to container reachable over vmbr1
  • LAN to Proxmox console 10.0.0.12:8006 returns 200, unaffected
  • NAT and FORWARD rules present live and persisted, in correct order
  • Host: running, zero failed units

CT 100

  • Debian 12 bookworm, fully upgraded, zero pending
  • en_US.UTF-8 generated and active; America/Chicago
  • running, zero failed units — after reboot (F-021 corrected)
  • /var/lib/mechcomp is a real separate ext4 filesystem
  • Service user mechcomp, home /var/www/mechcomp, correct ownership
  • All four data subdirectories present, mechcomp:mechcomp, 0750
  • mechcomp.env present, root:mechcomp, 0640, secret generated
  • Base packages installed; OpenSCAD, Qt and X11 absent
  • Docker active, overlay2 / systemd, no fallback needed
  • Repository cloned at e85c4f4e, verified as the owning user
  • Python 3.11.2 venv created, owned by mechcomp
  • openssh-server enabled and active
  • mechcomp-placeholder.service enabled, active, reboot-persistent (F-019)
  • Listener on 10.20.0.10:8770 only — not 0.0.0.0

CT 101

  • Debian 12 bookworm, fully upgraded, zero pending
  • en_US.UTF-8 generated and active; America/Chicago
  • running, zero failed units — after reboot (F-021 corrected)
  • nginx 1.22.1, nginx -t passes, enabled and active
  • Only the project vhost enabled; Debian default removed
  • listen 443 ssl default_server on both address families
  • HTTP to HTTPS 301 redirect
  • Local CA created, leaf issued, trusted on all three hosts
  • HTTPS end to end from srv-b, CT 100 and CT 101 without -k
  • X-Forwarded-Proto: https observed at the backend — re-proved after the topology change and again after CT 100's reboot
  • openssh-server enabled and active

3. Remaining — infrastructure

Immediate

Nothing. The infrastructure boundary is accepted.

Deferred by operator decision

Recorded here so they are visually distinct from failures and from forgotten work. None is a defect.

  • Mail end-to-end delivery (F-023). Handoff to wg-pk works; final delivery does not. Leading untested hypothesis: envelope sender root@srv-b.dev.infra rejected on sender-domain verification, since dev.infra does not resolve publicly. Candidate remedies myorigin or smtp_generic_maps. Check the postfix check chroot divergence first.
  • smartd explicit four-member configuration. The package is installed and the service is active, but DEVICESCAN currently monitors zero devices; the P410i members are visible only through explicit -d cciss,N. Active is not the same as monitoring. Deferred with mail.
  • Backup infrastructure, entirely. Postponed 2026-08-16 because the strategy may change: vzdump job, archive sizing, retention, free-space guard, host-side pull, mechcomp-backup, backup alerting, gold media, 3+ TB redundancy.

One request standing against the postponement: a single manual vzdump of both containers as a point-in-time snapshot before further change. It presupposes nothing about the eventual strategy and yields the archive size that has been an open question since revision 4. Four 2010-vintage disks currently hold the only instance with no monitoring, no alerting and no backup.

Optional, recorded not scheduled

  • Second nameserver line in container resolver configuration, as a mitigation for F-022. One line, no daemon. Deliberately not applied on a single unexplained event.
  • Staging certificate expires 2028-11-18 and nothing renews it.

4. Access model

Confirmed: no workstation access from the home LAN is required.

internet -> WireGuard -> srv-b -> containers

srv-b is the bastion, confirmed by the authorized key matching /root/.ssh/id_rsa.pub on the host. Root on the host can pct enter regardless, so SSH adds no privilege — but every path to a container runs through srv-b, which is deliberate.

Direct WireGuard-side access to the catalogue would be a route addition for 10.20.0.0/24 on the hub. Not now.

No DHCP anywhere. Two containers with fixed addresses on a portless bridge is the entire address space.


5. Blocked on application code

Repository state:

HEAD        e85c4f4e5bab9a4f032230c99aa6784ace4c80e2
top level   .git  LICENSE  README.md  venv
git status  ?? venv/
absent      requirements-base.txt, requirements-cad.txt, service units

None of the following can be honestly completed, and none may be fabricated by provisioning:

  • Dependency install from committed manifests
  • mechcomp.service, mechcomp-worker.service
  • Reference toolchain image mechcomp/reference-toolchain:8.0.0
  • Fixture reproduction against ddd0f154...
  • Application-runtime acceptance
  • Replacing the placeholder with the real service

Architect decision: venv/ belongs in .gitignore — it is a legitimate artifact inside install_dir per YunoHost convention, it simply should not be tracked. Applied with the first application commit.

The Shapely port gates all of the above.


6. Open questions

# Question Blocks
1 Why does delivery fail downstream of wg-pk? mail, smartd, backup alerting
2 What is the backup strategy? all backup work
3 Where is the 3+ TB USB disk attached? gold redundancy step

Question 1 is answered by diagnosis. Questions 2 and 3 need CIVICVS.


7. Closed questions

Question Answer
Kane County CA issuance path None on srv-b. Local staging CA generated.
Relay address and authentication 10.110.0.1:25 over wg0, unauthenticated from 10.110.0.12 accepted.
openssh-server present Yes, both containers.
ADMIN_USER / key sandor; root@srv-b key, bastion pattern confirmed.
DHCP pool on 10.0.0.0/24 Not applicable. Containers are not on that network.
Docker storage driver overlay2 / systemd. No fallback needed.
LAN workstation access Not required. WireGuard through srv-b.
MECHCOMP_BIND semantics Scalar, service address only. Loopback not bound.