Files
mechanical-compiler/docs/STAGING-STATE.md
T
2026-08-17 11:48:32 -04:00

17 KiB

STAGING-STATE.md

Live state of the Mechanical Compiler staging instance on srv-b.

Updated 2026-08-17, after work order 003 (monitoring, SMTP egress)
Instance Staging / development
Specification ENVIRONMENT.md revision 5
Failure log FAILURES.md
Method Manual, one command group at a time

0. How to use this file

This is the authoritative record of what is true on srv-b. Where it and ENVIRONMENT.md disagree, this file wins for facts and the specification is defective and must be corrected.

Completed and remaining work are in the same document deliberately. They are two halves of one boundary; separating them guarantees they drift.

Before any command: read this file, confirm the immediately relevant live state with a read-only command, then issue one command group. If it fails, record it in FAILURES.md before changing anything else.

Production gets its own PRODUCTION-STATE.md. The specification is shared; the state is not.

Acceptance boundary, 2026-08-16

Staging infrastructure is accepted. That claim is narrower than "the application is deployed," and deliberately so. Three subsystems sit outside it and must not be represented as either hidden failures or completed work:

Subsystem Status
Mail alert delivery Accepted 2026-08-17. Delivered end to end, twice, headers captured.
smartd monitoring Accepted 2026-08-17. Four members monitored, four alerts received.
Backup infrastructure Postponed by operator decision — strategy may change

1. Instance values

The srv-b bindings for the parameters in ENVIRONMENT.md.

Host

hostname          srv-b  /  srv-b.dev.infra
platform          Proxmox VE 8.4.0, Debian 12, kernel 6.8.12-9-pve
hardware          HP ProLiant DL360 G7, 2 x Xeon X5650, 24 threads, 31 GiB
storage           P410i, 4 x EG0146FAWHU, RAID 1+0, all members SMART OK
local             directory /var/lib/vz, ~70 GiB free   iso,vztmpl,backup
local-lvm         LVM-thin pve/data, 166.9 GiB          rootdir,images
vmbr0             10.0.0.12/24 on enp3s0f0, gw 10.0.0.1   management, LAN
vmbr1             10.20.0.1/24, bridge-ports none         service, portless
wg0               10.110.0.12/32, peer wg-pk.civicus.us:51820, allowed 10.110.0.0/22
resolver          75.75.75.75, search dev.infra
ip_forward        1
timezone          America/Chicago, NTP active, clock synchronised
systemd           running, zero failed units
spare NICs        enp3s0f1, enp4s0f0, enp4s0f1 — unconfigured, deliberately
template          local:vztmpl/debian-12-standard_12.12-1_amd64.tar.zst

Containers

CT 100 CT 101
hostname mechcomp mcproxy
role application, worker reverse proxy, TLS
features nesting=1,keyctl=1 nesting=1
cores / RAM / swap 16 / 16384 / 4096 2 / 1024 / 512
rootfs 40 GiB local-lvm 8 GiB local-lvm
mp0 60 GiB → /var/lib/mechcomp, backup=0 —
net0 eth0 on vmbr1, 10.20.0.10/24, gw 10.20.0.1 eth0 on vmbr1, 10.20.0.11/24, gw 10.20.0.1
onboot 1 1
Debian 12 bookworm, fully upgraded 12 bookworm, fully upgraded
systemd running, zero failed units running, zero failed units

Neither container has a LAN interface (F-017). Neither has a stale eth1 stanza (F-021) — verified persistent across reboot.

Names and identity

mechanical-compiler.dev.infra   ->  10.20.0.11    (CT 101, the proxy)
mechcomp.dev.infra              ->  10.20.0.10    PVE-generated
mcproxy.dev.infra               ->  10.20.0.11    PVE-generated

ADMIN_USER      sandor, uid 1000, groups sandor + mechcomp(996)
                shell /bin/bash, both containers
authorized key  SHA256:2pNffCscUUW5Wbs9uMepLEvMLKPqUV/7Lpk5tw7iWSY
                matches /root/.ssh/id_rsa.pub on srv-b — bastion, see section 4
service user    mechcomp, uid 999, gid 996
                home /var/www/mechcomp, shell /usr/sbin/nologin

Host NAT, persisted in /etc/iptables/rules.v4

nat POSTROUTING
  -s 10.20.0.0/24 -d 10.0.0.0/24            -j RETURN       # LAN not translated
  -s 10.0.0.0/24                  -o wg0    -j MASQUERADE   # pre-existing
  -s 10.20.0.0/24                 -o wg0    -j MASQUERADE   # containers -> WireGuard
  -s 10.20.0.0/24                 -o vmbr0  -j MASQUERADE   # containers -> internet

filter FORWARD
  -s 10.20.0.0/24 -d 10.110.0.0/22 -p tcp
     -m multiport --dports 25,465,587       -j DROP         # SMTP egress blocked
  -s 10.20.0.0/24 -d 10.20.0.0/24           -j ACCEPT       # container <-> container
  -s 10.20.0.0/24 -d 10.0.0.0/24            -j DROP         # LAN blocked

Rule order is load-bearing. RETURN must precede both service-network masquerades. Live and persisted states match.

Disk monitoring — accepted 2026-08-17

DEVICESCAN finds nothing behind the P410i. Members are addressed explicitly:

/etc/smartd.conf
DEFAULT -a -m root -M exec /usr/share/smartmontools/smartd-runner
/dev/sda -d cciss,0
/dev/sda -d cciss,1
/dev/sda -d cciss,2
/dev/sda -d cciss,3

before:  Monitoring 0 ATA/SATA, 0 SCSI/SAS and 0 NVMe devices
after:   Monitoring 0 ATA/SATA, 4 SCSI/SAS and 0 NVMe devices
unit:    smartmontools.service   (smartd.service is an alias with no journal)
alerts:  4 test alerts generated and received; headers captured
         -M test reverted; permanent config restored and reboot-proven

Physical identity of the four members. This is how a future SMART alert is matched to a caddy:

Member Model Serial
cciss,0 EG0146FAWHU 6SD0LETK0000B05009W1
cciss,1 EG0146FAWHU 3SD3K5P00000905094UN
cciss,2 EG0146FAWHU 3SD3GHJ600009047RRRG
cciss,3 EG0146FAWHU 3SD3GVP800009045X79B

Not yet captured: power-on hours and grown-defect counts. On 2010-vintage drives holding the only instance, those numbers say how much runway remains. smartctl -d cciss,N -A /dev/sda for each member.

TLS

CA        CN = Mechanical Compiler Staging CA   (locally generated)
leaf      CN = mechanical-compiler.dev.infra
          SAN = DNS:mechanical-compiler.dev.infra
validity  2026-08-16 -> 2028-11-18            <-- nothing renews this
trusted   srv-b, CT 100, CT 101
key       root:root 0600 /etc/ssl/mechcomp/server.key

No Kane County Civic Infrastructure CA issuance path exists on srv-b: no trust anchor, no step, no cfssl, no EasyRSA. Proxmox's own CA was deliberately not reused. Replacing this leaf later is two file copies and a reload.

Application environment

/etc/mechcomp/mechcomp.env, root:mechcomp, 0640:

MECHCOMP_ENV=staging
MECHCOMP_BIND=10.20.0.10          # scalar; loopback is NOT bound
MECHCOMP_PORT=8770
MECHCOMP_BASE_URL=https://mechanical-compiler.dev.infra
MECHCOMP_DATA_DIR=/var/lib/mechcomp
MECHCOMP_LOG_LEVEL=info
MECHCOMP_DB_URL=sqlite:////var/lib/mechcomp/db/mechcomp.sqlite3
MECHCOMP_WORKER_CONCURRENCY=4
MECHCOMP_CAD_BACKEND=none
MECHCOMP_ARTIFACT_RETENTION_DAYS=30
MECHCOMP_SECRET_KEY=<generated, not recorded>

Mail — accepted 2026-08-17

srv-b relayhost   [10.110.0.1]:25 over wg0
root alias        sandor@kane-il.us   (postalias verified)
inet_interfaces   loopback-only
delivery path     srv-b -> wg-pk -> mx1 -> kane-il.us
status            DELIVERED end to end, twice, full headers captured
                  queue empty, no bounced or deferred entries

F-023 was proven to be a relay-trust mismatch three hops downstream: mx1 rejected RCPT TO with 554 5.7.1 Access denied because wg-pk's public addresses were absent from its mynetworks. Corrected by adding only those two addresses.

Standing constraints on this path — do not violate:

Host Constraint
kane-il.us Runs YunoHost. Its mail configuration must not be altered. This is why mx1 exists.
mx1.diagnostics.kane-il.us Uses DANE. Its TLS and certificate configuration must not be broken. The mynetworks change is inbound client trust and does not interact with DANE.
wg-pk Must continue relaying to arbitrary external destinations. Hubzilla registration and notification mail depends on it, and the ISP blocks port 25. Do not restrict it by recipient.

Fragility to know about: delivery now depends on Linode not reassigning 198.58.111.109 or 2600:3c00::f03c:92ff:fe42:43d7. If either changes, mail stops silently and the cause is in mx1's mynetworks.

Open exposure — see F-025. The correction authorises wg-pk under permit_mynetworks, which grants relay to any destination. Combined with wg-pk's own mynetworks = 10.110.0.0/22, every tunnel peer — including both containers, which arrive as 10.110.0.12 through the host masquerade — can originate mail as kane-il.us infrastructure. Correction pending decision.

Logging note (F-024): /var/log/mail.log does not exist. Proxmox ships without rsyslog; Postfix logs to journald. Use journalctl -u postfix@-.

Placeholder backend — staging scaffold, not application code

/usr/local/libexec/mechcomp-placeholder.py
/etc/systemd/system/mechcomp-placeholder.service
runs as mechcomp:mechcomp, reads mechcomp.env, binds 10.20.0.10:8770
Restart=on-failure, hardening set applied, enabled at multi-user.target

Exists solely to prove the proxy chain independently of the application. It is removed when the real service arrives.

Rollback copies on disk

srv-b     /etc/network/interfaces.before-mechcomp
          /etc/hosts.before-mechcomp
          /etc/hosts.before-svcfqdn
          /etc/iptables/rules.v4.before-svcnat
          /etc/iptables/rules.v4.before-lanblock
          /root/pct-100.before-svcnet
          /root/pct-101.before-svcnet
CT 100    /etc/hosts.before-svcfqdn
          /etc/network/interfaces pre-F-021 copy
CT 101    /etc/hosts.before-svcfqdn
          /etc/network/interfaces pre-F-021 copy

2. Verified

Each line was demonstrated by command output, not inferred.

Host and network

  • vmbr1 created, active, 10.20.0.1/24, portless
  • Duplicate address detection run before each container creation
  • Container to internet reachable (deb.debian.org, gitea.barternetwork.us)
  • Container to WireGuard reachable (10.110.0.1)
  • Container to LAN gateway 10.0.0.1 = 100% packet loss — the requirement
  • Container to srv-b 10.0.0.12 reachable via INPUT path (required, not a leak)
  • Container to container reachable over vmbr1
  • LAN to Proxmox console 10.0.0.12:8006 returns 200, unaffected
  • NAT and FORWARD rules present live and persisted, in correct order
  • Host: running, zero failed units
  • Mail delivered end to end from root on srv-b, twice, headers captured
  • smartd monitoring 4 devices; 4 test alerts delivered and received
  • Container SMTP egress blocked; internet egress and host mail intact
  • Both containers autostart after a host reboot (onboot: 1)

CT 100

  • Debian 12 bookworm, fully upgraded, zero pending
  • en_US.UTF-8 generated and active; America/Chicago
  • running, zero failed units — after reboot (F-021 corrected)
  • /var/lib/mechcomp is a real separate ext4 filesystem
  • Service user mechcomp, home /var/www/mechcomp, correct ownership
  • All four data subdirectories present, mechcomp:mechcomp, 0750
  • mechcomp.env present, root:mechcomp, 0640, secret generated
  • Base packages installed; OpenSCAD, Qt and X11 absent
  • Docker active, overlay2 / systemd, no fallback needed
  • Repository cloned at e85c4f4e, verified as the owning user
  • Python 3.11.2 venv created, owned by mechcomp
  • openssh-server enabled and active
  • mechcomp-placeholder.service enabled, active, reboot-persistent (F-019)
  • Listener on 10.20.0.10:8770 only — not 0.0.0.0

CT 101

  • Debian 12 bookworm, fully upgraded, zero pending
  • en_US.UTF-8 generated and active; America/Chicago
  • running, zero failed units — after reboot (F-021 corrected)
  • nginx 1.22.1, nginx -t passes, enabled and active
  • Only the project vhost enabled; Debian default removed
  • listen 443 ssl default_server on both address families
  • HTTP to HTTPS 301 redirect
  • Local CA created, leaf issued, trusted on all three hosts
  • HTTPS end to end from srv-b, CT 100 and CT 101 without -k
  • X-Forwarded-Proto: https observed at the backend — re-proved after the topology change and again after CT 100's reboot
  • openssh-server enabled and active

3. Remaining — infrastructure

Immediate

Nothing. The infrastructure boundary is accepted.

Deferred by operator decision

Recorded here so they are visually distinct from failures and from forgotten work. None is a defect.

  • wg-pk mynetworks scope (F-025). Estate decision: whether every tunnel peer should originate mail as kane-il.us infrastructure, or only the hosts that legitimately do. Not this project's to change unilaterally.
  • SMART attribute baseline. Power-on hours and grown-defect counts for the four members, so drive age is a known quantity rather than an assumption. smartctl -d cciss,N -A /dev/sda. Read-only, one command.
  • Backup infrastructure, entirely. Postponed 2026-08-16 because the strategy may change: vzdump job, archive sizing, retention, free-space guard, host-side pull, mechcomp-backup, backup alerting, gold media, 3+ TB redundancy.

One request standing against the postponement: a single manual vzdump of both containers as a point-in-time snapshot before further change. It presupposes nothing about the eventual strategy and yields the archive size that has been an open question since revision 4. Four 2010-vintage disks currently hold the only instance with no monitoring, no alerting and no backup.

Optional, recorded not scheduled

  • Second nameserver line in container resolver configuration, as a mitigation for F-022. One line, no daemon. Deliberately not applied on a single unexplained event.
  • Staging certificate expires 2028-11-18 and nothing renews it.

4. Access model

Confirmed: no workstation access from the home LAN is required.

internet -> WireGuard -> srv-b -> containers

srv-b is the bastion, confirmed by the authorized key matching /root/.ssh/id_rsa.pub on the host. Root on the host can pct enter regardless, so SSH adds no privilege — but every path to a container runs through srv-b, which is deliberate.

Direct WireGuard-side access to the catalogue would be a route addition for 10.20.0.0/24 on the hub. Not now.

No DHCP anywhere. Two containers with fixed addresses on a portless bridge is the entire address space.


5. Blocked on application code

Repository state:

HEAD        e85c4f4e5bab9a4f032230c99aa6784ace4c80e2
top level   .git  LICENSE  README.md  venv
git status  ?? venv/
absent      requirements-base.txt, requirements-cad.txt, service units

None of the following can be honestly completed, and none may be fabricated by provisioning:

  • Dependency install from committed manifests
  • mechcomp.service, mechcomp-worker.service
  • Reference toolchain image mechcomp/reference-toolchain:8.0.0
  • Fixture reproduction against ddd0f154...
  • Application-runtime acceptance
  • Replacing the placeholder with the real service

Architect decision: venv/ belongs in .gitignore — it is a legitimate artifact inside install_dir per YunoHost convention, it simply should not be tracked. Applied with the first application commit.

The Shapely port gates all of the above.


6. Open questions

# Question Blocks
1 Should wg-pk mynetworks narrow to explicit hosts? F-025 estate half
2 What is the backup strategy? all backup work
3 Where is the 3+ TB USB disk attached? gold redundancy step

All three need CIVICVS.


7. Closed questions

Question Answer
Kane County CA issuance path None on srv-b. Local staging CA generated.
Relay address and authentication 10.110.0.1:25 over wg0, unauthenticated from 10.110.0.12 accepted.
openssh-server present Yes, both containers.
ADMIN_USER / key sandor; root@srv-b key, bastion pattern confirmed.
DHCP pool on 10.0.0.0/24 Not applicable. Containers are not on that network.
Docker storage driver overlay2 / systemd. No fallback needed.
LAN workstation access Not required. WireGuard through srv-b.
MECHCOMP_BIND semantics Scalar, service address only. Loopback not bound.
Why did delivery fail downstream? mx1 relay trust. Proven and corrected (F-023).
Where does Postfix log on this host? journald. No rsyslog, no /var/log/mail.log (F-024).
Should container SMTP egress be blocked? Yes. Implemented and reboot-proven (F-025 project-local half).
Do the containers autostart? Yes, onboot: 1 on both, proven by host reboot (F-026).