Files
mechanical-compiler/docs/STAGING-STATE.md
T
TheRON a8081e17dc state: the composer replaced the placeholder; open the public ingress question
PROCESS.md section 8 requires host changes to reach STAGING-STATE.md before a session ends. mechcomp-placeholder.service was disabled and stopped on 2026-09-11 and mechcomp.service took 10.20.0.10:8770 in its place. nginx on CT 101 needed no change because the real service took the address the placeholder occupied. The placeholder unit stays on disk, disabled, as the rollback.

Section 5 items closed: mechcomp.service, and replacing the placeholder. Application runtime acceptance is marked partial rather than done, because no acceptance criteria have been written for the composer and it is not reachable from outside.

WORK-ORDER-004 publishes the composer at dev.mechcomp.kane-il.us. The srv-b half needs no change, verified: forwarding on, FORWARD policy ACCEPT, a direct route on vmbr1, and none of the three FORWARD rules matches hub initiated inbound traffic. The single gate is AllowedIPs on the hub peer entry for srv-b, because WireGuard drops by cryptokey routing before consulting any routing table.

Design decision recorded: the public name terminates on the hub and proxies to CT 100 directly rather than through CT 101. Routing it through CT 101 would make the public path depend on a locally signed leaf that expires 2028-11-18 with nothing renewing it. The tunnel already provides the encryption that hop would add. CT 101 keeps serving the internal name.

Section 4 not now decision on the hub route is reopened. It was correct while there was no application to reach. The consequence of leaving it closed is that the operator cannot see the application at all.
2026-09-11 06:59:53 -05:00

23 KiB

STAGING-STATE.md

Live state of the Mechanical Compiler staging instance on srv-b.

Updated 2026-09-11, after the composer replaced the placeholder
Instance Staging / development
Specification ENVIRONMENT.md revision 5
Failure log FAILURES.md
Process PROCESS.md — read first
Method Manual, one command group at a time

0. How to use this file

This is the authoritative record of what is true on srv-b. Where it and ENVIRONMENT.md disagree, this file wins for facts and the specification is defective and must be corrected.

Completed and remaining work are in the same document deliberately. They are two halves of one boundary; separating them guarantees they drift.

Read PROCESS.md before your first command. It describes how work is done here — who runs what, on which machine, with which tools. This file describes only what is currently true.

Before any command: read this file, confirm the immediately relevant live state with a read-only command, then issue one command group. If it fails, record it in FAILURES.md before changing anything else.

Production gets its own PRODUCTION-STATE.md. The specification is shared; the state is not.

Acceptance boundary, 2026-08-16

Staging infrastructure is accepted. That claim is narrower than "the application is deployed," and deliberately so. Three subsystems sit outside it and must not be represented as either hidden failures or completed work:

Subsystem Status
Mail alert delivery Accepted 2026-08-17. Delivered end to end, twice, headers captured.
smartd monitoring Accepted 2026-08-17. Four members monitored, four alerts received.
Backup infrastructure Postponed by operator decision — strategy may change

1. Instance values

The srv-b bindings for the parameters in ENVIRONMENT.md.

Host

hostname          srv-b  /  srv-b.dev.infra
platform          Proxmox VE 8.4.0, Debian 12, kernel 6.8.12-9-pve
hardware          HP ProLiant DL360 G7, 2 x Xeon X5650, 24 threads, 31 GiB
storage           P410i, 4 x EG0146FAWHU, RAID 1+0, all members SMART OK
local             directory /var/lib/vz, ~70 GiB free   iso,vztmpl,backup
local-lvm         LVM-thin pve/data, 166.9 GiB          rootdir,images
vmbr0             10.0.0.12/24 on enp3s0f0, gw 10.0.0.1   management, LAN
vmbr1             10.20.0.1/24, bridge-ports none         service, portless
wg0               10.110.0.12/32, peer wg-pk.civicus.us:51820, allowed 10.110.0.0/22
resolver          75.75.75.75, search dev.infra
ip_forward        1
timezone          America/Chicago, NTP active, clock synchronised
systemd           running, zero failed units
spare NICs        enp3s0f1, enp4s0f0, enp4s0f1 — unconfigured, deliberately
template          local:vztmpl/debian-12-standard_12.12-1_amd64.tar.zst

Containers

CT 100 CT 101
hostname mechcomp mcproxy
role application, worker reverse proxy, TLS
features nesting=1,keyctl=1 nesting=1
cores / RAM / swap 16 / 16384 / 4096 2 / 1024 / 512
rootfs 40 GiB local-lvm 8 GiB local-lvm
mp0 60 GiB → /var/lib/mechcomp, backup=0 —
net0 eth0 on vmbr1, 10.20.0.10/24, gw 10.20.0.1 eth0 on vmbr1, 10.20.0.11/24, gw 10.20.0.1
onboot 1 1
Debian 12 bookworm, fully upgraded 12 bookworm, fully upgraded
systemd running, zero failed units running, zero failed units

Neither container has a LAN interface (F-017). Neither has a stale eth1 stanza (F-021) — verified persistent across reboot.

Names and identity

mechanical-compiler.dev.infra   ->  10.20.0.11    (CT 101, the proxy)
mechcomp.dev.infra              ->  10.20.0.10    PVE-generated
mcproxy.dev.infra               ->  10.20.0.11    PVE-generated

ADMIN_USER      sandor, uid 1000, groups sandor + mechcomp(996)
                shell /bin/bash, both containers
authorized key  SHA256:2pNffCscUUW5Wbs9uMepLEvMLKPqUV/7Lpk5tw7iWSY
                matches /root/.ssh/id_rsa.pub on srv-b — bastion, see section 4
service user    mechcomp, uid 999, gid 996
                home /var/www/mechcomp, shell /usr/sbin/nologin

Host NAT, persisted in /etc/iptables/rules.v4

nat POSTROUTING
  -s 10.20.0.0/24 -d 10.0.0.0/24            -j RETURN       # LAN not translated
  -s 10.0.0.0/24                  -o wg0    -j MASQUERADE   # pre-existing
  -s 10.20.0.0/24                 -o wg0    -j MASQUERADE   # containers -> WireGuard
  -s 10.20.0.0/24                 -o vmbr0  -j MASQUERADE   # containers -> internet

filter FORWARD
  -s 10.20.0.0/24 -d 10.110.0.0/22 -p tcp
     -m multiport --dports 25,465,587       -j DROP         # SMTP egress blocked
  -s 10.20.0.0/24 -d 10.20.0.0/24           -j ACCEPT       # container <-> container
  -s 10.20.0.0/24 -d 10.0.0.0/24            -j DROP         # LAN blocked

Rule order is load-bearing. RETURN must precede both service-network masquerades. Live and persisted states match.

Disk monitoring — accepted 2026-08-17

DEVICESCAN finds nothing behind the P410i. Members are addressed explicitly:

/etc/smartd.conf
DEFAULT -a -m root -M exec /usr/share/smartmontools/smartd-runner
/dev/sda -d cciss,0
/dev/sda -d cciss,1
/dev/sda -d cciss,2
/dev/sda -d cciss,3

before:  Monitoring 0 ATA/SATA, 0 SCSI/SAS and 0 NVMe devices
after:   Monitoring 0 ATA/SATA, 4 SCSI/SAS and 0 NVMe devices
unit:    smartmontools.service   (smartd.service is an alias with no journal)
alerts:  4 test alerts generated and received; headers captured
         -M test reverted; permanent config restored and reboot-proven

Physical identity of the four members. This is how a future SMART alert is matched to a caddy:

Member Model Serial
cciss,0 EG0146FAWHU 6SD0LETK0000B05009W1
cciss,1 EG0146FAWHU 3SD3K5P00000905094UN
cciss,2 EG0146FAWHU 3SD3GHJ600009047RRRG
cciss,3 EG0146FAWHU 3SD3GVP800009045X79B

Not yet captured: power-on hours and grown-defect counts. On 2010-vintage drives holding the only instance, those numbers say how much runway remains. smartctl -d cciss,N -A /dev/sda for each member.

TLS

CA        CN = Mechanical Compiler Staging CA   (locally generated)
leaf      CN = mechanical-compiler.dev.infra
          SAN = DNS:mechanical-compiler.dev.infra
validity  2026-08-16 -> 2028-11-18            <-- nothing renews this
trusted   srv-b, CT 100, CT 101
key       root:root 0600 /etc/ssl/mechcomp/server.key

No Kane County Civic Infrastructure CA issuance path exists on srv-b: no trust anchor, no step, no cfssl, no EasyRSA. Proxmox's own CA was deliberately not reused. Replacing this leaf later is two file copies and a reload.

Application environment

/etc/mechcomp/mechcomp.env, root:mechcomp, 0640:

MECHCOMP_ENV=staging
MECHCOMP_BIND=10.20.0.10          # scalar; loopback is NOT bound
MECHCOMP_PORT=8770
MECHCOMP_BASE_URL=https://mechanical-compiler.dev.infra
MECHCOMP_DATA_DIR=/var/lib/mechcomp
MECHCOMP_LOG_LEVEL=info
MECHCOMP_DB_URL=sqlite:////var/lib/mechcomp/db/mechcomp.sqlite3
MECHCOMP_WORKER_CONCURRENCY=4
MECHCOMP_CAD_BACKEND=none
MECHCOMP_ARTIFACT_RETENTION_DAYS=30
MECHCOMP_SECRET_KEY=<generated, not recorded>

Mail — accepted 2026-08-17

srv-b relayhost   [10.110.0.1]:25 over wg0
root alias        sandor@kane-il.us   (postalias verified)
inet_interfaces   loopback-only
delivery path     srv-b -> wg-pk -> mx1 -> kane-il.us
status            DELIVERED end to end, twice, full headers captured
                  queue empty, no bounced or deferred entries

F-023 was proven to be a relay-trust mismatch three hops downstream: mx1 rejected RCPT TO with 554 5.7.1 Access denied because wg-pk's public addresses were absent from its mynetworks. Corrected by adding only those two addresses.

Standing constraints on this path — do not violate:

Host Constraint
kane-il.us Runs YunoHost. Its mail configuration must not be altered. This is why mx1 exists.
mx1.diagnostics.kane-il.us Uses DANE. Its TLS and certificate configuration must not be broken. The mynetworks change is inbound client trust and does not interact with DANE.
wg-pk Must continue relaying to arbitrary external destinations. Hubzilla registration and notification mail depends on it, and the ISP blocks port 25. Do not restrict it by recipient.

Fragility to know about: delivery now depends on Linode not reassigning 198.58.111.109 or 2600:3c00::f03c:92ff:fe42:43d7. If either changes, mail stops silently and the cause is in mx1's mynetworks.

Open exposure — see F-025. The correction authorises wg-pk under permit_mynetworks, which grants relay to any destination. Combined with wg-pk's own mynetworks = 10.110.0.0/22, every tunnel peer — including both containers, which arrive as 10.110.0.12 through the host masquerade — can originate mail as kane-il.us infrastructure. Correction pending decision.

Logging note (F-024): /var/log/mail.log does not exist. Proxmox ships without rsyslog; Postfix logs to journald. Use journalctl -u postfix@-.

Placeholder backend — staging scaffold, not application code

/usr/local/libexec/mechcomp-placeholder.py
/etc/systemd/system/mechcomp-placeholder.service
runs as mechcomp:mechcomp, reads mechcomp.env, binds 10.20.0.10:8770
Restart=on-failure, hardening set applied, enabled at multi-user.target

Exists solely to prove the proxy chain independently of the application. It is removed when the real service arrives.

Retired 2026-09-11. mechcomp-placeholder.service is disabled and stopped. mechcomp.service holds 10.20.0.10:8770 in its place, running /var/www/mechcomp/venv/bin/python -m mechcomp.web as mechcomp, reading /etc/mechcomp/mechcomp.env for its binding. nginx on CT 101 required no change: the real service took the address the placeholder occupied.

The placeholder unit file is left on disk, disabled. It is the rollback: systemctl disable --now mechcomp.service then systemctl enable --now mechcomp-placeholder.service.

Proven end to end from srv-b at commit 14514b0: HTTPS 200 for the page and a real JSON build payload through the TLS chain.

Rollback copies on disk

srv-b     /etc/network/interfaces.before-mechcomp
          /etc/hosts.before-mechcomp
          /etc/hosts.before-svcfqdn
          /etc/iptables/rules.v4.before-svcnat
          /etc/iptables/rules.v4.before-lanblock
          /root/pct-100.before-svcnet
          /root/pct-101.before-svcnet
CT 100    /etc/hosts.before-svcfqdn
          /etc/network/interfaces pre-F-021 copy
CT 101    /etc/hosts.before-svcfqdn
          /etc/network/interfaces pre-F-021 copy

2. Verified

Each line was demonstrated by command output, not inferred.

Host and network

  • vmbr1 created, active, 10.20.0.1/24, portless
  • Duplicate address detection run before each container creation
  • Container to internet reachable (deb.debian.org, gitea.barternetwork.us)
  • Container to WireGuard reachable (10.110.0.1)
  • Container to LAN gateway 10.0.0.1 = 100% packet loss — the requirement
  • Container to srv-b 10.0.0.12 reachable via INPUT path (required, not a leak)
  • Container to container reachable over vmbr1
  • LAN to Proxmox console 10.0.0.12:8006 returns 200, unaffected
  • NAT and FORWARD rules present live and persisted, in correct order
  • Host: running, zero failed units
  • Mail delivered end to end from root on srv-b, twice, headers captured
  • smartd monitoring 4 devices; 4 test alerts delivered and received
  • Container SMTP egress blocked; internet egress and host mail intact
  • Both containers autostart after a host reboot (onboot: 1)

CT 100

  • Debian 12 bookworm, fully upgraded, zero pending
  • en_US.UTF-8 generated and active; America/Chicago
  • running, zero failed units — after reboot (F-021 corrected)
  • /var/lib/mechcomp is a real separate ext4 filesystem
  • Service user mechcomp, home /var/www/mechcomp, correct ownership
  • All four data subdirectories present, mechcomp:mechcomp, 0750
  • mechcomp.env present, root:mechcomp, 0640, secret generated
  • Base packages installed; OpenSCAD, Qt and X11 absent
  • Docker active, overlay2 / systemd, no fallback needed
  • Repository cloned at e85c4f4e, verified as the owning user
  • Python 3.11.2 venv created, owned by mechcomp
  • openssh-server enabled and active
  • mechcomp-placeholder.service enabled, active, reboot-persistent (F-019)
  • Listener on 10.20.0.10:8770 only — not 0.0.0.0

CT 101

  • Debian 12 bookworm, fully upgraded, zero pending
  • en_US.UTF-8 generated and active; America/Chicago
  • running, zero failed units — after reboot (F-021 corrected)
  • nginx 1.22.1, nginx -t passes, enabled and active
  • Only the project vhost enabled; Debian default removed
  • listen 443 ssl default_server on both address families
  • HTTP to HTTPS 301 redirect
  • Local CA created, leaf issued, trusted on all three hosts
  • HTTPS end to end from srv-b, CT 100 and CT 101 without -k
  • X-Forwarded-Proto: https observed at the backend — re-proved after the topology change and again after CT 100's reboot
  • openssh-server enabled and active

3. Remaining — infrastructure

Immediate

Nothing. The infrastructure boundary is accepted.

Deferred by operator decision

Recorded here so they are visually distinct from failures and from forgotten work. None is a defect.

  • wg-pk mynetworks scope (F-025). Estate decision: whether every tunnel peer should originate mail as kane-il.us infrastructure, or only the hosts that legitimately do. Not this project's to change unilaterally.
  • SMART attribute baseline. Power-on hours and grown-defect counts for the four members, so drive age is a known quantity rather than an assumption. smartctl -d cciss,N -A /dev/sda. Read-only, one command.
  • Backup infrastructure, entirely. Postponed 2026-08-16 because the strategy may change, and reaffirmed 2026-08-18 on stronger grounds: a backup of an unverified configuration restores the confusion along with the data. Entry condition: ct-baseline.sh exits 0. Scope when resumed: vzdump job, archive sizing, retention, free-space guard, host-side pull, mechcomp-backup, backup alerting, gold media, 3+ TB redundancy.

One request standing against the postponement: a single manual vzdump of both containers as a point-in-time snapshot before further change. It presupposes nothing about the eventual strategy and yields the archive size that has been an open question since revision 4. Four 2010-vintage disks currently hold the only instance with no monitoring, no alerting and no backup.

Optional, recorded not scheduled

  • Second nameserver line in container resolver configuration, as a mitigation for F-022. One line, no daemon. Deliberately not applied on a single unexplained event.
  • Staging certificate expires 2028-11-18 and nothing renews it.

3b. Container baseline

ct-baseline.sh is the definition of "standard configuration" on this host. It changes nothing, runs any time, and exits 0 only when every container conforms.

last run   2026-08-18
result     62 passed, 0 failed, 3 informational
scope      srv-b, CT 100, CT 101, CT 102

A property not checked by that script is not part of the standard. That is what makes it terminate. Divergence used to be discovered by asking, one property at a time, whenever something behaved oddly — which never finishes, because nothing stated what "the same" meant. If a property should be uniform, it belongs in the script rather than in someone's memory.

What it asserts, per container: onboot, unprivileged, nesting=1, a single interface on the service bridge, systemd running with zero failed units, one interface stanza, Debian 12, timezone, generated locale, curl, ca-certificates, openssh-server, sshd active, admin account with authorized_keys at 0600, no mail agent, SMTP egress blocked, LAN gateway unreachable. Host-level: systemd health, both firewall rules, smartd device count, bastion key.

Each failure cites the failure entry that established the requirement, so the reason survives the person.

The mail standard

Containers do not send mail. Only srv-b does, and only its own alerts, through [10.110.0.1]:25. A container with a mail agent installed but SMTP egress blocked is the worst case: it queues forever and delivers nothing.

CT 100 and CT 101 were in exactly that state until 2026-08-18 (F-031).

Ownership

The script checks containers belonging to two projects and encodes host requirements neither owns alone (F-032). It is host-level, not Mechanical Compiler property. Suggested home: /usr/local/sbin/ct-baseline.sh on srv-b, with the canonical copy outside this repository.

Relationship to backup

Conformance is a precondition for backup, not a companion to it. A backup taken before 2026-08-18 would have preserved two containers that queue mail forever, and a restore would have faithfully returned them to that state.

Backing up an unverified configuration does not preserve a system; it preserves a state of confusion. The entry condition for backup work is ct-baseline.sh exiting 0 — and it becomes part of restore verification.


3a. Parallel project on this host

Kane Fabric is a separate project sharing srv-b. It is where SASE, HOA Diagnostics, the Mechanical Compiler and other SASE-consuming projects are implemented.

Repository github.com/git64bit/Kane-Fabric
Container CT 102 kane-fabric
Address 10.20.0.12/24 on vmbr1, gateway 10.20.0.1
Ownership Separate project. Not Mechanical Compiler infrastructure.
Baseline Conforms as of 2026-08-18, verified by ct-baseline.sh

Recorded here for two operational reasons only.

It shares the service network. CT 102 sits on vmbr1 alongside CT 100 and CT 101, so host-level rules scoped to 10.20.0.0/24 apply to it. In particular it inherits the SMTP egress block (F-025) — if Kane Fabric ever needs to send mail, that will present as a Postfix failure with a non-obvious cause.

Host resources are shared, not pooled. Neither project consumes or repurposes the other's containers. Any integration is through an explicit contract, not through co-location on srv-b.

No interface between the two projects exists today, and none is assumed.


4. Access model

Confirmed: no workstation access from the home LAN is required.

internet -> WireGuard -> srv-b -> containers

srv-b is the bastion, confirmed by the authorized key matching /root/.ssh/id_rsa.pub on the host. Root on the host can pct enter regardless, so SSH adds no privilege — but every path to a container runs through srv-b, which is deliberate.

Direct WireGuard-side access to the catalogue would be a route addition for 10.20.0.0/24 on the hub. Not now.

Reopened 2026-09-11 by WORK-ORDER-004. That decision was correct while there was no application to reach. There is one now, and the consequence of the decision is that CIVICVS cannot see it: mechanical-compiler.dev.infra resolves only on srv-b, CT 100 and CT 101, and 10.20.0.0/24 sits behind a portless bridge with LAN traffic dropped.

The srv-b half needs no change, verified 2026-09-11: ip_forward=1, FORWARD policy ACCEPT, a direct route on vmbr1, and none of the three FORWARD rules matches hub-initiated inbound traffic. The single gate is AllowedIPs on the hub's peer entry for srv-b, currently 10.110.0.0/22 -- WireGuard drops by cryptokey routing before consulting any routing table, so a route without that entry does nothing.

No DHCP anywhere. Two containers with fixed addresses on a portless bridge is the entire address space.


5. Blocked on application code

Repository state:

HEAD        c7e32d8e0772d89d1117bab7ff08d61e5b2a497e
top level   .gitignore LICENSE Makefile README.md docs/ fixtures/ legacy/
            pyproject.toml requirements-*.txt src/ tests/ tools/ venv/
remote      ssh://git@gitea.barternetwork.us:42022/TheRON/mechanical-compiler.git
            deploy key srv-b-ct100, read/write
venv        Python 3.11.2, shapely 2.1.2, cadquery 2.8.0, mechcomp editable
tests       3 passed, 236 skipped
oracle      ddd0f154... verified in-container
absent      the Shapely port itself

None of the following can be honestly completed, and none may be fabricated by provisioning:

  • Dependency install from committed manifests — done 2026-08-18
  • mechcomp.service — done 2026-09-11, deploy/mechcomp.service
  • mechcomp-worker.service — no worker exists yet; src/mechcomp/worker/ is still a stub
  • Reference toolchain image mechcomp/reference-toolchain:8.0.0
  • Fixture reproduction against ddd0f154...
  • Application-runtime acceptance — partial. The composer serves and is reachable through CT 101; no acceptance criteria have been written for it, and it is not reachable from outside (see section 6, question 4)
  • Replacing the placeholder with the real service — done 2026-09-11

Applied 2026-08-18: venv/ is in .gitignore, along with .cache/, .local/, .ssh/, .gitconfig and .lesshst — the service user's home is install_dir, so anything writing to $HOME writes into the working tree (F-029).

The Shapely port gates all of the above.


6. Open questions

# Question Blocks
1 Should wg-pk mynetworks narrow to explicit hosts? F-025 estate half
2 What is the backup strategy? all backup work
3 Where is the 3+ TB USB disk attached? gold redundancy step
4 Public ingress for the composer at dev.mechcomp.kane-il.us WORK-ORDER-004; anyone outside srv-b seeing the application at all

All three need CIVICVS.


7. Closed questions

Question Answer
Kane County CA issuance path None on srv-b. Local staging CA generated.
Relay address and authentication 10.110.0.1:25 over wg0, unauthenticated from 10.110.0.12 accepted.
openssh-server present Yes, both containers.
ADMIN_USER / key sandor; root@srv-b key, bastion pattern confirmed.
DHCP pool on 10.0.0.0/24 Not applicable. Containers are not on that network.
Docker storage driver overlay2 / systemd. No fallback needed.
LAN workstation access Not required. WireGuard through srv-b.
MECHCOMP_BIND semantics Scalar, service address only. Loopback not bound.
Why did delivery fail downstream? mx1 relay trust. Proven and corrected (F-023).
Where does Postfix log on this host? journald. No rsyslog, no /var/log/mail.log (F-024).
Should container SMTP egress be blocked? Yes. Implemented and reboot-proven (F-025 project-local half).
Do the containers autostart? Yes, onboot: 1 on all three, proven by host reboot (F-026).
How do files reach a container? Upload to /root/incoming on srv-b, then pct push / pct exec. See PROCESS.md.
How does CT 100 push to Gitea? Deploy key srv-b-ct100, SSH on port 42022, read/write.
Do containers send mail? No. Only srv-b does. See section 3b.