From a8081e17dc2acb2c6b3a1c89841014de57bcc4f9 Mon Sep 17 00:00:00 2001 From: TheRON Date: Fri, 11 Sep 2026 06:59:53 -0500 Subject: [PATCH] state: the composer replaced the placeholder; open the public ingress question PROCESS.md section 8 requires host changes to reach STAGING-STATE.md before a session ends. mechcomp-placeholder.service was disabled and stopped on 2026-09-11 and mechcomp.service took 10.20.0.10:8770 in its place. nginx on CT 101 needed no change because the real service took the address the placeholder occupied. The placeholder unit stays on disk, disabled, as the rollback. Section 5 items closed: mechcomp.service, and replacing the placeholder. Application runtime acceptance is marked partial rather than done, because no acceptance criteria have been written for the composer and it is not reachable from outside. WORK-ORDER-004 publishes the composer at dev.mechcomp.kane-il.us. The srv-b half needs no change, verified: forwarding on, FORWARD policy ACCEPT, a direct route on vmbr1, and none of the three FORWARD rules matches hub initiated inbound traffic. The single gate is AllowedIPs on the hub peer entry for srv-b, because WireGuard drops by cryptokey routing before consulting any routing table. Design decision recorded: the public name terminates on the hub and proxies to CT 100 directly rather than through CT 101. Routing it through CT 101 would make the public path depend on a locally signed leaf that expires 2028-11-18 with nothing renewing it. The tunnel already provides the encryption that hop would add. CT 101 keeps serving the internal name. Section 4 not now decision on the hub route is reopened. It was correct while there was no application to reach. The consequence of leaving it closed is that the operator cannot see the application at all. --- docs/STAGING-STATE.md | 39 ++++- docs/WORK-ORDER-004-public-ingress.md | 226 ++++++++++++++++++++++++++ 2 files changed, 261 insertions(+), 4 deletions(-) create mode 100644 docs/WORK-ORDER-004-public-ingress.md diff --git a/docs/STAGING-STATE.md b/docs/STAGING-STATE.md index a2e365a..9e913e3 100644 --- a/docs/STAGING-STATE.md +++ b/docs/STAGING-STATE.md @@ -4,7 +4,7 @@ Live state of the Mechanical Compiler staging instance on `srv-b`. | | | |---|---| -| Updated | 2026-08-18, after repository seeding and container standardisation | +| Updated | 2026-09-11, after the composer replaced the placeholder | | Instance | Staging / development | | Specification | `ENVIRONMENT.md` revision 5 | | Failure log | `FAILURES.md` | @@ -238,6 +238,19 @@ Restart=on-failure, hardening set applied, enabled at multi-user.target Exists solely to prove the proxy chain independently of the application. It is removed when the real service arrives. +**Retired 2026-09-11.** `mechcomp-placeholder.service` is disabled and stopped. +`mechcomp.service` holds `10.20.0.10:8770` in its place, running +`/var/www/mechcomp/venv/bin/python -m mechcomp.web` as `mechcomp`, reading +`/etc/mechcomp/mechcomp.env` for its binding. nginx on CT 101 required no +change: the real service took the address the placeholder occupied. + +The placeholder unit file is left on disk, disabled. It is the rollback: +`systemctl disable --now mechcomp.service` then +`systemctl enable --now mechcomp-placeholder.service`. + +Proven end to end from `srv-b` at commit `14514b0`: `HTTPS 200` for the page +and a real JSON build payload through the TLS chain. + ### Rollback copies on disk ``` @@ -452,6 +465,19 @@ so SSH adds no privilege — but **every path to a container runs through Direct WireGuard-side access to the catalogue would be a route addition for `10.20.0.0/24` on the hub. Not now. +**Reopened 2026-09-11 by `WORK-ORDER-004`.** That decision was correct while +there was no application to reach. There is one now, and the consequence of +the decision is that CIVICVS cannot see it: `mechanical-compiler.dev.infra` +resolves only on `srv-b`, CT 100 and CT 101, and `10.20.0.0/24` sits behind a +portless bridge with LAN traffic dropped. + +The `srv-b` half needs no change, verified 2026-09-11: `ip_forward=1`, +`FORWARD` policy `ACCEPT`, a direct route on `vmbr1`, and none of the three +`FORWARD` rules matches hub-initiated inbound traffic. The single gate is +`AllowedIPs` on the hub's peer entry for `srv-b`, currently `10.110.0.0/22` -- +WireGuard drops by cryptokey routing before consulting any routing table, so a +route without that entry does nothing. + No DHCP anywhere. Two containers with fixed addresses on a portless bridge is the entire address space. @@ -477,11 +503,15 @@ None of the following can be honestly completed, and none may be fabricated by provisioning: - [x] ~~Dependency install from committed manifests~~ — done 2026-08-18 -- [ ] `mechcomp.service`, `mechcomp-worker.service` +- [x] ~~`mechcomp.service`~~ — done 2026-09-11, `deploy/mechcomp.service` +- [ ] `mechcomp-worker.service` — no worker exists yet; `src/mechcomp/worker/` + is still a stub - [ ] Reference toolchain image `mechcomp/reference-toolchain:8.0.0` - [ ] Fixture reproduction against `ddd0f154...` -- [ ] Application-runtime acceptance -- [ ] Replacing the placeholder with the real service +- [ ] Application-runtime acceptance — partial. The composer serves and is + reachable through CT 101; no acceptance criteria have been written for + it, and it is not reachable from outside (see section 6, question 4) +- [x] ~~Replacing the placeholder with the real service~~ — done 2026-09-11 **Applied 2026-08-18:** `venv/` is in `.gitignore`, along with `.cache/`, `.local/`, `.ssh/`, `.gitconfig` and `.lesshst` — the service user's home is @@ -499,6 +529,7 @@ The Shapely port gates all of the above. | 1 | Should `wg-pk` `mynetworks` narrow to explicit hosts? | F-025 estate half | | 2 | What is the backup strategy? | all backup work | | 3 | Where is the 3+ TB USB disk attached? | gold redundancy step | +| 4 | Public ingress for the composer at `dev.mechcomp.kane-il.us` | `WORK-ORDER-004`; anyone outside `srv-b` seeing the application at all | All three need CIVICVS. diff --git a/docs/WORK-ORDER-004-public-ingress.md b/docs/WORK-ORDER-004-public-ingress.md new file mode 100644 index 0000000..abbf74f --- /dev/null +++ b/docs/WORK-ORDER-004-public-ingress.md @@ -0,0 +1,226 @@ +# WORK-ORDER-004 — public ingress for the composer + +Publish the Mechanical Compiler composer at `dev.mechcomp.kane-il.us`, reachable +from an ordinary browser on the internet. + +| | | +|---|---| +| Mode | Infrastructure (`PROCESS.md` §2) | +| Created | 2026-09-11 | +| Executed on | `wg-pk` (the WireGuard hub), via its Webmin terminal | +| Depends on | `14514b0` — composer live on CT 100, proven through CT 101 | + +--- + +## 0. Read this before the first command + +**This work order is executed on a machine outside this project.** `wg-pk` +carries `kane-il.us` mail and Hubzilla. `PROCESS.md` §7 requires escalation for +exactly this, and the escalation is: CIVICVS executes, one group at a time, and +anything surprising stops the work rather than being worked around. + +### The lockout risk, and why it is bounded + +Step 2 changes a live WireGuard peer. If it goes wrong the tunnel drops, and +with it every path to `srv-b` — `PROCESS.md` §1 records that the `srv-b` shell +is the operator's entire working surface for this project. + +It is recoverable because **the operator is sitting on `wg-pk` itself**, in +Webmin, not reaching it through the tunnel. `wg set` is not persistent, so +`systemctl restart wg-quick@wg0` on the hub restores the on-disk configuration +and the tunnel with it. + +Do not perform step 2 from a shell that reaches the hub through the tunnel. + +### What `srv-b` contributes + +**Nothing. It is already correct and must not be touched.** Verified +2026-09-11: + +``` +net.ipv4.ip_forward = 1 +-P FORWARD ACCEPT +10.20.0.10 dev vmbr1 src 10.20.0.1 +``` + +The three `FORWARD` rules block container-sourced SMTP, container-to-container +(ACCEPT), and container-to-LAN. None matches hub-initiated inbound traffic. The +`-s 10.20.0.0/24 -o wg0 MASQUERADE` rule does not apply either: the first packet +of a hub-initiated flow is not container-sourced, so no NAT binding is created +and replies return through conntrack. + +**Explicit do-not-touch list:** `srv-b` iptables, `srv-b` WireGuard, CT 101 +nginx, CT 101 TLS, the `kane-il.us` MX records, anything on +`mx1.diagnostics.kane-il.us` (it publishes a TLSA record — confirmed +2026-09-11). + +--- + +## 1. Topology, and why + +``` +browser + -> DNS dev.mechcomp.kane-il.us -> wg-pk public address + -> nginx on wg-pk, Let's Encrypt TLS terminated here + -> WireGuard tunnel, already encrypted + -> srv-b 10.110.0.12, forwards, no configuration change + -> CT 100 10.20.0.10:8770, the composer +``` + +**The public path deliberately bypasses CT 101.** Routing it through CT 101 +would make the public name depend on a locally-signed leaf valid to 2028-11-18 +with nothing renewing it — a dated outage designed in from the start. The tunnel +already provides the encryption that hop would add. + +CT 101 continues to serve `mechanical-compiler.dev.infra` for work from inside. +Two ingresses, each with a distinct reason to exist. + +--- + +## 2. Success criteria, stated before the work + +1. `curl -I https://dev.mechcomp.kane-il.us/` returns `200` from a machine with + no WireGuard access and no special DNS. +2. The certificate is issued by Let's Encrypt and chains without `-k`. +3. `https://dev.mechcomp.kane-il.us/api/build?family=3x&profile=Y` returns JSON + whose `groups` end with a `Y` group and contain no `three_fin` key. +4. The page renders, controls change the drawing, and a changed control changes + the `design` identity shown beneath it. +5. **The negative:** `mechanical-compiler.dev.infra` still answers from `srv-b`, + and mail from `srv-b` still delivers. Neither path was in scope; both must be + proven unharmed. +6. Every change survives `systemctl restart wg-quick@wg0` and a hub reboot. + +--- + +## 3. Command groups + +One at a time. Paste output back before the next. + +### Group 1 — read-only, on `wg-pk` + +Nothing here changes anything. + +```bash +echo "=== peer entry for srv-b ===" && \ +wg show && \ +echo "=== on-disk wireguard config ===" && \ +grep -n "AllowedIPs\|PublicKey\|Address\|PostUp" /etc/wireguard/wg0.conf && \ +echo "=== routing toward the service network ===" && \ +ip route | grep -E "10\.20\.|10\.110\." ; \ +echo "=== forwarding ===" && \ +sysctl net.ipv4.ip_forward && \ +echo "=== nginx vhost conventions, following symlinks ===" && \ +grep -rn --dereference-recursive "server_name\|proxy_pass\|listen\|ssl_certificate " \ + /etc/nginx/sites-enabled/ | head -40 && \ +echo "=== certificate issuance ===" && \ +ls /etc/letsencrypt/live 2>/dev/null || echo "no certbot live dir" ; \ +which certbot ; \ +echo "=== does the name resolve yet ===" && \ +dig +short dev.mechcomp.kane-il.us A ; \ +dig +short dev.mechcomp.kane-il.us AAAA +``` + +What matters in the output: whether the peer's `AllowedIPs` is the only gate, +whether nginx already has `listen [::]:443` on its vhost pattern (it must, if +the name inherits an AAAA), how certificates are issued, and whether the name +resolves yet. + +**Report this before proceeding.** Groups 2 onward are written against what it +shows; they are deliberately not drafted in advance. + +### Group 2 — WireGuard AllowedIPs + +Drafted after group 1. The shape: + +- Back up `/etc/wireguard/wg0.conf` first. `PROCESS.md` §2: preserve a rollback + copy before editing any configuration file. +- Add `10.20.0.0/24` to the `srv-b` peer's `AllowedIPs`, **keeping + `10.110.0.0/22`**. Replacing rather than extending it is the way this breaks. +- Apply live, confirm the tunnel is still up and `srv-b` still reachable, then + persist. +- Prove: `ping -c2 10.20.0.10` from the hub. + +### Group 3 — route + +A route for `10.20.0.0/24` via `10.110.0.12`, persisted the way the hub already +persists routes — which group 1 reveals. Prove with +`curl -sS -o /dev/null -w '%{http_code}\n' http://10.20.0.10:8770/`. + +### Group 4 — certificate + +Let's Encrypt HTTP-01 for `dev.mechcomp.kane-il.us`. A third-level name needs no +wildcard and no DNS-01. Requires the name to resolve to the hub first. + +### Group 5 — vhost + +```nginx +# /etc/nginx/sites-available/dev.mechcomp.kane-il.us +# Matching whatever pattern group 1 shows the existing vhosts use. +server { + listen 80; + listen [::]:80; + server_name dev.mechcomp.kane-il.us; + return 301 https://$host$request_uri; +} + +server { + listen 443 ssl; + listen [::]:443 ssl; # required if the name has an AAAA + server_name dev.mechcomp.kane-il.us; + + ssl_certificate /etc/letsencrypt/live/dev.mechcomp.kane-il.us/fullchain.pem; + ssl_certificate_key /etc/letsencrypt/live/dev.mechcomp.kane-il.us/privkey.pem; + + location / { + proxy_pass http://10.20.0.10:8770; + proxy_set_header Host $host; + proxy_set_header X-Real-IP $remote_addr; + proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for; + proxy_set_header X-Forwarded-Proto $scheme; + + # The composer redraws on every control change. Small responses, many of + # them; buffering adds latency for no benefit. + proxy_buffering off; + proxy_read_timeout 30s; + } +} +``` + +`nginx -t` before reload, always. + +### Group 6 — acceptance + +All six criteria in §2, including the two negatives. + +--- + +## 4. Known unknowns + +Stated rather than assumed, because assuming is what produced three unreachable +URLs before this document existed. + +- **Whether the hub's nginx listens on v6.** If `dev.mechcomp` is a CNAME to a + name carrying an AAAA, v6 is published whether or not nginx answers on it. + v6-preferring browsers then fail while v4 ones succeed — a fault that looks + like anything except DNS. +- **How the hub persists routes.** Distribution- and tooling-dependent; group 1 + shows it. +- **Whether Webmin or YunoHost manages the nginx configuration.** If so, + hand-written vhosts may be overwritten on their next reconfiguration, and the + vhost belongs wherever that system expects it instead. + +--- + +## 5. Acceptance and reporting + +Closed when all six criteria in §2 hold and survive a hub reboot. + +Back to the architect: the output of every group, and any surprise about the +environment in `FAILURES.md` format — `PROCESS.md` §7. A surprise on the hub is +still a finding even though the hub is not this project's property. + +`STAGING-STATE.md` §4 records "direct WireGuard-side access to the catalogue +would be a route addition for `10.20.0.0/24` on the hub. Not now." That decision +was correct when there was no application to reach. This work order is its +reversal, and §4 should say so when this closes.