diff --git a/docs/STAGING-STATE.md b/docs/STAGING-STATE.md index a2e365a..9e913e3 100644 --- a/docs/STAGING-STATE.md +++ b/docs/STAGING-STATE.md @@ -4,7 +4,7 @@ Live state of the Mechanical Compiler staging instance on `srv-b`. | | | |---|---| -| Updated | 2026-08-18, after repository seeding and container standardisation | +| Updated | 2026-09-11, after the composer replaced the placeholder | | Instance | Staging / development | | Specification | `ENVIRONMENT.md` revision 5 | | Failure log | `FAILURES.md` | @@ -238,6 +238,19 @@ Restart=on-failure, hardening set applied, enabled at multi-user.target Exists solely to prove the proxy chain independently of the application. It is removed when the real service arrives. +**Retired 2026-09-11.** `mechcomp-placeholder.service` is disabled and stopped. +`mechcomp.service` holds `10.20.0.10:8770` in its place, running +`/var/www/mechcomp/venv/bin/python -m mechcomp.web` as `mechcomp`, reading +`/etc/mechcomp/mechcomp.env` for its binding. nginx on CT 101 required no +change: the real service took the address the placeholder occupied. + +The placeholder unit file is left on disk, disabled. It is the rollback: +`systemctl disable --now mechcomp.service` then +`systemctl enable --now mechcomp-placeholder.service`. + +Proven end to end from `srv-b` at commit `14514b0`: `HTTPS 200` for the page +and a real JSON build payload through the TLS chain. + ### Rollback copies on disk ``` @@ -452,6 +465,19 @@ so SSH adds no privilege — but **every path to a container runs through Direct WireGuard-side access to the catalogue would be a route addition for `10.20.0.0/24` on the hub. Not now. +**Reopened 2026-09-11 by `WORK-ORDER-004`.** That decision was correct while +there was no application to reach. There is one now, and the consequence of +the decision is that CIVICVS cannot see it: `mechanical-compiler.dev.infra` +resolves only on `srv-b`, CT 100 and CT 101, and `10.20.0.0/24` sits behind a +portless bridge with LAN traffic dropped. + +The `srv-b` half needs no change, verified 2026-09-11: `ip_forward=1`, +`FORWARD` policy `ACCEPT`, a direct route on `vmbr1`, and none of the three +`FORWARD` rules matches hub-initiated inbound traffic. The single gate is +`AllowedIPs` on the hub's peer entry for `srv-b`, currently `10.110.0.0/22` -- +WireGuard drops by cryptokey routing before consulting any routing table, so a +route without that entry does nothing. + No DHCP anywhere. Two containers with fixed addresses on a portless bridge is the entire address space. @@ -477,11 +503,15 @@ None of the following can be honestly completed, and none may be fabricated by provisioning: - [x] ~~Dependency install from committed manifests~~ — done 2026-08-18 -- [ ] `mechcomp.service`, `mechcomp-worker.service` +- [x] ~~`mechcomp.service`~~ — done 2026-09-11, `deploy/mechcomp.service` +- [ ] `mechcomp-worker.service` — no worker exists yet; `src/mechcomp/worker/` + is still a stub - [ ] Reference toolchain image `mechcomp/reference-toolchain:8.0.0` - [ ] Fixture reproduction against `ddd0f154...` -- [ ] Application-runtime acceptance -- [ ] Replacing the placeholder with the real service +- [ ] Application-runtime acceptance — partial. The composer serves and is + reachable through CT 101; no acceptance criteria have been written for + it, and it is not reachable from outside (see section 6, question 4) +- [x] ~~Replacing the placeholder with the real service~~ — done 2026-09-11 **Applied 2026-08-18:** `venv/` is in `.gitignore`, along with `.cache/`, `.local/`, `.ssh/`, `.gitconfig` and `.lesshst` — the service user's home is @@ -499,6 +529,7 @@ The Shapely port gates all of the above. | 1 | Should `wg-pk` `mynetworks` narrow to explicit hosts? | F-025 estate half | | 2 | What is the backup strategy? | all backup work | | 3 | Where is the 3+ TB USB disk attached? | gold redundancy step | +| 4 | Public ingress for the composer at `dev.mechcomp.kane-il.us` | `WORK-ORDER-004`; anyone outside `srv-b` seeing the application at all | All three need CIVICVS. diff --git a/docs/WORK-ORDER-004-public-ingress.md b/docs/WORK-ORDER-004-public-ingress.md new file mode 100644 index 0000000..abbf74f --- /dev/null +++ b/docs/WORK-ORDER-004-public-ingress.md @@ -0,0 +1,226 @@ +# WORK-ORDER-004 — public ingress for the composer + +Publish the Mechanical Compiler composer at `dev.mechcomp.kane-il.us`, reachable +from an ordinary browser on the internet. + +| | | +|---|---| +| Mode | Infrastructure (`PROCESS.md` §2) | +| Created | 2026-09-11 | +| Executed on | `wg-pk` (the WireGuard hub), via its Webmin terminal | +| Depends on | `14514b0` — composer live on CT 100, proven through CT 101 | + +--- + +## 0. Read this before the first command + +**This work order is executed on a machine outside this project.** `wg-pk` +carries `kane-il.us` mail and Hubzilla. `PROCESS.md` §7 requires escalation for +exactly this, and the escalation is: CIVICVS executes, one group at a time, and +anything surprising stops the work rather than being worked around. + +### The lockout risk, and why it is bounded + +Step 2 changes a live WireGuard peer. If it goes wrong the tunnel drops, and +with it every path to `srv-b` — `PROCESS.md` §1 records that the `srv-b` shell +is the operator's entire working surface for this project. + +It is recoverable because **the operator is sitting on `wg-pk` itself**, in +Webmin, not reaching it through the tunnel. `wg set` is not persistent, so +`systemctl restart wg-quick@wg0` on the hub restores the on-disk configuration +and the tunnel with it. + +Do not perform step 2 from a shell that reaches the hub through the tunnel. + +### What `srv-b` contributes + +**Nothing. It is already correct and must not be touched.** Verified +2026-09-11: + +``` +net.ipv4.ip_forward = 1 +-P FORWARD ACCEPT +10.20.0.10 dev vmbr1 src 10.20.0.1 +``` + +The three `FORWARD` rules block container-sourced SMTP, container-to-container +(ACCEPT), and container-to-LAN. None matches hub-initiated inbound traffic. The +`-s 10.20.0.0/24 -o wg0 MASQUERADE` rule does not apply either: the first packet +of a hub-initiated flow is not container-sourced, so no NAT binding is created +and replies return through conntrack. + +**Explicit do-not-touch list:** `srv-b` iptables, `srv-b` WireGuard, CT 101 +nginx, CT 101 TLS, the `kane-il.us` MX records, anything on +`mx1.diagnostics.kane-il.us` (it publishes a TLSA record — confirmed +2026-09-11). + +--- + +## 1. Topology, and why + +``` +browser + -> DNS dev.mechcomp.kane-il.us -> wg-pk public address + -> nginx on wg-pk, Let's Encrypt TLS terminated here + -> WireGuard tunnel, already encrypted + -> srv-b 10.110.0.12, forwards, no configuration change + -> CT 100 10.20.0.10:8770, the composer +``` + +**The public path deliberately bypasses CT 101.** Routing it through CT 101 +would make the public name depend on a locally-signed leaf valid to 2028-11-18 +with nothing renewing it — a dated outage designed in from the start. The tunnel +already provides the encryption that hop would add. + +CT 101 continues to serve `mechanical-compiler.dev.infra` for work from inside. +Two ingresses, each with a distinct reason to exist. + +--- + +## 2. Success criteria, stated before the work + +1. `curl -I https://dev.mechcomp.kane-il.us/` returns `200` from a machine with + no WireGuard access and no special DNS. +2. The certificate is issued by Let's Encrypt and chains without `-k`. +3. `https://dev.mechcomp.kane-il.us/api/build?family=3x&profile=Y` returns JSON + whose `groups` end with a `Y` group and contain no `three_fin` key. +4. The page renders, controls change the drawing, and a changed control changes + the `design` identity shown beneath it. +5. **The negative:** `mechanical-compiler.dev.infra` still answers from `srv-b`, + and mail from `srv-b` still delivers. Neither path was in scope; both must be + proven unharmed. +6. Every change survives `systemctl restart wg-quick@wg0` and a hub reboot. + +--- + +## 3. Command groups + +One at a time. Paste output back before the next. + +### Group 1 — read-only, on `wg-pk` + +Nothing here changes anything. + +```bash +echo "=== peer entry for srv-b ===" && \ +wg show && \ +echo "=== on-disk wireguard config ===" && \ +grep -n "AllowedIPs\|PublicKey\|Address\|PostUp" /etc/wireguard/wg0.conf && \ +echo "=== routing toward the service network ===" && \ +ip route | grep -E "10\.20\.|10\.110\." ; \ +echo "=== forwarding ===" && \ +sysctl net.ipv4.ip_forward && \ +echo "=== nginx vhost conventions, following symlinks ===" && \ +grep -rn --dereference-recursive "server_name\|proxy_pass\|listen\|ssl_certificate " \ + /etc/nginx/sites-enabled/ | head -40 && \ +echo "=== certificate issuance ===" && \ +ls /etc/letsencrypt/live 2>/dev/null || echo "no certbot live dir" ; \ +which certbot ; \ +echo "=== does the name resolve yet ===" && \ +dig +short dev.mechcomp.kane-il.us A ; \ +dig +short dev.mechcomp.kane-il.us AAAA +``` + +What matters in the output: whether the peer's `AllowedIPs` is the only gate, +whether nginx already has `listen [::]:443` on its vhost pattern (it must, if +the name inherits an AAAA), how certificates are issued, and whether the name +resolves yet. + +**Report this before proceeding.** Groups 2 onward are written against what it +shows; they are deliberately not drafted in advance. + +### Group 2 — WireGuard AllowedIPs + +Drafted after group 1. The shape: + +- Back up `/etc/wireguard/wg0.conf` first. `PROCESS.md` §2: preserve a rollback + copy before editing any configuration file. +- Add `10.20.0.0/24` to the `srv-b` peer's `AllowedIPs`, **keeping + `10.110.0.0/22`**. Replacing rather than extending it is the way this breaks. +- Apply live, confirm the tunnel is still up and `srv-b` still reachable, then + persist. +- Prove: `ping -c2 10.20.0.10` from the hub. + +### Group 3 — route + +A route for `10.20.0.0/24` via `10.110.0.12`, persisted the way the hub already +persists routes — which group 1 reveals. Prove with +`curl -sS -o /dev/null -w '%{http_code}\n' http://10.20.0.10:8770/`. + +### Group 4 — certificate + +Let's Encrypt HTTP-01 for `dev.mechcomp.kane-il.us`. A third-level name needs no +wildcard and no DNS-01. Requires the name to resolve to the hub first. + +### Group 5 — vhost + +```nginx +# /etc/nginx/sites-available/dev.mechcomp.kane-il.us +# Matching whatever pattern group 1 shows the existing vhosts use. +server { + listen 80; + listen [::]:80; + server_name dev.mechcomp.kane-il.us; + return 301 https://$host$request_uri; +} + +server { + listen 443 ssl; + listen [::]:443 ssl; # required if the name has an AAAA + server_name dev.mechcomp.kane-il.us; + + ssl_certificate /etc/letsencrypt/live/dev.mechcomp.kane-il.us/fullchain.pem; + ssl_certificate_key /etc/letsencrypt/live/dev.mechcomp.kane-il.us/privkey.pem; + + location / { + proxy_pass http://10.20.0.10:8770; + proxy_set_header Host $host; + proxy_set_header X-Real-IP $remote_addr; + proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for; + proxy_set_header X-Forwarded-Proto $scheme; + + # The composer redraws on every control change. Small responses, many of + # them; buffering adds latency for no benefit. + proxy_buffering off; + proxy_read_timeout 30s; + } +} +``` + +`nginx -t` before reload, always. + +### Group 6 — acceptance + +All six criteria in §2, including the two negatives. + +--- + +## 4. Known unknowns + +Stated rather than assumed, because assuming is what produced three unreachable +URLs before this document existed. + +- **Whether the hub's nginx listens on v6.** If `dev.mechcomp` is a CNAME to a + name carrying an AAAA, v6 is published whether or not nginx answers on it. + v6-preferring browsers then fail while v4 ones succeed — a fault that looks + like anything except DNS. +- **How the hub persists routes.** Distribution- and tooling-dependent; group 1 + shows it. +- **Whether Webmin or YunoHost manages the nginx configuration.** If so, + hand-written vhosts may be overwritten on their next reconfiguration, and the + vhost belongs wherever that system expects it instead. + +--- + +## 5. Acceptance and reporting + +Closed when all six criteria in §2 hold and survive a hub reboot. + +Back to the architect: the output of every group, and any surprise about the +environment in `FAILURES.md` format — `PROCESS.md` §7. A surprise on the hub is +still a finding even though the hub is not this project's property. + +`STAGING-STATE.md` §4 records "direct WireGuard-side access to the catalogue +would be a route addition for `10.20.0.0/24` on the hub. Not now." That decision +was correct when there was no application to reach. This work order is its +reversal, and §4 should say so when this closes.