diff --git a/docs/STAGING-STATE.md b/docs/STAGING-STATE.md index 9e913e3..8917edc 100644 --- a/docs/STAGING-STATE.md +++ b/docs/STAGING-STATE.md @@ -4,7 +4,7 @@ Live state of the Mechanical Compiler staging instance on `srv-b`. | | | |---|---| -| Updated | 2026-09-11, after the composer replaced the placeholder | +| Updated | 2026-09-11, after the composer was published at `dev.mechcomp.kane-il.us` | | Instance | Staging / development | | Specification | `ENVIRONMENT.md` revision 5 | | Failure log | `FAILURES.md` | @@ -465,6 +465,49 @@ so SSH adds no privilege — but **every path to a container runs through Direct WireGuard-side access to the catalogue would be a route addition for `10.20.0.0/24` on the hub. Not now. +**Superseded 2026-09-11. A public path now exists, and it is not a route.** + +``` +browser + -> dev.mechcomp.kane-il.us A 198.58.111.109 + AAAA 2600:3c00::f03c:92ff:fe42:43d7 + -> nginx on wg-pk, Let's Encrypt TLS, expires 2026-12-10 + -> proxy_pass http://10.110.0.12:8770 (srv-b's own tunnel address) + -> DNAT on srv-b -> 10.20.0.10:8770 + -> mechcomp.service in CT 100 +``` + +**No WireGuard change was made and no route was added.** The hub's peer entry +for `srv-b` is still `AllowedIPs = 10.110.0.12/32`, as all twenty peers are. +An earlier plan widened it to carry `10.20.0.0/24`; that was abandoned once the +hub's own convention was read. Every existing vhost there -- +`witness.diagnostics.kane-il.us`, `otium.civicus.us`, `shell.infra.civicus.us`, +`corpusdb.infra.civicus.us` -- proxies to a `10.110.0.x` tunnel address +directly. Following that pattern removed the only step with a lockout risk. + +The rule on `srv-b`, in `nat PREROUTING`, persisted: + +``` +-A PREROUTING -s 10.110.0.1/32 -d 10.110.0.12/32 -i wg0 \ + -p tcp -m tcp --dport 8770 -j DNAT --to-destination 10.20.0.10:8770 +``` + +**Scoped to source `10.110.0.1` only** -- the hub. The other nineteen tunnel +peers cannot reach the service network through it. Widening that is one field. + +Rollback copy at `/etc/iptables/rules.v4.before-mechcomp-dnat`. `POSTROUTING` +order was verified unchanged after `iptables-save` rewrote the file: the +`RETURN` still precedes both masquerades. + +**`curl http://10.110.0.12:8770` from `srv-b` itself fails, and that is +correct.** Locally-originated traffic traverses `OUTPUT`, not `PREROUTING`, so +it never meets the rule. Test from the hub. + +Proven 2026-09-11: `https 200` over v4 and v6, `http` returning 301, real build +JSON through the full chain, and the mail path unaffected -- a test message +from `srv-b` was delivered to `theron@kane-il.us` via `wg-pk` and `mx1` after +the ruleset was rewritten. + **Reopened 2026-09-11 by `WORK-ORDER-004`.** That decision was correct while there was no application to reach. There is one now, and the consequence of the decision is that CIVICVS cannot see it: `mechanical-compiler.dev.infra` @@ -508,9 +551,12 @@ provisioning: is still a stub - [ ] Reference toolchain image `mechcomp/reference-toolchain:8.0.0` - [ ] Fixture reproduction against `ddd0f154...` -- [ ] Application-runtime acceptance — partial. The composer serves and is - reachable through CT 101; no acceptance criteria have been written for - it, and it is not reachable from outside (see section 6, question 4) +- [ ] Application-runtime acceptance — partial. The composer serves, is + reachable through CT 101 internally and at + `https://dev.mechcomp.kane-il.us` publicly, and no acceptance criteria + have been written for it. Certificate renewal is unattended via certbot + and **has not yet been observed to succeed** for this name; first + renewal is due before 2026-12-10. - [x] ~~Replacing the placeholder with the real service~~ — done 2026-09-11 **Applied 2026-08-18:** `venv/` is in `.gitignore`, along with `.cache/`, @@ -529,7 +575,8 @@ The Shapely port gates all of the above. | 1 | Should `wg-pk` `mynetworks` narrow to explicit hosts? | F-025 estate half | | 2 | What is the backup strategy? | all backup work | | 3 | Where is the 3+ TB USB disk attached? | gold redundancy step | -| 4 | Public ingress for the composer at `dev.mechcomp.kane-il.us` | `WORK-ORDER-004`; anyone outside `srv-b` seeing the application at all | +| 4 | ~~Public ingress for the composer~~ | **Closed 2026-09-11.** Live at `https://dev.mechcomp.kane-il.us`; see §4 | +| 5 | `shell.infra.civicus.us` and `corpusdb.infra.civicus.us` publish AAAA `2600:3c00:e000:365::`, but `wg-pk` holds `2600:4c00:e000:365::` -- one hex digit apart. Both names are broken for v6-preferring clients. | CIVICVS, estate; not this project | All three need CIVICVS. diff --git a/docs/WORK-ORDER-004-public-ingress.md b/docs/WORK-ORDER-004-public-ingress.md index abbf74f..5e54235 100644 --- a/docs/WORK-ORDER-004-public-ingress.md +++ b/docs/WORK-ORDER-004-public-ingress.md @@ -1,226 +1,140 @@ # WORK-ORDER-004 — public ingress for the composer -Publish the Mechanical Compiler composer at `dev.mechcomp.kane-il.us`, reachable -from an ordinary browser on the internet. +**CLOSED 2026-09-11.** The composer is live at `https://dev.mechcomp.kane-il.us`. | | | |---|---| | Mode | Infrastructure (`PROCESS.md` §2) | | Created | 2026-09-11 | -| Executed on | `wg-pk` (the WireGuard hub), via its Webmin terminal | -| Depends on | `14514b0` — composer live on CT 100, proven through CT 101 | +| Closed | 2026-09-11 | +| Executed on | `wg-pk` and `srv-b`, by CIVICVS | --- -## 0. Read this before the first command +## 0. This document was rewritten mid-execution -**This work order is executed on a machine outside this project.** `wg-pk` -carries `kane-il.us` mail and Hubzilla. `PROCESS.md` §7 requires escalation for -exactly this, and the escalation is: CIVICVS executes, one group at a time, and -anything surprising stops the work rather than being worked around. +The original proposed widening `AllowedIPs` on the hub's peer entry for `srv-b` +to carry `10.20.0.0/24`, adding a route, and warned at length about the lockout +risk of changing a live WireGuard peer. -### The lockout risk, and why it is bounded +**None of that was done, and none of it was necessary.** It was written before +reading the hub, from an assumption about how the estate must be wired. -Step 2 changes a live WireGuard peer. If it goes wrong the tunnel drops, and -with it every path to `srv-b` — `PROCESS.md` §1 records that the `srv-b` shell -is the operator's entire working surface for this project. +Reading it showed the convention: every vhost on `wg-pk` — +`witness.diagnostics.kane-il.us`, `otium.civicus.us`, `shell.infra.civicus.us`, +`corpusdb.infra.civicus.us` — proxies to a `10.110.0.x` tunnel address directly. +Not one routes into a subnet behind a peer. All twenty peers carry a `/32`. -It is recoverable because **the operator is sitting on `wg-pk` itself**, in -Webmin, not reaching it through the tunnel. `wg set` is not persistent, so -`systemctl restart wg-quick@wg0` on the hub restores the on-disk configuration -and the tunnel with it. +Following that convention removed the WireGuard change entirely, and with it the +only step that could have locked the operator out of `srv-b`. -Do not perform step 2 from a shell that reaches the hub through the tunnel. - -### What `srv-b` contributes - -**Nothing. It is already correct and must not be touched.** Verified -2026-09-11: - -``` -net.ipv4.ip_forward = 1 --P FORWARD ACCEPT -10.20.0.10 dev vmbr1 src 10.20.0.1 -``` - -The three `FORWARD` rules block container-sourced SMTP, container-to-container -(ACCEPT), and container-to-LAN. None matches hub-initiated inbound traffic. The -`-s 10.20.0.0/24 -o wg0 MASQUERADE` rule does not apply either: the first packet -of a hub-initiated flow is not container-sourced, so no NAT binding is created -and replies return through conntrack. - -**Explicit do-not-touch list:** `srv-b` iptables, `srv-b` WireGuard, CT 101 -nginx, CT 101 TLS, the `kane-il.us` MX records, anything on -`mx1.diagnostics.kane-il.us` (it publishes a TLSA record — confirmed -2026-09-11). +The original text is in git history at `a8081e1`. It is kept there rather than +here, because a work order describing a plan nobody executed is exactly the stale +record this project does not tolerate. --- -## 1. Topology, and why +## 1. What was actually built ``` browser - -> DNS dev.mechcomp.kane-il.us -> wg-pk public address - -> nginx on wg-pk, Let's Encrypt TLS terminated here - -> WireGuard tunnel, already encrypted - -> srv-b 10.110.0.12, forwards, no configuration change - -> CT 100 10.20.0.10:8770, the composer + -> dev.mechcomp.kane-il.us A 198.58.111.109 + AAAA 2600:3c00::f03c:92ff:fe42:43d7 + -> nginx on wg-pk, Let's Encrypt TLS, expires 2026-12-10 + -> proxy_pass http://10.110.0.12:8770 srv-b's own tunnel address + -> DNAT on srv-b -> 10.20.0.10:8770 + -> mechcomp.service in CT 100 ``` -**The public path deliberately bypasses CT 101.** Routing it through CT 101 -would make the public name depend on a locally-signed leaf valid to 2028-11-18 -with nothing renewing it — a dated outage designed in from the start. The tunnel -already provides the encryption that hop would add. +Four changes, in the order made: -CT 101 continues to serve `mechanical-compiler.dev.infra` for work from inside. -Two ingresses, each with a distinct reason to exist. +**1. DNAT on `srv-b`**, added live, persisted only after being proven from the +hub. Rollback copy at `/etc/iptables/rules.v4.before-mechcomp-dnat`. ---- - -## 2. Success criteria, stated before the work - -1. `curl -I https://dev.mechcomp.kane-il.us/` returns `200` from a machine with - no WireGuard access and no special DNS. -2. The certificate is issued by Let's Encrypt and chains without `-k`. -3. `https://dev.mechcomp.kane-il.us/api/build?family=3x&profile=Y` returns JSON - whose `groups` end with a `Y` group and contain no `three_fin` key. -4. The page renders, controls change the drawing, and a changed control changes - the `design` identity shown beneath it. -5. **The negative:** `mechanical-compiler.dev.infra` still answers from `srv-b`, - and mail from `srv-b` still delivers. Neither path was in scope; both must be - proven unharmed. -6. Every change survives `systemctl restart wg-quick@wg0` and a hub reboot. - ---- - -## 3. Command groups - -One at a time. Paste output back before the next. - -### Group 1 — read-only, on `wg-pk` - -Nothing here changes anything. - -```bash -echo "=== peer entry for srv-b ===" && \ -wg show && \ -echo "=== on-disk wireguard config ===" && \ -grep -n "AllowedIPs\|PublicKey\|Address\|PostUp" /etc/wireguard/wg0.conf && \ -echo "=== routing toward the service network ===" && \ -ip route | grep -E "10\.20\.|10\.110\." ; \ -echo "=== forwarding ===" && \ -sysctl net.ipv4.ip_forward && \ -echo "=== nginx vhost conventions, following symlinks ===" && \ -grep -rn --dereference-recursive "server_name\|proxy_pass\|listen\|ssl_certificate " \ - /etc/nginx/sites-enabled/ | head -40 && \ -echo "=== certificate issuance ===" && \ -ls /etc/letsencrypt/live 2>/dev/null || echo "no certbot live dir" ; \ -which certbot ; \ -echo "=== does the name resolve yet ===" && \ -dig +short dev.mechcomp.kane-il.us A ; \ -dig +short dev.mechcomp.kane-il.us AAAA +``` +-A PREROUTING -s 10.110.0.1/32 -d 10.110.0.12/32 -i wg0 \ + -p tcp -m tcp --dport 8770 -j DNAT --to-destination 10.20.0.10:8770 ``` -What matters in the output: whether the peer's `AllowedIPs` is the only gate, -whether nginx already has `listen [::]:443` on its vhost pattern (it must, if -the name inherits an AAAA), how certificates are issued, and whether the name -resolves yet. +Scoped to the hub as source. The other nineteen tunnel peers cannot reach the +service network through it. -**Report this before proceeding.** Groups 2 onward are written against what it -shows; they are deliberately not drafted in advance. +`curl http://10.110.0.12:8770` from `srv-b` itself fails, correctly: +locally-originated traffic traverses `OUTPUT`, not `PREROUTING`. Test from the +hub. -### Group 2 — WireGuard AllowedIPs +**2. DNS.** `A` and `AAAA` for `dev.mechcomp.kane-il.us`, TTL 300. Not a CNAME — +the hub's other names are all A/AAAA, and matching the zone's own convention beat +the marginal tidiness of an alias. The `AAAA` was safe to publish immediately +because `witness` and `otium` already carry `listen [::]:443 ssl`. -Drafted after group 1. The shape: +Nothing was created at `mechcomp.kane-il.us`: CIVICVS never serves at the root of +a subdomain. -- Back up `/etc/wireguard/wg0.conf` first. `PROCESS.md` §2: preserve a rollback - copy before editing any configuration file. -- Add `10.20.0.0/24` to the `srv-b` peer's `AllowedIPs`, **keeping - `10.110.0.0/22`**. Replacing rather than extending it is the way this breaks. -- Apply live, confirm the tunnel is still up and `srv-b` still reachable, then - persist. -- Prove: `ping -c2 10.20.0.10` from the hub. +**3. Port-80 vhost on `wg-pk`**, no TLS. HTTP-01 needs a live vhost to validate +against, and a `listen 443` block naming certificate paths that do not exist yet +fails `nginx -t` and takes the reload down with every other site. -### Group 3 — route - -A route for `10.20.0.0/24` via `10.110.0.12`, persisted the way the hub already -persists routes — which group 1 reveals. Prove with -`curl -sS -o /dev/null -w '%{http_code}\n' http://10.20.0.10:8770/`. - -### Group 4 — certificate - -Let's Encrypt HTTP-01 for `dev.mechcomp.kane-il.us`. A third-level name needs no -wildcard and no DNS-01. Requires the name to resolve to the hub first. - -### Group 5 — vhost - -```nginx -# /etc/nginx/sites-available/dev.mechcomp.kane-il.us -# Matching whatever pattern group 1 shows the existing vhosts use. -server { - listen 80; - listen [::]:80; - server_name dev.mechcomp.kane-il.us; - return 301 https://$host$request_uri; -} - -server { - listen 443 ssl; - listen [::]:443 ssl; # required if the name has an AAAA - server_name dev.mechcomp.kane-il.us; - - ssl_certificate /etc/letsencrypt/live/dev.mechcomp.kane-il.us/fullchain.pem; - ssl_certificate_key /etc/letsencrypt/live/dev.mechcomp.kane-il.us/privkey.pem; - - location / { - proxy_pass http://10.20.0.10:8770; - proxy_set_header Host $host; - proxy_set_header X-Real-IP $remote_addr; - proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for; - proxy_set_header X-Forwarded-Proto $scheme; - - # The composer redraws on every control change. Small responses, many of - # them; buffering adds latency for no benefit. - proxy_buffering off; - proxy_read_timeout 30s; - } -} -``` - -`nginx -t` before reload, always. - -### Group 6 — acceptance - -All six criteria in §2, including the two negatives. +**4. `certbot --nginx -d dev.mechcomp.kane-il.us`**, run once. HTTP-01 through +the nginx plugin, matching all four existing certificates — verified by reading +`/etc/letsencrypt/renewal/*.conf` rather than assuming. No DNS plugin is +installed and BIND was not involved. Certbot wrote the 443 block, the redirect +and the certificate lines into the same file. --- -## 4. Known unknowns +## 2. Acceptance — all met 2026-09-11 -Stated rather than assumed, because assuming is what produced three unreachable -URLs before this document existed. +| | Criterion | Result | +|---|---|---| +| 1 | `https://dev.mechcomp.kane-il.us/` from outside | `200` | +| 2 | Let's Encrypt chain, no `-k` | issued, expires 2026-12-10 | +| 3 | `/api/build` returns real JSON | confirmed | +| 4 | Reachable over IPv6 | `200` | +| 5 | HTTP redirects rather than serving plaintext | `301` | +| 6 | **Negative:** `mechanical-compiler.dev.infra` still answers | `200` | +| 7 | **Negative:** CT 100 direct still answers | `200` | +| 8 | **Negative:** mail from `srv-b` still delivers | delivered to `theron@` via `wg-pk` and `mx1` | -- **Whether the hub's nginx listens on v6.** If `dev.mechcomp` is a CNAME to a - name carrying an AAAA, v6 is published whether or not nginx answers on it. - v6-preferring browsers then fail while v4 ones succeed — a fault that looks - like anything except DNS. -- **How the hub persists routes.** Distribution- and tooling-dependent; group 1 - shows it. -- **Whether Webmin or YunoHost manages the nginx configuration.** If so, - hand-written vhosts may be overwritten on their next reconfiguration, and the - vhost belongs wherever that system expects it instead. +The three negatives matter as much as the positives. `iptables-save` rewrote the +whole ruleset on `srv-b`, and mail traverses the same `wg0` interface the DNAT +was added to. `POSTROUTING` order was verified unchanged — the `RETURN` still +precedes both masquerades. --- -## 5. Acceptance and reporting +## 3. What this did not settle -Closed when all six criteria in §2 hold and survive a hub reboot. +- **Renewal has never been observed to succeed for this name.** The certbot timer + is scheduled and the other four certificates are managed identically, but the + first renewal for `dev.mechcomp` is due before 2026-12-10 and nobody has + watched one complete. `certbot renew --dry-run` counts against a rate limit and + was not repeated after one attempt timed out. +- **No acceptance criteria exist for the composer itself**, only for the path to + it. A `200` says the chain works, not that the page is right. +- **The service has no authentication.** It is world-reachable and computes + geometry for anyone who asks. Acceptable for a development name; it should be a + conscious decision before anything else is published this way. -Back to the architect: the output of every group, and any surprise about the -environment in `FAILURES.md` format — `PROCESS.md` §7. A surprise on the hub is -still a finding even though the hub is not this project's property. +--- -`STAGING-STATE.md` §4 records "direct WireGuard-side access to the catalogue -would be a route addition for `10.20.0.0/24` on the hub. Not now." That decision -was correct when there was no application to reach. This work order is its -reversal, and §4 should say so when this closes. +## 4. Found while working, not this project's to fix + +`shell.infra.civicus.us` and `corpusdb.infra.civicus.us` publish `AAAA` +`2600:3c00:e000:365::`. `wg-pk` holds `2600:4c00:e000:365::` — one hex digit +apart, `3c` against `4c`. **Both names are unreachable for v6-preferring clients +right now**, while v4 clients see working sites: the intermittent fault that +looks like anything except DNS. + +Recorded in `STAGING-STATE.md` §6 so it does not evaporate with the scrollback. + +--- + +## 5. Method note + +The one thing that made this go quickly, after several turns of it not going +quickly at all: **read the estate's own conventions before designing against +it.** Every wrong turn in this work order's history — a route that was not +needed, a lockout risk that did not exist, a CNAME where the zone uses A/AAAA, a +DNS-01 challenge where HTTP-01 was already standard — came from proposing a +design before reading what four working services already did.